diff --git a/CLAUDE.md b/CLAUDE.md index a12e2f044..4afed0117 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -2233,6 +2233,306 @@ a `jvm.config` takes no comments, so REUSE can only read its metadata from that `REUSE Compliance Check` job fails on `main` — which is how it was found, the PR run having been cancelled. Spotless (palantir) is configured in its own pom; the model-free CI job runs `spotless:check`. +**The REPL layer (commands, approval, status line, rendering).** Eight small classes and one dependency +(`org.jline:jline`, one jar, no transitive deps): +`SlashCommands` (a line starting with `/` whose first word names a command is handled locally — +`/help /status /tools /mode /compact /clear /exit`; **an unknown `/command` goes to the model**, which +is why no escape syntax is needed for `/usr/bin/…`), `ApprovalMode` + `ConsoleApprovalStrategy` +(`[y]es/[n]o/[a]uto` per gated call), `StatusLine`, and `Ansi` + `MarkdownConsole`. Four points that +are decisions, not details: + +1. **The approval gate is Atmosphere's, not ours.** `AgentRunner.approval(strategy, policy)` attaches + `ToolApprovalPolicy.custom(...)` (gating `run_command`, `write_file`, `edit_file`, `delete`, + `rename` — reading tools never ask) and a `ConsoleApprovalStrategy`; `ToolExecutionHelper` then + blocks the tool loop before the executor runs and turns a denial into the tool result + `{"status":"cancelled","message":"Action cancelled by user"}` for the model. Do not reimplement + that message. `ApprovalWireTest` pins both halves over the real server. **The reading tools + (`ls`, `read_file`, `glob`, `grep`) never ask, and that is a decision, not an omission**: gating + them would make the question so frequent it stops being read. Because the gate is a list of + *names*, a tool upstream adds or renames would drop out of it and then run unasked — so + `READ_ONLY_TOOLS` names the other half explicitly and + `ConsoleApprovalStrategyTest.everyOfferedToolIsEitherGatedOrDeclaredReadOnly` asserts every offered + tool is in exactly one of the two sets, and that neither set names a tool nobody offers. Same class + as the stale `spotbugs-exclude.xml` entries: an allowlist that silently stops matching. +2. **One-shot (`--prompt`) denies a gated call** instead of auto-approving it — `--auto` is the + deliberate opt-in. Atmosphere itself fails closed when no strategy is wired, and this keeps that + direction: an unattended run must not be the most permissive one. +3. **`AgentTerminal` has exactly two implementations, chosen once at startup, and `--plain` picks the + line-oriented one on purpose.** `LocalAgent.usesFullTerminal(options, interactive)` is the single + place that decides: the cursor-controlling console needs someone typing **and** permission to move + the cursor, and `--plain` withholds the second even on a real terminal. That is not a fallback but + a supported mode — for a session that is piped, logged, recorded, or carried by something that + forwards lines rather than a screen. It gives up the pinned block, the spinner, history/completion + and typing-during-a-turn (`PlainTerminal.hasPendingInput()` is always false), and gains being + correct when the output is a file. A normal SSH session needs none of this: a remote terminal + reports its size and handles cursor control like a local one. `JLineTerminal` (a real + terminal: line editing, history, Tab completion of the command names, a status line pinned to the + bottom via JLine's `Status`, single-key answers through `enterRawMode`, streamed output via + `LineReader.printAbove` so the bottom block stays put) and `PlainTerminal` (a `PrintStream` plus a + `BufferedReader`: no cursor control at all, correct when the output is a file). `JLineTerminal.open` + returns **null** instead of throwing when there is no usable terminal — piped input, a dumb + terminal, a missing native provider — and the caller falls back. Every test drives `PlainTerminal`, + which is why none of them needs a TTY. Verified on Windows: JLine picks the `windows-vtp` provider, + so ANSI works there without the registry caveat. +4. **The answer is rendered append-only, one completed line at a time** (`MarkdownConsole`). Redrawing + on every token is what produces the known overdraw/truncation bugs in the Ink/Bubble-Tea based + clients and breaks when the output is piped. Only headings, bullets, fences and inline + `**bold**`/`` `code` `` are handled; italics deliberately are not (`*` is more often a glob than + emphasis). Colour is decided once in `Ansi.detect()` — `CLICOLOR_FORCE`, then `NO_COLOR`, then + `TERM=dumb`/`CLICOLOR=0`, else "is a terminal" via `Console.isTerminal()` (reflective: JDK 22+; + below that `System.console() != null`). +5. **`/loop` keeps its state in a file, not in the context, and stops on a text marker** (`TaskLoop`, + `LoopOptions`). Every step re-sends the task verbatim and drops the history, so the context cannot + grow (Claude Code's ralph-wiggum plugin does the same, and `AGENT-LOOP.md` in the workspace is the + memory). The stop signal is a line that is **exactly** `<>`, never a substring — + deliberately **not** a "done" tool: below 7B a model emits a malformed tool call far more often + than a malformed line, and mini-SWE-agent's SWE-bench results come from exactly this plain-sentinel + design. `--check ''` re-verifies the claim and feeds a failure back. Four guards, none of them + trusted to the model: step cap, wall-clock budget, stall detection (three steps with no file change + and no tool call), interval. **Order matters and a test pins it**: the marker is checked *before* + the stall detector, because the step that only answers "done" changes nothing and would otherwise + be reported as no progress. +6. **Three of the file tools are this project's, not Atmosphere's** (`WorkspaceTools`, `TextEdits`, + `WorkspaceSearch`) — `read_file`, `edit_file`, `grep`. They are **replacements, not additions**: + two tools that both claim to read a file is the worst case for tool selection. They call the same + `AgentFileSystem`, so workspace confinement, path validation and size limits stay Atmosphere's. + Each replacement has a measured reason, and all three are pinned by tests: + - `edit_file`: the framework matches against the **raw** file content, so a model's LF text never + matches a CRLF file — on Windows *every* edit fails silently. `TextEdits` normalizes before + matching and restores the file's own ending and byte-order mark. It also shows the nearest lines + on a miss and the line numbers on an ambiguous match (a failed edit drops the eventual success + rate from 90.5 % to 57.2 %), offers `replace_all`, and applies a batch of edits **all-or-nothing** + — a deviation from every shipping agent, which apply sequentially and leave a half-edited file. + - `grep`: the framework walks **alphabetically** with one global 2-second deadline and one global + 500-hit budget, so `.git`, `target` and `node_modules` consume both before `src` is reached. + `WorkspaceSearch` excludes them, groups by file with line numbers, and **states** truncation. + - `read_file`: `offset`/`limit` and numbered lines (whole-file reads measure 12.7 % against 18.0 % + task success in the SWE-agent ablations). The numbers are display only, which both the tool + description and the system prompt say — leaked line numbers in `old_string` are a known failure. + - `edit_file` **refuses a file that was not read** in this session (`WorkspaceTools.ReadTracker`). + Not a staleness check: an exact unambiguous match is safe regardless; this catches the model + inventing the text. + **Rejected on evidence, do not add later without new numbers:** a unified-diff/patch tool (Meta's + ablation: search-replace 42–53 % vs 26–30 % unified diff vs 20–26 % line diff on one model; a 7B + model collapses 54 → 33 → 14 %), fuzzy matching (turns a loud miss into a silent wrong-place edit), + an embedding index (Cursor's production effect is +0.3 %), and LSP tools (the one isolation study + finds them token-negative and *worse* at multi-file rename, because renames touch comments and + strings that semantic references exclude). +7. **The session transcript (`Transcript`) is not the conversation the model is sent, and must not be + merged with it.** The model's history is rewritten by `/compact` — a summary replaces the turns — + and has never carried a timestamp; the transcript only grows and stamps every entry. `/compact` + adds a note to it and changes nothing else, `/clear` empties it (the command means "forget this + session"), `/save [name]` writes it into the workspace, and `--transcript ` appends live so a + killed session still leaves what it had — that write failing is swallowed, because a record that + exists to survive a bad ending may not cause one. **A list, not a map keyed by the timestamp**: a + tool result and the answer after it regularly share a millisecond and a map would drop one + silently; insertion order already is time order. `ToolCallLog` stays as the separate, hard-cut + receipt for `/calls` — it answers "did that really run", which prose cannot. + **`/load` replays only `USER` and `AGENT` entries** as messages: a tool result outside its round is + not something a chat template has a place for, and inventing a shape for it would be worse than + letting the model call the tool again. The parser treats a line without a stamp as a continuation + of the entry above it, because an entry is not a line — an answer keeps its newlines when written, + and reading line by line would turn one answer into several. A file that is not a transcript yields + **no** entries rather than one wrong one, since anything it yielded would be replayed to the model + as if it had been said. + +8. **Tool calls are carried into the conversation as a text note, and logged for `/calls`.** + `LocalAgent.withToolNotes` prefixes each turn's answer in the history with + `(tools I actually ran this turn: -> )`, and `ToolCallLog` + keeps the same data for the `/calls` command. **Why it is a note and not real `tool_calls` + messages:** `AbstractAgentRuntime.assembleMessages` rebuilds every history entry as + `new ChatMessage(h.role(), h.content())` — the tool-call array and the tool-call id never leave the + framework, so protocol-faithful replay through `context.history()` is impossible; content is what + survives. **Why it exists at all:** with only user text and assistant prose in the history, a 4B + model stopped calling tools after the third turn of a real session and *described* the work instead + — inventing JUnit tests, a Maven build and a `.bat` script, complete with exit codes, while the + workspace stayed empty. The system prompt also forbids claiming an action without the call. + **Placement was found by failing twice, so do not "simplify" it:** in front of the assistant's + answer made the model copy the record into its own replies (the user saw `(tools I actually ran + this turn: …)` as the first line of an answer); real `tool_calls` messages are impossible (see + above); a mid-history system message is cleanest but Mistral's template requires strict + user/assistant alternation and Gemma has no system role. It therefore rides in front of the **next + user message**, which every template accepts. + Pinned by `LocalAgentTest.aToolCallStaysInTheHistorySoTheNextTurnSeesItHappened`. + **Live feedback while a turn runs** (`LocalAgent.activityLine`, `ShellTool`'s line-by-line output): + the block's first row names the running tool and its own elapsed time, and shell output is printed + as it arrives. **The turn runs on its own thread, and it has to:** Atmosphere's `execute()` is + synchronous — it returns only once the whole turn including every tool round is done — so running + it on the console thread leaves nobody to refresh the line, and the block sits on + "… waiting for input …" for the entire turn (exactly the symptom that was reported). The approval + prompt then reads a key in raw mode on that worker thread while the console thread redraws four + times a second, so `TurnActivity` pauses the redraw for as long as the question is open. Reading the pipe incrementally is not only cosmetic — an unread pipe blocks the child once + it is full, which on Windows is roughly 4 KB. +9. **The context number in the status line is an estimate, marked `~`, and it moves during the turn.** + llama.cpp emits its usage chunk only when the client sets `stream_options.include_usage`, and + Atmosphere's client does not; `ConsoleSession.usage()` takes the real count when one arrives, + otherwise `LocalAgent.estimateTokens` uses four characters per token. The window size is + `--ctx-size` (in-process) or the server's `/props` (`ServerProps`), and is omitted rather than + guessed when neither answers. **The state row is a `Function`, not a + string**, and `awaitWithActivity` asks it again on every redraw: it used to be rendered once before + the turn and handed over fixed, so the figure stood still through every tool round and only moved + at the next `you>` — which is exactly when it no longer helps anyone decide whether to `/compact`. + `LocalAgent.liveTokens` adds `ConsoleSession.producedChars()` (streamed text **plus** every tool + call and result — all of it is in the prompt of the next model call of the *same* turn) to what the + request carried when it was sent, and yields to the server's own count as soon as one arrives. + `TaskLoop` passes a constant function, and its step label must be copied into a local first: a + lambda may not close over the loop counter. +10. **One call to `AgentTerminal.line` is one screen line**, and `ConsoleSessionTest` is what defends + it. The pinned block is reserved in **lines**, so a single "line" carrying twenty newlines moves the + screen twenty rows further than the terminal accounted for and the block is then drawn across the + output — reported twice, both times from a `write_file` call whose `content` argument was the file. + `ConsoleSession.describeArguments` folds and cuts **each argument value on its own** (80 chars) + before cutting the whole rendering (200), so a call carrying a whole file still shows the file + *name*; results and errors are folded the same way. `JLineTerminal.line` splits a multi-line string + as a backstop for a caller that forgets. Only the console is cut — the model gets everything, and + `ConsoleSession.rounds()` keeps the full arguments for the history note and `/calls`. +11. **The approval mode carries a glyph, and shift+tab switches it**: `ApprovalMode.symbol()` / + `badge()` render `⏸ manual` and `⏵⏵ auto` on the status line and in `/mode`, the transport symbols + the established terminal agents use for the same distinction; `ApprovalMode.next()` is the cycle + the key walks. The binding is `AgentTerminal.onCycleMode(Runnable)`, which **defaults to declining** + — only `JLineTerminal` overrides it, and the startup line advertises the key only when the bind + succeeded. Two details are not obvious: it is bound **both** through terminfo + (`InfoCmp.Capability.key_btab`) **and** to the literal `ESC [ Z`, because JLine's + `windows-vtp.caps` declares no `key_btab` at all while the terminal in virtual-terminal input mode + does send the sequence; and the shortcut fires **only while a line is being read**, so it switches + the mode between turns — which is when it is decided anyway. The REPL therefore holds the two + status numbers in an `AtomicLong`/`AtomicBoolean` rather than locals, so the widget (which runs + inside the reader) can re-render the pinned row with what the last turn left behind. + +12. **The input is framed into the pinned block, and the prompt stays there during a turn; typing + stops the turn.** The frame is two halves that must be read together: the **top** rule is the first + line of the reader's *prompt* (`rule() + newline + "> "`, rebuilt on every read because the window + can be resized) — **no, and the second attempt was wrong too.** Both are recorded because the + obvious fix is the one that fails. (1) Rule as the first line of a **two-line prompt**: + `ERASE_LINE_ON_FINISH` erases exactly **one** line, so every Enter leaves the rule behind and + holding Enter draws a column of them. (2) Rule as **ordinary output before each read**: nothing is + left behind on Enter any more, but one rule now stays in the scrollback per turn and travels up + with it. **There is no third option** — JLine's status region is below the prompt and never above + it, so a rule above the input can only be part of the prompt (1) or part of the scrollback (2). + The settled shape is therefore **one** rule, the first line of the status block, directly under the + input line; the prompt is `"> "`, one line, and `AgentTerminal.readLine`'s `prompt` argument is + consequently **ignored** here. + `ERASE_LINE_ON_FINISH` removes the input line on Enter and the reader thread echoes it above as + `› text`, so the transcript keeps what was asked. + + **`JLineTerminalTest` is how any of this is checkable**: `JLineTerminal.over(Terminal, …)` takes a + terminal built over two streams, which renders exactly like a TTY, so the screen can be asserted on + the emitted bytes. Two things that cost an hour each and are not guessable: the test terminal needs + **`stdoutEncoding`** as well as `encoding`, or every `─` arrives as `?`; and a box character is + written as UTF-8 from `printAbove` but as the **DEC line-drawing set** (`ESC(0` + `q`s + `ESC(B`) + inside a *prompt*, so a counter that looks only for `─` passes against the exact bug it was + written for — verified by putting the two-line prompt back and watching the test go from 1 rule to 5. + + **Escape sequences drawn as text** (`[?1h` above the prompt, then a `1H` inside the rule) were two + further defects of the same family, and the second is the one that eventually **destroyed the + block**. First: `line()` wrote straight to the terminal whenever the reader was not inside + `readLine`, which is exactly when the next read emits its init sequence — once the reader thread + exists, **everything** now goes through `printAbove`. Second: three threads write to this terminal + as a matter of course — the turn (Atmosphere's thread) prints tool lines, the console thread + refreshes the block four times a second, and with an in-process model **llama.cpp logs to stderr**, + which is the same console and goes around JLine entirely. The first two are serialised by a + `writing` lock held across `line()` and `status()`; the third is fixed by + `LocalAgent.captureNativeLog`, which routes the native log through `LlamaModel.setLogger` + (the callback sink `patches/0014` added) into `terminal.line`, so it scrolls in above the prompt + like any other output instead of scrolling lines JLine never sees. **Honest limit:** the lock is + reasoned, not test-covered. Two attempts to pin it are recorded in the history of + `JLineTerminalTest` and both passed with the lock removed — even one that sliced every + `OutputStream.write` in half — because `PrintWriter` already makes a single call atomic and the + interleaving happens *between* calls, inside JLine. A test that is green either way is worse than + none, so it was deleted rather than kept. + + **`/cls` wipes the screen, `/clear` wipes it and the history.** `AgentTerminal.clearScreen()` + defaults to doing nothing (a stream has no screen); `JLineTerminal` expands the terminal's + `clear_screen` capability and sends it **through `printAbove`**, like every other write, then + refills the blank rows and redraws the block. One trap, caught by the test rather than by reading: + `getStringCapability` returns **terminfo source** (`\E[H\E[2J`, with the escape spelled out), so + writing it as it comes prints that text on the screen — `Curses.tputs` expands it. A second one, and + the reason the bar went missing after a `/cls`: **`Status.redraw()` writes nothing after a wipe.** + It draws what has *changed*, and a wipe changes nothing about its content — it only removes it from + the screen, which the object has no way of knowing. So the block is kept in a field as it was last + rendered and put back with `status.reset()` (forget what is believed to be on screen) followed by + `status.update(block)`; `redraw()` alone is a no-op, verified by putting it back and watching the + test go red. Ctrl-L already + did this before the command existed, bound by JLine's own keymap; a test pins that too, so a keymap + option cannot quietly remove it. + + **The screen is scrolled to the bottom once, before the first prompt** (`scrollToBottom`). The + reader draws its prompt at the cursor, i.e. after the last line printed, while only the status + block is pinned to the window — so on a half-empty screen the input floats in the middle with the + block far below it, and they only meet once output has scrolled the cursor down by itself. That is + why it looked right after a few turns and like an ordinary prompt at the start. Emitting + `rows - 1` newlines once makes it the state from the first prompt on; from then on every printed + line scrolls and the cursor stays on the last row. The cost is a screenful of blank lines above the + session, which is what a program that wants its input at the bottom *without* taking over the + screen has to pay. + + **Do not add a `WINCH` handler, and the reason is measured.** A resize drawing a row of + `> > > > >` across the screen looks like the pinned region not being told about the new size, so a + `Signal.WINCH` handler that resized and re-rendered it was added — and the user reported it + **worse**, not better. `LineReaderImpl.handleSignal(WINCH)` already calls `Status.resize(Size)`, + and the reader installs its own handler for as long as it is reading, which is the whole session; + ours therefore either never ran or ran *in addition*, putting a second writer on the terminal from + the signal thread at the exact moment the reader was redrawing. A probe driving a real + `terminal.raise(WINCH)` against a pipe-backed terminal shows JLine doing it correctly on its own: + scroll region reset, the rule re-cut to the new width, **one** prompt. So the remaining report is + not reproducible in the harness and has no fix here yet — stated rather than papered over. + `fit()` measuring in **screen columns** (`AttributedString.columnLength`) rather than characters is + ours and does matter: an icon is one character and two columns, and a row wider than the window + wraps onto a second screen line, which the reserved region cannot survive. + + **Two things tried and removed, both because measurement said they did nothing.** (1) Skipping a + status write when the block is unchanged — removing the guard again left the emitted bytes + identical, because JLine already skips an unchanged block. (2) Re-cutting the rows when the window + width changed — `Status.resize()` re-cuts the rows it holds itself. Neither was kept with a comment + claiming a benefit it does not have. The stray `?1h` consequently still has **no established + cause**: the lock covers our writes, the reader's own are inside JLine. + + **Restoring the block after a wipe takes three steps, found by measurement not by reading**: + `status.reset()`, then `status.update(List.of())`, then render it again from `requested` (the text + the caller gave, not the rendered rows). With only the first two, `Status` draws the difference it + computes against a belief the wipe invalidated — observed as a single character emitted where a + whole block was missing. `JLineTerminalTest.theBlockIsBackOnScreenAfterAClear` is what says so. + + **The turn after an interrupted one is pinned end to end** (`InterruptedTurnTest`): with a scripted + backend behind the real `OpenAiCompatServer`, a turn is cut short by pending input and the next one + is then driven through — it reaches the server, carries the interrupted question in its history, + and answers. Reported as "it does not carry on by itself"; the mechanism works, so the cause of + that report is elsewhere and is **not** claimed to be fixed. Two things the writing of it settled: + a turn that finishes inside one activity tick is never even looked at for interruption (correct — + there is nothing to cut short), which is why the scripted backend has to be made slow or the test + proves nothing; and a second line typed during the replacement turn stops that one too, which is + the design and not a defect, but looks from the outside exactly like a turn that never started. + + **A blank line must not count as pending input.** `hasPendingInput()` ignores blank lines but leaves + them queued: counting them meant that holding Enter cancelled one turn per keystroke and produced + nothing, while dropping them would break the approval prompt, where an empty answer means yes. + `DISABLE_EVENT_EXPANSION` is set in the same builder because the reader's default treats `!` as a + shell history expansion, which silently rewrites a request like `git commit -m "fixed!"`. + **What cannot be done, asked and answered:** keep the block visible while the *user* scrolls the + terminal's scrollback. That needs the alternate screen buffer, i.e. a full-screen application, which + would give up the scrollback and the append-only property the whole console design rests on. + + **The prompt stays at the bottom during a turn, and typing stops the turn.** One thread inside + `JLineTerminal` (`startReading`) sits in `readLine` for the whole session and fills a queue; + **every** read in that class is served from it, because a terminal has one keyboard and two threads + reading it take turns at random. That is why `readKey` no longer reads a single key in raw mode: the + question is printed above the prompt and answered in the same input line (`y` + Enter). End of input + cannot be a queue value, so a sentinel is queued and **put back on every take** — otherwise the + second reader after Ctrl-D would see "nothing typed yet" instead of "no more input". + `AgentTerminal.hasPendingInput()` is the peek the turn loop peeks with; it deliberately does **not** + consume, so the line the interruption was triggered by is still there for the next `readLine` and + becomes the next message. The stop itself is `AgentRunner.start(...)` → + `runtime.executeWithHandle(...)`, whose handle closes the in-flight SSE stream (Atmosphere's own + "D-6 built-in hard-cancel"); that also replaced `turn()`'s hand-rolled worker thread, since + `executeWithHandle` dispatches on a virtual thread and returns at once. `awaitWithActivity` returns a + three-valued `TurnEnd` rather than a boolean, because *interrupted* must not be reported as the + *timed out* error the old `false` produced. **Order in that loop is load-bearing and a test pins the + behaviour**: the `activity.isPaused()` check comes first, so while an approval question is open a + typed line is its answer and not an interruption. **What this is not:** Claude Code injects a + mid-turn message into the running loop; `AgentExecutionContext` is a record whose request is built + once from `message()` + `history()`, with nothing to append to, so stop-and-resend is the achievable + equivalent — and it acts immediately instead of waiting out a tool loop. + **The default system prompt is general-purpose on purpose — do not narrow it back.** Every model-facing text is a resource, not a Java literal: `src/main/resources/net/ladenthin/llama/atmosphere/` holds `system-prompt.txt`, `system-prompt-shell.txt`, `system-prompt-no-shell.txt` and `run-command-tool.txt` diff --git a/README.md b/README.md index 551c09723..daa243ebd 100644 --- a/README.md +++ b/README.md @@ -1042,8 +1042,14 @@ mvn -q compile exec:java \ ``` > [!WARNING] -> With `--allow-shell` the model runs any command it decides to run, with your user's rights and -> without asking. Use a machine and a workspace you are willing to hand to the model. +> `--allow-shell` lets the model run any command with your user's rights. By default every write and +> every command is confirmed on the console (`[y]es / [n]o / [a]uto`); `--auto` turns that off. Use a +> workspace you are willing to hand to the model. + +In the REPL, `/help` lists the commands (`/status`, `/tools`, `/mode manual|auto`, `/compact`, +`/clear`, `/exit`); anything else goes to the model. A status line shows the approval mode and the +context used (`[manual · ctx ~3.1k/16k · 9 tools · local-model]`), and the answer is rendered with +headings, bullets and code spans. On Windows PowerShell quote the whole argument (`"-Dexec.args=--model models\… --allow-shell"`); for the GPU add e.g. `-Dllama.classifier=vulkan-windows-x86-64` and `--ngl 99`. The agent's diff --git a/llama-atmosphere-agent/README.md b/llama-atmosphere-agent/README.md index 293595abe..30c5c8ef1 100644 --- a/llama-atmosphere-agent/README.md +++ b/llama-atmosphere-agent/README.md @@ -15,9 +15,10 @@ offline: you start yourself, or the GGUF loaded **in this process**. - **Agent:** [Atmosphere](https://github.com/Atmosphere/atmosphere)'s built-in OpenAI-compatible runtime (`org.atmosphere:atmosphere-ai`): streaming, the model→tool→model loop, and its - workspace-confined file tools (`ls`, `read_file`, `write_file`, `edit_file`, `glob`, `grep`, - `delete`, `rename`). Driven **headless** — no Spring Boot, no servlet container, no `@Agent` - scanning — through `BuiltInAgentRuntime`. + workspace-confined file tools (`ls`, `write_file`, `glob`, `delete`, `rename`). Driven **headless** + — no Spring Boot, no servlet container, no `@Agent` scanning — through `BuiltInAgentRuntime`. + `read_file`, `edit_file` and `grep` are this project's own (see [Tools](#tools)), on the same + workspace-confined filesystem. - **Shell:** an opt-in `run_command` tool (`--allow-shell`) that runs any command line through the system shell (`cmd.exe` on Windows, `sh` elsewhere). @@ -27,10 +28,10 @@ run it. Its `pom.xml` pins `llama.version` to the release these instructions des pass `-Dllama.version=…` to run against another core, e.g. a `-SNAPSHOT` before a release. > [!WARNING] -> With `--allow-shell` the model runs **any** command it decides to run, with **your** user's -> rights, without asking — deleting files, pushing to git, stopping containers included. Start it on -> a machine and account you are willing to hand to the model, and point `--workspace` at a copy of -> a project, not your only one. Without the flag it can only use the file tools inside `--workspace`. +> `--allow-shell` lets the model run **any** command with **your** user's rights. By default it asks +> first (`[y]es / [n]o / [a]uto` per call) and only writes and commands are gated — but `--auto`, and +> the `[a]` answer, turn that off for the rest of the session. Point `--workspace` at a copy of a +> project, not your only one. ## Getting started from scratch @@ -102,9 +103,9 @@ Metal with the default jar already. The root README's classifier table lists eve -Dexec.args="--base-url http://127.0.0.1:8080/v1 --workspace /path/to/project --allow-shell" ``` -A `you>` prompt appears. Type a request; the answer streams as it is generated, and every tool call -and its result are printed as `⚙ read_file {path=…}` / `↳ …` lines. `/clear` drops the history, -`/exit` quits. +A `you>` prompt appears, with a status line pinned to the bottom of the window. The answer streams as it is generated, and every +tool call and its result are printed as `● read_file {path=…}` / `↳ …` lines. See +[Commands, approval and the status line](#commands-approval-and-the-status-line). A single turn without the REPL: @@ -150,6 +151,8 @@ irrelevant: inference stays in the running server, the agent's JVM loads no mode | `--log-verbosity ` / `--verbose` | llama.cpp log threshold for `--model` (1 errors, 2 warnings, 3 info, 4 trace, 5 debug) / log everything | `2` / off | | `--workspace ` | directory the file tools are confined to, and where `run_command` starts | cwd | | `--allow-shell` | register `run_command`: any command line, starting in the workspace | off | +| `--auto` | run tools without asking (otherwise every write and command is confirmed) | off | +| `--auto-compact ` / `--compact-at ` | summarize the history before it overflows the context, and how full it may get first | `true` / `70` | | `--system ` | replace the default system prompt | built-in | | `--prompt `, `-p` | one turn, then exit | interactive | | `--temperature ` / `--max-tokens ` | sampling / per-call budget | `0.2` / `2048` | @@ -172,6 +175,404 @@ code page it saw at startup, so umlauts and emoji in the answer would turn into project's `.mvn/jvm.config` pins `-Dstdout.encoding=UTF-8 -Dstderr.encoding=UTF-8` for the `mvn` JVM so both sides agree. +### Two consoles: the full one, and `--plain` + +There are two, and both stay. The default is the **full console** described below: a block pinned to +the bottom of the window, an input line that is there while the agent works, a spinner, single-key +navigation. It positions the cursor, so it needs a terminal that reports its size and understands the +sequences. + +`--plain` chooses the **line-oriented console** instead. It only ever appends lines: the status is +printed as an ordinary line before the prompt and scrolls away with everything else, there is no +pinned block and no spinner, and nothing on screen is ever rewritten. That makes a session readable +when it is piped, logged, recorded, or carried by anything that forwards lines rather than a screen: + +```bash +mvn -q compile exec:java -Dexec.args="--base-url http://127.0.0.1:8080/v1 --plain" | tee session.log +``` + +The same console is what the agent falls back to on its own when there is no usable terminal — a pipe, +a `dumb` terminal, an editor's run window — so `--plain` only *forces* what would otherwise be +detected. Note that a normal SSH session does **not** need it: a remote terminal reports its size and +handles cursor control like a local one. It is for the cases where that is not true. + +What the line-oriented console gives up, so the choice is an informed one: + +| | full (default) | `--plain` | +|---|---|---| +| input while the agent works | yes, and a typed line stops the turn | no, the prompt appears between turns | +| status | pinned at the bottom | printed once before each prompt | +| activity / spinner | yes | dropped rather than repeated into the log | +| approvals | `y` + Enter | `y` + Enter | +| history, Tab completion, Ctrl-L, Shift+Tab | yes | no | +| correct when the output is a file | — | yes | + +### Asking again, and your own system prompt + +`/retry` (`/again`) sends the last question once more and **drops the answer that came back** from the +conversation first. That is the whole point: leaving it in place would show the model what it said +last time, and a model that sees its own answer repeats it. Before anything has been asked it says so +rather than sending an empty turn. + +`--system-file ` replaces the system prompt with the content of a file — the same as `--system` +but without fighting the shell over quoting and newlines. It is read at startup, so a wrong path is a +usage error immediately instead of a surprise on the first turn. + +### Multi-line input + +A console sends on Enter, so pasting a stack trace or a function into the prompt would turn it into +several questions. A line that is exactly `"""` opens a block and the next one closes it; everything +between is **one** message, newlines and all: + +``` +you> """ +... public int add(int a, int b) { +... return a - b; +... } +... """ +``` + +Chosen over a key combination because it works in both consoles, survives a paste — the fences arrive +as part of the pasted text — and asks nothing of the terminal. End of input inside an unfinished block +ends the session without sending it. + +### The session transcript + +What was said is recorded separately from the conversation the model is sent, because those are two +different things: `/compact` **rewrites** the model's conversation (a summary replaces the turns it +summarises) and it never carried a timestamp at all. The transcript only grows. + +- `/save [name]` writes it into the workspace, one stamped line per entry: + `[2026-09-24 18:41:07] you: what does hello.txt say?`. Without a name the file is named after the + time. +- `--transcript ` appends every entry **as it is said**, so a session that is killed still + leaves what it had. A failure to write is swallowed on purpose: a record that exists to survive a + bad ending must not cause one. +- `/compact` keeps the record and notes that it happened. `/clear` empties it, because that command + means "forget this session" and leaving the text behind would make that untrue. +- `/calls` is the shorter, tool-only receipt and is unchanged. + +`/load ` (`/resume`) reads one back **as the conversation**: the questions and the answers +become messages again, so the model can be asked to carry on rather than to start over. Tool calls and +session notes stay in the record but are **not** replayed — a tool result outside its round is not +something a chat template has a place for, and inventing one would be worse than letting the model +call the tool again. A path without a directory is resolved in the workspace, so `/load session.txt` +finds what `/save session.txt` wrote; a file that cannot be read, or that is not a transcript, says so +and changes nothing. + +It is a list, not a map keyed by the timestamp: a tool result and the answer that follows it regularly +land in the same millisecond, and a map would keep one and drop the other silently. Insertion order +already is time order. + +### Commands, approval and the status line + +A line that starts with `/` and names a command is answered by the agent itself; anything else — an +unknown `/command` included — goes to the model: + +| Command | | +|---|---| +| `/help` (`/?`, `/commands`) | the overview below | +| `/status` | mode, context use, tools, model, workspace, history size | +| `/tools` | the tools offered, and which of them ask first | +| `/calls` (`/log`) | every tool call of this session with its result — the receipt | +| `/mode [manual\|auto]` (`/approve`) | show or set the approval mode (`⏸ manual` / `⏵⏵ auto`) | +| `/compact [focus]` | summarize the conversation and continue from the summary | +| `/loop [--every 5m] [--max 20] [--check ''] ` | keep working on one task until it is done | +| `/clear` (`/reset`, `/new`) | drop the history, and wipe the screen with it | +| `/cls` (`/clear-screen`) | wipe the screen, keep the conversation — Ctrl-L does the same | +| `/exit` (`/quit`) | leave | + +**The prompt.** On a real terminal the agent uses [JLine](https://github.com/jline/jline3): arrow keys +and the usual editing shortcuts work, ↑ recalls earlier lines, Tab completes the commands, Ctrl-C +drops the current line and Ctrl-D leaves. The status line and the rule above it stay at the bottom +while the answer scrolls past, and while a turn runs that line shows what is going on +(`⠙ working… (12s · 2 tool calls)`). With piped input, in one-shot mode and wherever JLine finds no +terminal, everything falls back to plain `println`/`readLine` — same features, no cursor tricks. + +**Approval.** In the default `manual` mode every tool that writes or runs a command — +`run_command`, `write_file`, `edit_file`, `delete`, `rename` — asks before it runs: + +``` +● run_command {command=rm -rf build} +? run_command {command=rm -rf build} + allow? [y]es / [n]o / [a]uto (no more questions): +``` + +`[y]` runs it once, `[n]` cancels it *and tells the model*, so it replans instead of assuming the +command ran, `[a]` switches to `auto` for the rest of the session (`/mode manual` switches back). On a +terminal a single key is enough — no Enter; Enter alone also means yes, and Ctrl-C means no. +Reading tools (`ls`, `read_file`, `glob`, `grep`) never ask. **In one-shot mode (`--prompt`) nobody +can answer, so a gated call is denied** — pass `--auto` to run unattended. The gate itself is +Atmosphere's (`ToolApprovalPolicy` + `ApprovalStrategy`); the agent only supplies the question and +the answer. + +**`/calls` is the receipt.** It lists every tool call of the session, one line each, with the +arguments and a short result. Use it when an answer sounds too good: a model that has drifted starts +*describing* work — "the tests passed, the jar was created" — while calling nothing at all. The +scrollback reads the same either way; this list only grows when something really ran. + +To make that drift less likely, each turn's calls ride along with the **next** message as a short +record. Without it the history holds only the user's messages and the model's own prose, and a small +model then continues the pattern it sees — prose. This is not theoretical: in a real session a 4B +model invented JUnit tests, a Maven build and a `.bat` script it had never written, three turns in a +row. + +Two placements were tried and discarded, both visible failures: in front of the assistant's answer +(the model copied it into its next reply, so the record appeared as the first line of an answer), and +as real `tool_calls` messages (impossible — Atmosphere's `assembleMessages` rebuilds every history +entry as `new ChatMessage(role, content)` and drops the rest). A system message mid-history would be +cleaner, but Mistral's template requires strict user/assistant alternation and Gemma has no system +role at all. + +**Auto-compaction.** Once the next request would fill more than `--compact-at` percent of the context +(70 by default), the history is summarized **before that request is sent** rather than after it — the +oversized request is the one thing worth avoiding, and afterwards it has already gone out. You see +`(context nearly full — compacting first)`, then your message is answered with the summary as its +context. `--auto-compact false` turns it off; `/compact` remains available at any time. With an +unknown context size — a foreign endpoint whose `/props` answers nothing — nothing is triggered at +all rather than guessed. The threshold sits below the ~85 % a hosted agent uses because our token +number is usually an estimate and the reply still has to fit next to the prompt. + +Calling `/compact` twice in a row answers `(the history is already a summary — nothing to compact)`: +after a compaction the history *is* the summary plus its acknowledgement, so summarizing it again +returns the same text for another model call. It also re-sends a byte-identical prompt, which is what +makes llama.cpp log `need to evaluate at least 1 token for each active slot` — a harmless note from +the server about a prompt it has already cached in full, not an error on our side. + +**`/compact`** asks the model to summarize the conversation (goal, facts, work done, problems, state, +next step; `/compact ` adds an emphasis), then replaces the history with that summary. Use it +when the context fills up. Note the history only ever held the user texts and the final answers — +tool rounds are not replayed across turns — so nothing else is lost. + +**`/loop`** keeps working on one task without you typing anything between steps: + +``` +/loop --check 'mvn -q test' make ShellToolTest pass on Windows +``` + +Every step sends the **same** message — the task verbatim plus "read `AGENT-LOOP.md`, do one concrete +step, write down what happened". The conversation history is **dropped between steps**: the file in +the workspace is the memory, so the context never grows and the loop can run for a long time. The +loop ends when a line of the answer is exactly + +``` +<> +``` + +A marker *mentioned* inside a sentence does not count, only a line of its own. This is a text marker +rather than a "done" tool on purpose: small local models produce a well-formed tool call far less +reliably than a line of text — mini-SWE-agent reaches its SWE-bench results with a plain sentinel and +no tool-call API at all, and Claude Code's own ralph-wiggum plugin matches an exact string too. + +With `--check ''` the marker is only believed when that command succeeds; otherwise its +output goes into the next step. That is the cheapest defence against a small model declaring victory +after one edit. Four limits stop a runaway loop, all enforced by the agent, none of them trusted to +the model: `--max` steps (20 by default), a two-hour wall-clock budget, a stall detector (three steps +in a row that write nothing and call no tool), and `--every ` for a paced run. A loop needs +the `auto` approval mode — it asks once and switches, or leaves you alone if you say no. + +**While a turn runs, the pinned line says what is happening**: `⠙ Fettling… (5s)` while the model +generates, and `⠙ Fettling… (run_command 47s of 61s · 2 tool calls)` while a tool is executing. The +word is drawn once per turn from +[`spinner-words.txt`](src/main/resources/net/ladenthin/llama/atmosphere/spinner-words.txt) — our own +two dozen, because Claude Code's list is extracted from a proprietary binary and the public copies of +it are either unlicensed or CC BY-NC-SA, neither of which fits an MIT project. Edit the file to +change them. The numbers stay next to the word on purpose: with a local model, "which tool, for how +long" is worth more than the joke. A build that +takes two minutes is otherwise indistinguishable from a hang. `run_command` additionally prints its +output **line by line while it runs** (dimmed, `│ `-prefixed) instead of dumping it at the end — which +also keeps the pipe drained; a process whose output nobody reads blocks once the buffer is full, and +on Windows that buffer is about 4 KB. + +**The bottom of the window is a pinned block**: the input line, then a rule, then what the agent is +doing and the session state. + +``` +› add a test for the parser +… the answer … +> +──────────────────────────────────────────────────────────────────── +⠙ Fettling… (run_command 47s of 61s · 2 tool calls) +[/path/to/project · ⏸ manual · ctx ~3.1k/16k · 9 tools · local-model] +``` + +There is no `you>`: the block already says where the input is. On Enter the input line is erased and +echoed above as `› your text`, so the transcript keeps what was asked. (`!` history expansion is +switched off in the same place, or a request like `git commit -m "fixed!"` would be rewritten silently.) + +**Why there is no second rule above the input**, although that is what this looked like at first: JLine's +pinned region sits below the prompt and never above it, so a rule above the input can only be part of the +prompt or part of the scrollback — and both were tried and both were wrong. In the prompt it survives +every Enter, because the reader erases exactly one line (hold Enter, get a column of rules). As output it +leaves one rule per turn behind, travelling up the scrollback. One rule, below the input, is the shape +that has neither problem. + +**Wiping the screen.** `/cls` clears the window and leaves the input and the block at the bottom — +Ctrl-L does the same, bound by the line reader itself rather than by this project (a test pins that, so +a keymap change cannot quietly take it away). `/clear` wipes the screen *and* drops the history: what is +still on screen after a `/clear` is a conversation the model no longer has, which reads as if it were +still in play. Neither touches the terminal emulator's own scrollback — what was written stays where +the scrollbar can reach it. The block at the bottom is redrawn from a kept copy afterwards: JLine draws +the pinned region only when its *content* changes, and a wipe does not change the content, it only takes +it off the screen — so asking it to redraw does nothing and the bottom of the window stays empty. + +**Why the screen is scrolled once at startup.** The line reader draws its prompt where the cursor is, +which is directly after the last thing printed; only the block below it is pinned to the window. On a +half-empty screen that leaves the input floating in the middle with the block far below, and the two +only meet once enough output has scrolled the cursor down by itself — which is why it looks right after +a few turns and wrong at the start. Pushing the cursor to the last row before the first prompt makes +that the state from the beginning. The cost is a screen of blank lines above the session: the +alternative is taking over the whole screen (alternate buffer), which costs the scrollback. + +**The status line is icons and values**: `📁` workspace, the mode glyph, `📊` context, +`🔧` tools, and `🤖` for a model this process loaded or `🌐` for one reached over the +network. Each is an icon, a space, its value. Rows are cut by **screen columns** rather than characters, +because an icon is one character and two columns — counting characters lets a row come out wider than +the window, wrap, and push the pinned block out of place. + +**Resizing the window** is left to JLine, which resizes the pinned region and re-cuts its rows itself. +A handler of our own was tried for a reported row of `> > > > >` after dragging the window smaller and +made it worse: the line reader installs its own handler for as long as it is reading, so ours only +added a second writer on the terminal while the reader was redrawing. A test drives the real path — a +size change plus the resize signal, with the reader reading as it does all session — across shrinking, +growing and a changed row count, and every one draws exactly one prompt. The leftover `> ` row that is +still reported is therefore not produced there; the remaining suspect is the console reflowing its own +screen buffer on a resize, which moves lines the program never wrote again and which nothing on this +side can reproduce. `/cls` or Ctrl-L cleans it up. + +**Three threads write to this console** and all of them had to be brought into line, because a write that +goes around the line reader scrolls the screen without JLine noticing and the pinned block ends up +somewhere else than it believes — first as a stray `1H` drawn into the rule, then as no block at all. +The turn thread and the console thread share a lock; llama.cpp's own log, which with `--model` goes to +stderr and therefore straight past everything, is routed through the console with `LlamaModel.setLogger`. + + +**On scrolling.** The block stays put while the agent writes: JLine keeps those lines out of the +terminal's scroll region. It cannot stay while *you* scroll the terminal's own scrollback with the +mouse — then the whole viewport moves and no program on this side of the terminal has a say. Staying +visible through that needs the alternate screen buffer, i.e. a full-screen application, which would +trade away the scrollback and the "written once, never redrawn" property this console is built on. The +established terminal agents behave the same way. + +On a plain stream (piped input, a one-shot run) nothing can be pinned, so the state line is printed +above the prompt instead and the activity row is dropped rather than repeated into the log. + +**The status line** above the prompt reads +`[/path/to/project · ⏸ manual · ctx ~3.1k/16k · 9 tools · local-model]`: the workspace, the approval +mode, the context used out of the window, the number of tools and the model id. + +The mode carries a glyph as well as its name — **`⏸ manual`** stops at every gated call, **`⏵⏵ auto`** +runs through — so the one thing that decides whether the next command asks first is findable without +reading the line. **Shift+Tab at the prompt switches it**, without typing `/mode`; the status line +updates on the key. The shortcut needs a real terminal and works between turns (while the prompt is +waiting), which is when the mode matters — during a turn nobody is reading keys. Where the terminal +does not send backtab, the startup line simply does not offer it and `/mode` still works. + +The context figure **moves while the turn runs**, not only at the next prompt: every tool round appends +the call and its output to the conversation the next model call of the same turn is sent, so a turn that +reads three files and runs a build can add thousands of tokens before you get the prompt back. A `~` +means the number is an estimate from the text length: llama.cpp reports token counts only to clients +that ask for them (`stream_options.include_usage`), which Atmosphere's client does not. The window size +comes from `--ctx-size` with `--model`, and from the server's `/props` with `--base-url`; when neither +answers, the line shows the count alone. + +**The prompt is at the bottom the whole time, and you can type while the agent works.** One thread owns +the keyboard and sits in the line reader for the entire session; everything else is written *above* the +prompt. A line typed **during** a turn **stops that turn** — Atmosphere's cancellable entry point closes +the HTTP stream the model is answering on — and is then sent as the next message, with whatever the model +had already produced kept in the history. That is not quite what Claude Code does (it feeds the message +into the running loop); Atmosphere builds its request once from the message plus the history and has no +place to append to, so stop-and-resend is the honest equivalent, and it is immediate rather than waiting +out a tool loop that may run for minutes. + +The cost is one keystroke: the approval question is answered in that same input line, so it is +`y` + Enter rather than a bare `y`. A single-key read needs a second reader on the same keyboard, and two +readers on one terminal take turns at random. While a question is open the typing-interrupts rule is +suspended, so an answer is an answer and not an interruption. + +**One printed line is one screen line.** A tool call and its result are shown as +`● write_file {file_path=notes.md, content=# Notes ## Build … (4812 chars)}` — every argument is +folded onto one line and cut *on its own* before the whole thing is cut, so a call carrying a whole +file still shows the file *name*. The reason is not tidiness: the block at the bottom is reserved in +*lines*, so a single "line" carrying twenty newlines moves the screen twenty rows further than the +terminal accounted for and the block ends up drawn across the output — which is what a `write_file` +call did. The model still receives every argument and every result in full; only the console is cut. + +**Colours and Markdown.** The answer is rendered line by line as it streams: headings, bullets, +fenced code blocks and inline `**bold**` / `` `code` ``. Nothing is ever redrawn, so piping the output +into a file stays correct. Colour is on only on a real terminal and obeys `NO_COLOR`, `TERM=dumb`, +`CLICOLOR=0` and `CLICOLOR_FORCE=1`. On the classic Windows `conhost.exe` escape sequences may show up +literally unless `HKCU\Console\VirtualTerminalLevel` is 1 — Windows Terminal needs nothing. + +### Try it + +Start the agent with shell access (add `-Dllama.classifier=…` and `--ngl 99` for a GPU; leave both +out to stay on the CPU): + +```bash +mvn -q compile exec:java \ + -Dexec.args="--model models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ctx-size 16384 --workspace /path/to/project --allow-shell" +``` + +Then, in this order: + +| Type this | What should happen | +|---|---| +| `/help` | the command overview — the agent answers, the model never sees the line | +| `/status` | mode, context use, tools, model, workspace, history size | +| `/tools` | every tool, and which of them ask before running | +| `docker is running locally, list the images` | `? run_command {command=docker images}` and the prompt `[y]es / [n]o / [a]uto` | +| answer `n` | the command does **not** run; the model is told it was cancelled and offers an alternative | +| ask again, answer `y` | the command runs and its output goes back to the model | +| `/mode auto` | the status line flips to `⏵⏵ auto`; nothing asks any more | +| press Shift+Tab at the prompt | the same switch without a command; the status line updates immediately | +| type a sentence while it is still working, press Enter | the turn stops at once and your message is the next one | +| `explain Markdown with a heading, a list, bold text and a code block` | the answer arrives rendered: heading bold, `•` bullets, code in colour | +| press ↑ | the previous line comes back; Tab after `/` completes the commands | +| `/compact` | the conversation is summarized and replaces the history; `ctx` drops | +| `/loop --max 3 add a line with the current date to notes.txt, then stop` | three steps at most, with `AGENT-LOOP.md` appearing in the workspace | +| `/exit` | leave | + +`--auto` starts in auto mode, `--verbose` brings llama.cpp's own log back, and `NO_COLOR=1` turns +the styling off. + +### Tools + +Eight file tools plus the optional shell. Five are Atmosphere's; three are replaced here because what +they return decides how well a model can work: + +| Tool | | | +|---|---|---| +| `ls`, `write_file`, `glob`, `delete`, `rename` | Atmosphere | unchanged | +| `read_file(file_path, offset, limit)` | **ours** | a numbered window, ` 12: text`, 400 lines at a time, and it says what it left out. Reading whole files costs context and measurably lowers task success (SWE-agent: 12.7 % with whole files against 18.0 % with a 100-line window) | +| `edit_file(file_path, old_string, new_string, replace_all, edits[])` | **ours** | see below | +| `grep(pattern, dir, glob, files_only)` | **ours** | skips `.git`, `target`, `build`, `node_modules` and friends, groups matches by file with line numbers, caps at 100 matches and **says so** when it truncates | +| `run_command` | ours | opt-in via `--allow-shell` | + +**Why `edit_file` is not Atmosphere's.** Four things it does that the framework's does not, each for a +measured reason: + +1. **Line endings.** The framework compares the raw file content, so a model's LF text never matches a + CRLF file — on Windows every edit fails silently. Here the file is normalized before matching and + written back in its own ending (byte-order mark included). +2. **A miss explains itself.** Instead of "not found", it shows the closest lines in the file with + their numbers. A failed edit is not a free retry: measured on SWE-agent trajectories, an edit + attempt eventually succeeds in 90.5 % of cases, but only 57.2 % once one edit has failed. +3. **An ambiguous match names the lines** (`occurs 2 times, on lines 1, 3`) instead of asking for + "more context", and `replace_all` is offered. +4. **Several edits in one call** via `edits: [{old_string, new_string}]`, applied **all or nothing** — + every shipping agent applies them sequentially and leaves a half-edited file behind. + +`edit_file` also **refuses to edit a file that was not read** in this session, so `old_string` comes +from the file rather than from the model's memory. + +**Formats that were considered and rejected.** A unified-diff or patch tool: Meta's ablation measures +search-replace at 42–53 % against 26–30 % for unified diff and 20–26 % for line diffs on the same +model, and a 7B model drops from 54 % to 33 % to 14 % across those three. Fuzzy matching (a similarity +threshold instead of an exact match): it turns a loud "not found" into a silent edit in the wrong +place. Whole-file rewriting stays available as `write_file` — for a small model that is often the most +reliable route, and the system prompt says so. + ### The system prompt Without `--system` the agent uses a built-in **general-purpose** prompt: it names the file tools and @@ -261,7 +662,14 @@ starter are the *deployment* layer on top of the same runtime — not needed for ## Limitations / next steps - Tool rounds are not kept in the cross-turn history (only `user`/`assistant` text is replayed). -- No approval prompts for destructive tools yet (`ToolDefinition.requiresApproval` exists in Atmosphere). +- Approval is per call, not per command prefix: there is no "always allow `git status`" rule yet. + Atmosphere's `ApprovalResolution` also supports approve-with-edited-arguments, which the console + does not offer. +- No auto-compaction when the context fills up; `/compact` is manual. +- A small model still drifts into describing instead of doing, especially after several turns; + `/calls` makes it visible, the note in the history makes it rarer, a bigger model makes it go away. +- `/loop` cannot be interrupted in the middle of a step — Ctrl-C ends the process; the loop file + survives, so restarting the same `/loop` continues where it left off. - An engine error after the stream started ends the turn silently (see the table). - **One in-process agent per machine at a time.** The core extracts its native library to a fixed name (`jllama.dll` / `libjllama.so` in the temp directory); on Windows a second JVM cannot replace diff --git a/llama-atmosphere-agent/pom.xml b/llama-atmosphere-agent/pom.xml index 34c19ba1f..fa2fc4d04 100644 --- a/llama-atmosphere-agent/pom.xml +++ b/llama-atmosphere-agent/pom.xml @@ -48,6 +48,7 @@ SPDX-License-Identifier: MIT e.g. -Dllama.classifier=cuda13-linux-x86-64 or vulkan-windows-x86-64 (runtime on PATH). --> 4.0.70 + 4.4.5 2.0.19 1.0.1 6.1.3 @@ -91,6 +92,15 @@ SPDX-License-Identifier: MIT ${atmosphere.version} + + + org.jline + jline + ${jline.version} + + org.jspecify jspecify diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java index cdc63af39..018401895 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentOptions.java @@ -37,6 +37,18 @@ public final class AgentOptions { /** Context size for the in-process model ({@code --model}). */ public static final int DEFAULT_CTX_SIZE = 8192; + /** Whether the history is summarized on its own before it overflows the context. */ + public static final boolean DEFAULT_AUTO_COMPACT = true; + + /** + * How full the context may get before that happens, in percent. + * + *

Lower than the ~85 % a hosted agent uses, and deliberately so: this number is usually an + * estimate from the text length (llama.cpp reports its own count only to clients that ask for it), + * and the reply still has to fit next to the prompt. + */ + public static final int DEFAULT_COMPACT_AT = 70; + /** * Log verbosity threshold of the in-process model ({@code --model}): llama.cpp's {@code -lv} * scale, {@code 0} output only, {@code 1} errors, {@code 2} warnings, {@code 3} info, {@code 4} @@ -56,6 +68,11 @@ public final class AgentOptions { private final String modelId; private final Path workspace; private final boolean allowShell; + private final boolean plain; + private final java.nio.file.@org.jspecify.annotations.Nullable Path transcript; + private final boolean auto; + private final boolean autoCompact; + private final int compactAt; private final double temperature; private final int maxTokens; private final int maxToolRounds; @@ -74,6 +91,11 @@ private AgentOptions(Builder b) { this.modelId = b.modelId; this.workspace = b.workspace; this.allowShell = b.allowShell; + this.plain = b.plain; + this.transcript = b.transcript; + this.auto = b.auto; + this.autoCompact = b.autoCompact; + this.compactAt = b.compactAt; this.temperature = b.temperature; this.maxTokens = b.maxTokens; this.maxToolRounds = b.maxToolRounds; @@ -97,6 +119,11 @@ public static AgentOptions parse(String[] args) { switch (a) { case "-h", "--help" -> b.help = true; case "--allow-shell" -> b.allowShell = true; + case "--plain" -> b.plain = true; + case "--transcript" -> b.transcript = java.nio.file.Path.of(value(args, ++i, a)); + case "--auto" -> b.auto = true; + case "--auto-compact" -> b.autoCompact = booleanValue(args, ++i, a); + case "--compact-at" -> b.compactAt = percentValue(args, ++i, a); case "--base-url" -> b.baseUrl = stripTrailingSlash(value(args, ++i, a)); case "--model" -> b.modelPath = value(args, ++i, a); case "--ngl", "--gpu-layers" -> b.gpuLayers = intValue(args, ++i, a); @@ -112,6 +139,7 @@ public static AgentOptions parse(String[] args) { case "--max-tokens" -> b.maxTokens = intValue(args, ++i, a); case "--max-tool-rounds" -> b.maxToolRounds = intValue(args, ++i, a); case "--system" -> b.systemPrompt = value(args, ++i, a); + case "--system-file" -> b.systemPrompt = readSystemPrompt(value(args, ++i, a)); case "--prompt", "-p" -> b.prompt = value(args, ++i, a); default -> throw new IllegalArgumentException("Unknown argument: " + a); } @@ -131,6 +159,25 @@ private static String value(String[] args, int index, String flag) { return args[index]; } + private static boolean booleanValue(String[] args, int index, String flag) { + String raw = value(args, index, flag).trim(); + if ("true".equalsIgnoreCase(raw) || "yes".equalsIgnoreCase(raw) || "on".equalsIgnoreCase(raw)) { + return true; + } + if ("false".equalsIgnoreCase(raw) || "no".equalsIgnoreCase(raw) || "off".equalsIgnoreCase(raw)) { + return false; + } + throw new IllegalArgumentException("Expected true or false for " + flag + ", got: " + raw); + } + + private static int percentValue(String[] args, int index, String flag) { + int percent = intValue(args, index, flag); + if (percent < 10 || percent > 95) { + throw new IllegalArgumentException(flag + " must be between 10 and 95, got: " + percent); + } + return percent; + } + private static int intValue(String[] args, int index, String flag) { String raw = value(args, index, flag); try { @@ -168,6 +215,13 @@ public static String usage() { "Agent:", " --workspace

directory the file tools are confined to (default: cwd)", " --allow-shell add the run_command tool (runs any command line, starting in the workspace)", + " --system-file replace the system prompt with the content of a file", + " --plain line-oriented console: no pinned block, no cursor control", + " --transcript append what is said, with timestamps, as it happens", + " --auto run tools without asking (default: ask before writes and commands)", + " --auto-compact summarize the history before it overflows the context (default " + + DEFAULT_AUTO_COMPACT + ")", + " --compact-at how full the context may get first (default " + DEFAULT_COMPACT_AT + ")", " --system replace the default system prompt", " --prompt , -p run one turn and exit (default: interactive; /exit to quit)", " --temperature sampling temperature (default " + DEFAULT_TEMPERATURE + ")", @@ -259,6 +313,34 @@ public Path getWorkspace() { return workspace; } + /** + * Whether the history is summarized before it overflows the context. + * + * @return {@code true} when auto-compaction is on + */ + public boolean isAutoCompact() { + return autoCompact; + } + + /** + * How full the context may get before the history is summarized. + * + * @return the threshold in percent + */ + public int getCompactAt() { + return compactAt; + } + + /** + * Whether tool calls run without asking. + * + * @return {@code true} when {@code --auto} was given, i.e. the session starts in + * {@link ApprovalMode#AUTO} + */ + public boolean isAuto() { + return auto; + } + /** * Whether the {@code run_command} tool is registered. * @@ -268,6 +350,35 @@ public boolean isAllowShell() { return allowShell; } + /** + * Whether to use the line-oriented console even when a full terminal is available. + * + *

The rich console positions the cursor: it pins a block to the bottom of the window and keeps + * the input line there while output scrolls above it. That needs a terminal that reports its size + * and understands the sequences, which is the normal case over SSH as well — but not in a plain + * pipe, a CI log, a `dumb` terminal, an editor's run window or a serial console, and not when the + * session is being recorded as text. This flag chooses the console that only ever appends lines, + * which is also what the agent falls back to on its own when there is no usable terminal. + * + * @return {@code true} when {@code --plain} was passed + */ + public boolean isPlain() { + return plain; + } + + /** + * Where to append the session transcript as it happens, if anywhere. + * + *

{@code /save} writes the whole thing on request; this writes each line as it is said, so a + * session that is killed still leaves what it had. A failure to write is swallowed: a record that + * exists to survive a bad ending must not cause one. + * + * @return the file, or {@code null} when the transcript is kept in memory only + */ + public java.nio.file.@org.jspecify.annotations.Nullable Path getTranscript() { + return transcript; + } + /** * Sampling temperature. * @@ -304,6 +415,25 @@ public int getMaxToolRounds() { return systemPrompt; } + /** + * Read a system prompt from a file. + * + *

Read here rather than when it is used, so a path that does not exist is a usage error at + * startup instead of a surprise on the first turn. A prompt long enough to be worth a file is also + * long enough that a typo in the path is easy to miss. + * + * @param path the file + * @return its content + * @throws IllegalArgumentException when it cannot be read + */ + private static String readSystemPrompt(String path) { + try { + return java.nio.file.Files.readString(java.nio.file.Path.of(path), java.nio.charset.StandardCharsets.UTF_8); + } catch (java.io.IOException | RuntimeException e) { + throw new IllegalArgumentException("--system-file cannot be read: " + path + " (" + e.getMessage() + ")"); + } + } + /** * One-shot prompt. * @@ -327,7 +457,11 @@ public String toString() { return "AgentOptions{baseUrl=" + baseUrl + ", modelPath=" + modelPath + ", gpuLayers=" + gpuLayers + ", ctxSize=" + ctxSize + ", logVerbosity=" + (verbose ? "verbose" : logVerbosity) + ", modelId=" + modelId + ", workspace=" + workspace - + ", allowShell=" + allowShell + ", temperature=" + temperature + ", maxTokens=" + maxTokens + + ", allowShell=" + allowShell + ", plain=" + plain + ", transcript=" + transcript + ", auto=" + auto + + ", autoCompact=" + autoCompact + + ", temperature=" + + temperature + ", maxTokens=" + + maxTokens + ", maxToolRounds=" + maxToolRounds + ", prompt=" + (prompt == null ? "" : "") + "}"; } @@ -347,6 +481,11 @@ private static final class Builder { String modelId = DEFAULT_MODEL_ID; Path workspace = Paths.get("").toAbsolutePath().normalize(); boolean allowShell; + boolean plain; + java.nio.file.@org.jspecify.annotations.Nullable Path transcript; + boolean auto; + boolean autoCompact = DEFAULT_AUTO_COMPACT; + int compactAt = DEFAULT_COMPACT_AT; double temperature = DEFAULT_TEMPERATURE; int maxTokens = DEFAULT_MAX_TOKENS; int maxToolRounds = DEFAULT_MAX_TOOL_ROUNDS; diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentRunner.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentRunner.java index 3ca6c9bd0..574b420f0 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentRunner.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentRunner.java @@ -8,13 +8,17 @@ import java.util.Map; import org.atmosphere.ai.AgentExecutionContext; import org.atmosphere.ai.AiConfig; +import org.atmosphere.ai.ExecutionHandle; import org.atmosphere.ai.RetryPolicy; import org.atmosphere.ai.StreamingSession; +import org.atmosphere.ai.approval.ApprovalStrategy; +import org.atmosphere.ai.approval.ToolApprovalPolicy; import org.atmosphere.ai.llm.BuiltInAgentRuntime; import org.atmosphere.ai.llm.ChatMessage; import org.atmosphere.ai.llm.ToolLoopPolicies; import org.atmosphere.ai.llm.ToolLoopPolicy; import org.atmosphere.ai.tool.ToolDefinition; +import org.jspecify.annotations.Nullable; /** * The minimal wiring between Atmosphere's built-in OpenAI-compatible agent runtime and an @@ -36,6 +40,8 @@ public final class AgentRunner { private final String systemPrompt; private final int maxToolRounds; private RetryPolicy retryPolicy = RetryPolicy.DEFAULT; + private @Nullable ApprovalStrategy approvalStrategy; + private @Nullable ToolApprovalPolicy approvalPolicy; /** * Configure the runtime for one endpoint. @@ -83,6 +89,24 @@ public AgentRunner retryPolicy(RetryPolicy retryPolicy) { return this; } + /** + * Gate the tools {@code policy} selects behind {@code strategy}: Atmosphere's tool loop then blocks + * on the strategy before such a tool runs, and turns a denial into a {@code cancelled} tool result + * for the model on its own. + * + *

Without this, no tool is gated — Atmosphere's default policy honours a tool's own + * {@code requiresApproval()}, and none of this agent's tools set it. + * + * @param strategy what asks the user, e.g. {@link ConsoleApprovalStrategy} + * @param policy which tools it is asked about, e.g. {@link ConsoleApprovalStrategy#policy()} + * @return this runner + */ + public AgentRunner approval(ApprovalStrategy strategy, ToolApprovalPolicy policy) { + this.approvalStrategy = strategy; + this.approvalPolicy = policy; + return this; + } + /** * The model ids the endpoint advertises on {@code GET /v1/models}, falling back to the configured * id when enumeration fails. @@ -110,6 +134,47 @@ public List toolNames() { * @param session receives streamed text, tool events and the terminal complete/error */ public void run(String message, List history, StreamingSession session) { + runtime.execute(context(message, history, tools, systemPrompt, session), session); + } + + /** + * Start one user turn and return at once, with a handle that can stop it. + * + *

This is the same turn {@link #run} performs, on Atmosphere's cancellation-aware entry point: + * the turn runs on a virtual thread of the framework's, and {@link ExecutionHandle#cancel()} + * closes the HTTP stream the model is answering on, which unblocks the read loop. That is what + * lets a request typed while the agent is working take effect immediately instead of at the end + * of a tool loop that may run for minutes. + * + * @param message the user message + * @param history prior turns, replayed before the message + * @param session receives streamed text, tool events and the terminal complete/error + * @return the handle; the session's own completion stays the signal that the turn is over + */ + public ExecutionHandle start(String message, List history, StreamingSession session) { + return runtime.executeWithHandle(context(message, history, tools, systemPrompt, session), session); + } + + /** + * Run one turn with no tools at all and a system prompt of its own — what {@code /compact} needs: + * a summary must not read files or run commands, it must only condense what is already there. + * + * @param message the user message + * @param history prior turns, replayed before the message + * @param session receives the streamed summary + * @param systemPrompt the system prompt for this one turn + */ + public void runWithoutTools( + String message, List history, StreamingSession session, String systemPrompt) { + runtime.execute(context(message, history, List.of(), systemPrompt, session), session); + } + + private AgentExecutionContext context( + String message, + List history, + List tools, + String systemPrompt, + StreamingSession session) { AgentExecutionContext context = new AgentExecutionContext( message, systemPrompt, @@ -127,7 +192,9 @@ public void run(String message, List history, StreamingSession sess null, null); context = context.withRetryPolicy(retryPolicy); - context = ToolLoopPolicies.attach(context, ToolLoopPolicy.maxIterations(maxToolRounds)); - runtime.execute(context, session); + if (approvalStrategy != null && approvalPolicy != null) { + context = context.withApprovalStrategy(approvalStrategy).withApprovalPolicy(approvalPolicy); + } + return ToolLoopPolicies.attach(context, ToolLoopPolicy.maxIterations(maxToolRounds)); } } diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java new file mode 100644 index 000000000..6358a420e --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/AgentTerminal.java @@ -0,0 +1,129 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import org.jspecify.annotations.Nullable; + +/** + * Everything the agent needs from a console, so that a real terminal and a plain stream are + * interchangeable. + * + *

Two implementations: {@link JLineTerminal} for an interactive session (line editing, history, + * tab completion, a status line pinned to the bottom, single-key answers) and {@link PlainTerminal} + * for everything else — one-shot runs, piped input, and the tests. The agent picks one at startup and + * never branches again. + * + *

Output is line-oriented on purpose: a line is written once, complete, and is never touched + * again. That is what lets the pinned status line coexist with streamed output without redrawing + * anything the reader has already seen. + */ +public interface AgentTerminal extends AutoCloseable { + + /** + * Write one completed line, above the prompt and the status line. + * + * @param text the line, without a terminator + */ + void line(String text); + + /** + * Read one line from the user. + * + * @param prompt the prompt to show, e.g. {@code "you> "} + * @return the line, or {@code null} at end of input + */ + @Nullable + String readLine(String prompt); + + /** + * Ask a question and read the answer. + * + *

The answer is a line, terminated with Enter, on both consoles. A real terminal could read a + * single key instead, but only by opening a second reader on the keyboard, and the one reader it + * has is busy offering the prompt that stays visible while the agent works — which is worth more + * than saving an Enter on a question that is asked a few times a session. + * + * @param prompt the question to show + * @return the answer in lower case, or {@code null} at end of input + */ + @Nullable + String readKey(String prompt); + + /** + * Set the block kept at the bottom of the window. + * + *

Two lines in practice: what the agent is doing right now, and the session's state. Keeping + * both there at all times is what stops the block from changing height, which would make the + * output above it jump on every update. + * + * @param lines the lines, top to bottom; an empty list removes the block + */ + void status(java.util.List lines); + + /** + * Whether {@link #status} really pins the line to the bottom of the window. + * + *

{@code false} on a plain stream, where the caller has to print the status itself. + * + * @return {@code true} on a real terminal + */ + boolean pinsStatus(); + + /** + * Wipe the window, leaving the input line and the block at the bottom. + * + *

The scrollback of the terminal emulator is not touched — what was written stays where the + * scrollbar can reach it. This only clears what is on screen, the way {@code clear} or Ctrl-L does. + * + *

A plain stream has no screen to clear, so the default does nothing. + */ + default void clearScreen() { + // nothing to wipe on a stream + } + + /** + * Whether the user has already typed a line that nobody has read yet. + * + *

This is what makes the prompt useful during a turn: a console that keeps reading while the + * agent works can say so, and the turn is then cut short and the line answered instead of being + * made to wait for an answer nobody wants any more. The line stays queued — the caller reads it + * with {@link #readLine} as usual. + * + *

A console that reads only when asked has nothing pending by definition, which is the default. + * + * @return {@code true} when a line is waiting + */ + default boolean hasPendingInput() { + return false; + } + + /** + * Ask the terminal to run {@code action} when the user presses shift+tab at the prompt. + * + *

Only a real terminal can offer this: it needs to own the keyboard and see a key that is not + * a line of text. It also only fires **while a line is being read** — during a turn nobody is + * reading keys, so the mode is switched between turns, which is when it matters. + * + *

The default is to decline, which every non-interactive console does; the caller uses that + * answer to decide whether to advertise the shortcut, and {@code /mode} remains either way. + * + * @param action what to run on the key; it must not block + * @return {@code true} when the key was bound + */ + default boolean onCycleMode(Runnable action) { + return false; + } + + /** + * The styles to use for this console. + * + * @return a colouring instance on a terminal, a plain one otherwise + */ + Ansi ansi(); + + /** Restore the terminal. */ + @Override + void close(); +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Ansi.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Ansi.java new file mode 100644 index 000000000..aac4bed09 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Ansi.java @@ -0,0 +1,173 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.util.Locale; +import java.util.function.Function; +import org.jspecify.annotations.Nullable; + +/** + * The handful of ANSI styles this agent uses, and the one decision of whether to emit them at all. + * + *

The decision is made once at startup, not per line: when colour is off every helper returns its + * argument unchanged, so no other class branches on it. The order follows the conventions the + * terminal ecosystem actually honours: + * + *

    + *
  1. {@code CLICOLOR_FORCE=1} forces colour on, even when the output is piped (for {@code less -R} + * and CI logs); + *
  2. {@code NO_COLOR} set to any non-empty value turns it off — that is the whole + * NO_COLOR specification; + *
  3. {@code TERM=dumb} or {@code CLICOLOR=0} turns it off; + *
  4. otherwise colour is on only when the output really is a terminal. + *
+ * + *

The terminal test is {@code Console.isTerminal()} where it exists (JDK 22+) and + * {@code System.console() != null} below that. The distinction matters: on JDK 22 to 24 + * {@code System.console()} also returns a console for redirected output, so the older test alone + * would colour a file. Reflection keeps the code compiling and running on JDK 21. + * + *

Windows: Windows Terminal and the VS Code terminal process escape sequences without any setup. + * The classic {@code conhost.exe} does not unless {@code HKCU\Console\VirtualTerminalLevel} is 1 — + * enabling it from the process needs native code, which this agent deliberately does not use, so + * there a user may see the raw sequences and can set {@code NO_COLOR=1}. + */ +public final class Ansi { + + private static final String RESET = "\u001b[0m"; + + /** Never emits escape sequences. */ + public static final Ansi PLAIN = new Ansi(false); + + private final boolean enabled; + + private Ansi(boolean enabled) { + this.enabled = enabled; + } + + /** + * Decide from the environment whether to use colour. + * + * @return a colouring or a plain instance + */ + public static Ansi detect() { + return detect(System::getenv, Ansi::consoleIsTerminal); + } + + /** + * The decision itself, with the environment injected so it can be tested. + * + * @param env reads an environment variable + * @param isTerminal whether standard output is a terminal + * @return a colouring or a plain instance + */ + static Ansi detect(Function env, java.util.function.BooleanSupplier isTerminal) { + if ("1".equals(trimmed(env.apply("CLICOLOR_FORCE")))) { + return new Ansi(true); + } + String noColor = env.apply("NO_COLOR"); + if (noColor != null && !noColor.isEmpty()) { + return PLAIN; + } + if ("dumb".equals(lower(env.apply("TERM"))) || "0".equals(trimmed(env.apply("CLICOLOR")))) { + return PLAIN; + } + return isTerminal.getAsBoolean() ? new Ansi(true) : PLAIN; + } + + private static boolean consoleIsTerminal() { + java.io.Console console = System.console(); + if (console == null) { + return false; + } + try { + // JDK 22+: the only reliable "is a terminal" test; JDK 21 has no such method. + return (boolean) java.io.Console.class.getMethod("isTerminal").invoke(console); + } catch (ReflectiveOperationException | RuntimeException e) { + return true; + } + } + + private static @Nullable String trimmed(@Nullable String value) { + return value == null ? null : value.trim(); + } + + private static @Nullable String lower(@Nullable String value) { + return value == null ? null : value.trim().toLowerCase(Locale.ROOT); + } + + /** + * Whether escape sequences are emitted. + * + * @return {@code true} when styling is on + */ + public boolean isEnabled() { + return enabled; + } + + private String style(String code, String text) { + return enabled ? "\u001b[" + code + "m" + text + RESET : text; + } + + /** + * Bold text. + * + * @param text the text + * @return the styled text + */ + public String bold(String text) { + return style("1", text); + } + + /** + * Dimmed text, for secondary output such as tool results. + * + * @param text the text + * @return the styled text + */ + public String dim(String text) { + return style("2", text); + } + + /** + * Cyan text, for code and tool names. + * + * @param text the text + * @return the styled text + */ + public String cyan(String text) { + return style("36", text); + } + + /** + * Green text, for the marker of a running tool. + * + * @param text the text + * @return the styled text + */ + public String green(String text) { + return style("32", text); + } + + /** + * Yellow text, for questions that need an answer. + * + * @param text the text + * @return the styled text + */ + public String yellow(String text) { + return style("33", text); + } + + /** + * Red text, for errors and denials. + * + * @param text the text + * @return the styled text + */ + public String red(String text) { + return style("31", text); + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java new file mode 100644 index 000000000..9ef88ca99 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ApprovalMode.java @@ -0,0 +1,93 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.util.Locale; + +/** + * Whether a tool call that changes something asks before it runs. + * + *

The default is {@link #MANUAL}: the model may read freely, but every write and every shell + * command is confirmed on the console. {@link #AUTO} is what the {@code --auto} flag and the + * {@code [a]} answer of a single approval prompt switch to — it stays on for the rest of the session + * until {@code /mode manual} switches back. + */ +public enum ApprovalMode { + + /** Ask before every gated tool call. */ + MANUAL("⏸"), + + /** Run every tool call without asking. */ + AUTO("⏵⏵"); + + private final String symbol; + + ApprovalMode(String symbol) { + this.symbol = symbol; + } + + /** + * The lower-case name used on the console and in {@code /mode}. + * + * @return {@code "manual"} or {@code "auto"} + */ + public String label() { + return name().toLowerCase(Locale.ROOT); + } + + /** + * The glyph shown in front of the name on the status line. + * + *

Two transport symbols, the way the established terminal agents mark the same distinction: + * {@code ⏸} for a session that stops at every gated call, {@code ⏵⏵} for one that runs through. The + * name stays next to it — the glyph makes the mode findable at a glance, it does not replace the + * word. + * + * @return {@code "⏸"} or {@code "⏵⏵"} + */ + public String symbol() { + return symbol; + } + + /** + * Symbol and name together, as the status line and {@code /mode} print them. + * + * @return e.g. {@code "⏸ manual"} + */ + public String badge() { + return symbol + " " + label(); + } + + /** + * The next mode in the cycle, which is what the shift+tab shortcut switches to. + * + *

With two modes this is a toggle; it is written as a cycle so a third mode would need no + * change here. The order follows the declaration order, so it is the same order {@code /mode} + * lists. + * + * @return the following mode, wrapping around at the end + */ + public ApprovalMode next() { + ApprovalMode[] all = values(); + return all[(ordinal() + 1) % all.length]; + } + + /** + * Parse a mode name as typed by the user. + * + * @param text the name, in any case, optionally surrounded by whitespace + * @return the mode + * @throws IllegalArgumentException when {@code text} names no mode + */ + public static ApprovalMode parse(String text) { + String normalized = text == null ? "" : text.trim().toLowerCase(Locale.ROOT); + for (ApprovalMode mode : values()) { + if (mode.label().equals(normalized)) { + return mode; + } + } + throw new IllegalArgumentException("Unknown mode: " + text + " (expected manual or auto)"); + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java new file mode 100644 index 000000000..2955d33bd --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategy.java @@ -0,0 +1,157 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.util.List; +import java.util.Set; +import java.util.concurrent.atomic.AtomicReference; +import org.atmosphere.ai.StreamingSession; +import org.atmosphere.ai.approval.ApprovalResolution; +import org.atmosphere.ai.approval.ApprovalStrategy; +import org.atmosphere.ai.approval.PendingApproval; +import org.atmosphere.ai.approval.ToolApprovalPolicy; + +/** + * Asks on the console before a tool that changes something runs: {@code [y]es / [n]o / [a]uto}. + * + *

Atmosphere does the gating itself — {@code ToolExecutionHelper} consults the + * {@link ToolApprovalPolicy} and, when it says the call needs approval, blocks the tool loop on + * {@link #awaitApprovalDetailed} before the executor runs. A {@link ApprovalResolution#deny() denial} + * is handed to the model as the tool result {@code {"status":"cancelled","message":"Action cancelled + * by user"}}, so it can replan instead of believing the command ran; a timeout becomes + * {@code {"status":"timeout",…}}. Nothing of that is reimplemented here. + * + *

Answering {@code a} switches the whole session to {@link ApprovalMode#AUTO} and approves — the + * remaining calls of the running turn included, because the mode is read per call. {@code /mode + * manual} switches back. + * + *

Without a console the answer is "no". In one-shot mode ({@code --prompt}) there is no one + * to ask, so a gated call is denied and the reason is printed. That is deliberate (and what Claude + * Code's non-interactive mode does): silently auto-approving would make an unattended run the most + * permissive one. Pass {@code --auto} to run unattended. + */ +public final class ConsoleApprovalStrategy implements ApprovalStrategy { + + /** + * The tools that ask before they run: the shell plus everything that writes. Reading ( + * {@code ls}, {@code read_file}, {@code glob}, {@code grep}) is never gated — it cannot change + * the machine, and gating it would make the prompt so frequent that it stops being read. + */ + public static final Set GATED_TOOLS = + Set.of(ShellTool.TOOL_NAME, "write_file", "edit_file", "delete", "rename"); + + /** + * The tools that deliberately never ask, because they only read. + * + *

It exists so that "does not ask" is a decision rather than the absence of one. + * {@link #GATED_TOOLS} is a list of names, so a tool that is added or renamed upstream falls out + * of it silently and then runs unasked — the same class of quiet breakage as a stale exclusion + * file. {@code ConsoleApprovalStrategyTest} asserts every registered tool is in exactly one of the + * two sets, so such a change fails the build instead of the session. + */ + public static final Set READ_ONLY_TOOLS = Set.of("ls", "read_file", "glob", "grep"); + + private static final int ARGUMENT_PREVIEW_CHARS = 300; + + private final AtomicReference mode; + private final AgentTerminal terminal; + private final boolean interactive; + private final Ansi ansi; + private final TurnActivity activity; + + /** + * Create the strategy. + * + * @param mode the shared, mutable approval mode (also written by {@code /mode} and by an + * {@code [a]} answer) + * @param terminal where the question is asked + * @param interactive whether anybody can answer at all ({@code false} for a one-shot run) + * @param activity paused while the question is open, so the spinner does not redraw over it + */ + public ConsoleApprovalStrategy( + AtomicReference mode, AgentTerminal terminal, boolean interactive, TurnActivity activity) { + this.mode = mode; + this.terminal = terminal; + this.interactive = interactive; + this.ansi = terminal.ansi(); + this.activity = activity; + } + + /** + * The policy that decides which tools this strategy is asked about. + * + * @return a policy gating {@link #GATED_TOOLS} + */ + public static ToolApprovalPolicy policy() { + return ToolApprovalPolicy.custom(tool -> tool != null && GATED_TOOLS.contains(tool.name())); + } + + /** + * The gated tool names among {@code tools}, for the console. + * + * @param toolNames every offered tool name + * @return the names that will ask before running + */ + public static List gated(List toolNames) { + return toolNames.stream().filter(GATED_TOOLS::contains).toList(); + } + + @Override + public ApprovalOutcome awaitApproval(PendingApproval approval, StreamingSession session) { + return awaitApprovalDetailed(approval, session).outcome(); + } + + @Override + public ApprovalResolution awaitApprovalDetailed(PendingApproval approval, StreamingSession session) { + if (mode.get() == ApprovalMode.AUTO) { + return ApprovalResolution.approve(); + } + if (!interactive) { + terminal.line(ansi.red("✗ " + approval.toolName() + " " + preview(approval) + + " — denied: no console to ask (run with --auto to allow tools unattended)")); + return ApprovalResolution.deny(); + } + terminal.line(ansi.yellow("? " + approval.toolName()) + " " + ansi.dim(preview(approval))); + activity.pause(); + try { + return ask(approval); + } finally { + activity.resume(); + } + } + + private ApprovalResolution ask(PendingApproval approval) { + while (true) { + String answer = terminal.readKey(ansi.yellow(" allow? [y]es / [n]o / [a]uto (no more questions): ")); + if (answer == null) { + // input closed mid-turn: the same situation as having no console at all + terminal.line(" denied (input closed)"); + return ApprovalResolution.deny(); + } + switch (answer) { + // "\r" / "\n": Enter in raw mode, taken as yes like an empty line on a plain stream + case "y", "yes", "", "\r", "\n" -> { + return ApprovalResolution.approve(); + } + case "n", "no" -> { + return ApprovalResolution.deny(); + } + case "a", "auto" -> { + mode.set(ApprovalMode.AUTO); + terminal.line(" approval mode: auto (use /mode manual to ask again)"); + return ApprovalResolution.approve(); + } + default -> terminal.line(" please answer y, n or a"); + } + } + } + + private static String preview(PendingApproval approval) { + String arguments = String.valueOf(approval.arguments()); + return arguments.length() <= ARGUMENT_PREVIEW_CHARS + ? arguments + : arguments.substring(0, ARGUMENT_PREVIEW_CHARS) + "… (" + arguments.length() + " chars)"; + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java index 2da678424..63acb8835 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ConsoleSession.java @@ -4,7 +4,6 @@ package net.ladenthin.llama.atmosphere; -import java.io.PrintStream; import java.time.Duration; import java.util.List; import java.util.Map; @@ -13,6 +12,7 @@ import java.util.concurrent.TimeUnit; import org.atmosphere.ai.AiEvent; import org.atmosphere.ai.StreamingSession; +import org.atmosphere.ai.TokenUsage; import org.atmosphere.ai.fs.AgentFileSystem; import org.jspecify.annotations.Nullable; @@ -29,25 +29,57 @@ public final class ConsoleSession implements StreamingSession { private static final int RESULT_PREVIEW_CHARS = 400; - private final PrintStream out; + /** How much of a call's arguments the console shows; the model still gets them in full. */ + private static final int ARGUMENT_PREVIEW_CHARS = 200; + + /** How much of a single argument value survives, so one big one cannot hide the others. */ + private static final int ARGUMENT_VALUE_PREVIEW_CHARS = 80; + + private final AgentTerminal terminal; + private final Ansi ansi; + private final MarkdownConsole markdown; private final Map, Object> injectables; private final StringBuilder text = new StringBuilder(); private final List chunks = new CopyOnWriteArrayList<>(); private final CountDownLatch done = new CountDownLatch(1); private volatile @Nullable Throwable failure; private volatile int toolCalls; + private volatile long inputTokens; + private final List rounds = new CopyOnWriteArrayList<>(); + private volatile @Nullable String runningTool; + private volatile long runningSince; /** - * Create a session printing to {@code out}. + * Create a session writing to {@code terminal}. * - * @param out where streamed text and tool lines go + * @param terminal where streamed text and tool lines go * @param fileSystem the workspace-confined filesystem handed to the file tools */ - public ConsoleSession(PrintStream out, AgentFileSystem fileSystem) { - this.out = out; + public ConsoleSession(AgentTerminal terminal, AgentFileSystem fileSystem) { + this.terminal = terminal; + this.ansi = terminal.ansi(); + this.markdown = new MarkdownConsole(terminal::line, ansi); this.injectables = Map.of(AgentFileSystem.class, fileSystem); } + /** + * One tool call of this turn, kept so the next turn can see that it happened. + * + * @param name the tool + * @param argumentsJson the arguments as JSON + * @param result what the tool returned, already shortened + */ + public record ToolRound(String name, String argumentsJson, String result) {} + + /** + * The tool calls of this turn, in order. + * + * @return the rounds, empty when the model only wrote text + */ + public List rounds() { + return List.copyOf(rounds); + } + @Override public String sessionId() { return "console"; @@ -61,14 +93,52 @@ public Map, Object> injectables() { @Override public void send(String chunk) { chunks.add(chunk); + // The history keeps the raw text; only the console sees the rendered form. text.append(chunk); - out.print(chunk); - out.flush(); + markdown.append(chunk); } @Override public void sendMetadata(String key, Object value) { - // token usage, model id, tool-call argument deltas: not shown on the console + // model id, tool-call argument deltas: not shown on the console + } + + @Override + public void usage(TokenUsage usage) { + // The prompt of the last model call is what fills the context window -- the tokens generated + // in that call are part of the next call's input. Several calls happen per turn (one per tool + // round); the last one wins, which is the largest and the one the next turn continues from. + if (usage != null && usage.input() > 0) { + inputTokens = usage.input(); + } + } + + /** + * Everything this turn has produced so far, in characters: the streamed text plus every tool call + * with its result. + * + *

All of it is in the prompt of the next model call of the same turn — a tool round + * appends the call and its output to the conversation the server is sent. So this is what makes the + * context grow while the turn runs, and the status line adds it to the count it showed before the + * turn started instead of standing still until the next prompt. + * + * @return the character count + */ + public long producedChars() { + long chars = text.length(); + for (ToolRound round : rounds) { + chars += round.argumentsJson().length() + round.result().length(); + } + return chars; + } + + /** + * The input tokens of the last model call of this turn. + * + * @return the count, or {@code 0} when the endpoint reported no usage + */ + public long inputTokens() { + return inputTokens; } @Override @@ -78,8 +148,7 @@ public void progress(String message) { @Override public void complete() { - out.println(); - out.flush(); + markdown.flush(); done.countDown(); } @@ -94,9 +163,8 @@ public void complete(String summary) { @Override public void error(Throwable t) { failure = t; - out.println(); - out.println("[error] " + t); - out.flush(); + markdown.flush(); + terminal.line(ansi.red("[error] " + t)); done.countDown(); } @@ -110,29 +178,102 @@ public void emit(AiEvent event) { switch (event) { case AiEvent.ToolStart start -> { toolCalls++; - if (text.length() > 0 && text.charAt(text.length() - 1) != '\n') { - out.println(); - } - out.println("⚙ " + start.toolName() + " " + start.arguments()); - out.flush(); + runningTool = start.toolName(); + runningSince = System.nanoTime(); + rounds.add(new ToolRound(start.toolName(), String.valueOf(start.arguments()), "")); + markdown.flush(); + terminal.line(ansi.green("●") + " " + ansi.bold(start.toolName()) + " " + + ansi.dim(describeArguments(start.arguments()))); } case AiEvent.ToolResult result -> { - out.println("↳ " + preview(String.valueOf(result.result()))); - out.flush(); + runningTool = null; + recordResult(String.valueOf(result.result())); + terminal.line(ansi.dim(" ↳ " + preview(String.valueOf(result.result())))); } case AiEvent.ToolError error -> { - out.println("↳ error: " + error.error()); - out.flush(); + runningTool = null; + recordResult("error: " + error.error()); + terminal.line(ansi.red(" ↳ error: " + cut(String.valueOf(error.error()), RESULT_PREVIEW_CHARS))); } default -> StreamingSession.super.emit(event); } } + /** + * The tool that is executing right now, if any. + * + * @return the tool name, or {@code null} when the model is generating rather than running something + */ + public @Nullable String runningTool() { + return runningTool; + } + + /** + * How long the running tool has been running. + * + * @return the seconds since it started, or {@code 0} when nothing runs + */ + public long runningSeconds() { + return runningTool == null ? 0 : (System.nanoTime() - runningSince) / 1_000_000_000L; + } + + /** Attach a result to the round that is still waiting for one. */ + private void recordResult(String result) { + for (int i = rounds.size() - 1; i >= 0; i--) { + ToolRound round = rounds.get(i); + if (round.result().isEmpty()) { + rounds.set(i, new ToolRound(round.name(), round.argumentsJson(), result)); + return; + } + } + } + private static String preview(String value) { - String oneLine = value.replace("\r\n", "\n").replace('\n', ' '); - return oneLine.length() <= RESULT_PREVIEW_CHARS - ? oneLine - : oneLine.substring(0, RESULT_PREVIEW_CHARS) + " … (" + value.length() + " chars)"; + return cut(value, RESULT_PREVIEW_CHARS); + } + + /** + * The arguments of a call, short enough for one console line. + * + *

Every value is cut on its own before the whole thing is. Cutting only the rendered + * map would let one big argument push the others out of the line — a {@code write_file} call would + * then show half of the file and not the name of the file, which is the one thing worth seeing. + * + * @param arguments what the model passed, usually a map + * @return one line + */ + private static String describeArguments(Object arguments) { + if (!(arguments instanceof Map map)) { + return cut(String.valueOf(arguments), ARGUMENT_PREVIEW_CHARS); + } + StringBuilder rendered = new StringBuilder("{"); + for (Map.Entry entry : map.entrySet()) { + if (rendered.length() > 1) { + rendered.append(", "); + } + rendered.append(entry.getKey()) + .append('=') + .append(cut(String.valueOf(entry.getValue()), ARGUMENT_VALUE_PREVIEW_CHARS)); + } + return cut(rendered.append('}').toString(), ARGUMENT_PREVIEW_CHARS); + } + + /** + * Fold a value onto one line and cut it. + * + *

Both halves matter. The cut keeps a whole file out of the scrollback, and the folding keeps + * the pinned block intact: that block is sized in lines, so a single printed "line" + * carrying twenty newlines moves the screen twenty rows further than the terminal accounted for, + * and the block ends up drawn across the output. A {@code write_file} call whose arguments contain + * the file did exactly that. + * + * @param value the raw text + * @param max how many characters survive + * @return one line, with a note about what was left out + */ + private static String cut(String value, int max) { + String oneLine = value.replace("\r\n", " ").replace('\n', ' ').replace('\r', ' '); + return oneLine.length() <= max ? oneLine : oneLine.substring(0, max) + " … (" + value.length() + " chars)"; } /** diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java new file mode 100644 index 000000000..3eb1e935b --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/JLineTerminal.java @@ -0,0 +1,476 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.io.IOException; +import java.util.List; +import java.util.Locale; +import java.util.concurrent.BlockingQueue; +import java.util.concurrent.LinkedBlockingQueue; +import org.jline.keymap.KeyMap; +import org.jline.reader.Binding; +import org.jline.reader.EndOfFileException; +import org.jline.reader.LineReader; +import org.jline.reader.LineReaderBuilder; +import org.jline.reader.Reference; +import org.jline.reader.UserInterruptException; +import org.jline.reader.impl.completer.StringsCompleter; +import org.jline.terminal.Terminal; +import org.jline.terminal.TerminalBuilder; +import org.jline.utils.AttributedString; +import org.jline.utils.AttributedStyle; +import org.jline.utils.InfoCmp; +import org.jline.utils.Status; +import org.jspecify.annotations.Nullable; + +/** + * An {@link AgentTerminal} on a real terminal, via JLine: line editing and history at the prompt, tab + * completion of the commands, a status line pinned to the bottom of the window, and a prompt that is + * there at all times — including while the agent is working. + * + *

That last part is why a thread of its own owns the keyboard (see {@code startReading}): it sits + * in {@code readLine} for the whole session and puts what is typed on a queue, and every read in this + * class is served from that queue. Streamed output goes through {@link LineReader#printAbove(String)}, + * which scrolls it in above the prompt while the bottom block stays where it is — the one thing a + * plain {@code println} cannot do. Nothing is ever redrawn above that block, so the scrollback stays + * exactly as it was written. + * + *

{@link #open} returns {@code null} instead of throwing when there is no usable terminal (piped + * input, a "dumb" terminal, a missing native provider); the caller then uses {@link PlainTerminal}. + * Ctrl-C at the prompt clears the line and returns an empty one — it does not end the session; Ctrl-D + * ends input like end-of-file. + */ +public final class JLineTerminal implements AgentTerminal { + + /** + * What is queued in place of a line when input ends. + * + *

A queue of lines cannot carry "no more lines" as a value, and the reader thread is the only + * one that learns it. The sentinel is put back on every take, so end of input stays end of input + * for every later caller instead of turning back into "nothing typed yet". + */ + private static final String END_OF_INPUT = "\u0000end-of-input"; + + private final Terminal terminal; + private final LineReader reader; + private final Status status; + private final Ansi ansi; + private final BlockingQueue typed = new LinkedBlockingQueue<>(); + /** + * Held for the length of every write this class makes. + * + *

Two threads write here as a matter of course: the turn runs on a thread of Atmosphere's and + * prints its tool lines and streamed text, while the console thread refreshes the pinned block + * four times a second. Neither JLine's {@code printAbove} nor {@code Status.update} knows about + * the other, so without this their escape sequences interleave and a fragment lands on screen as + * text — a stray {@code 1H}, the tail of a cursor-position sequence, drawn into the middle of the + * rule. + */ + private final Object writing = new Object(); + + /** + * The block as it was last handed over, so it can be put back after the screen is wiped. + * + *

{@code Status} draws only what has changed, and a wipe does not change its content — it just + * removes it from the screen. Without keeping a copy there is nothing to redraw it from, and the + * bottom of the window stays empty until the next refresh happens to differ. + */ + private volatile List block = List.of(); + + /** + * The block as the caller asked for it, before being cut to the window. + * + *

Kept so it can be drawn from scratch after the screen is wiped. Handing the rendered rows + * back is not enough: {@code Status} draws the difference between them and what it believes is on + * screen, and after a wipe that belief is wrong in a way it cannot detect — measured, it emitted a + * single character where a whole block was missing. + */ + private volatile List requested = List.of(); + + private volatile boolean closed; + private @Nullable Thread input; + + private JLineTerminal(Terminal terminal, LineReader reader, Status status, Ansi ansi) { + this.terminal = terminal; + this.reader = reader; + this.status = status; + this.ansi = ansi; + } + + /** + * Open the system terminal. + * + * @param completions the words tab completes, e.g. the command names + * @return the terminal, or {@code null} when this is not an interactive terminal + */ + public static @Nullable JLineTerminal open(List completions) { + try { + Terminal terminal = TerminalBuilder.builder().system(true).build(); + if (terminal.getType().startsWith(Terminal.TYPE_DUMB)) { + terminal.close(); + return null; + } + return over(terminal, completions); + } catch (IOException | RuntimeException e) { + // No terminal, no native provider, a restricted environment: the plain console still works. + return null; + } + } + + /** + * Wrap a terminal that has already been built. + * + *

The seam the tests use: a terminal over a pair of streams renders exactly like a real one — + * same escape sequences, same line reader — so what the screen would look like can be asserted on + * the emitted bytes, without a TTY. That is the only way to catch a drawing bug like a prompt whose + * height does not match what the reader erases when the line is submitted. + * + * @param terminal the terminal to drive + * @param completions the words tab completes + * @return the wrapper + */ + static JLineTerminal over(Terminal terminal, List completions) { + LineReader reader = LineReaderBuilder.builder() + .terminal(terminal) + .completer(new StringsCompleter(completions)) + // The input line is erased when submitted and echoed above instead, so it does + // not pile up in the scrollback. It erases exactly ONE line, which is why the + // prompt has to stay one line -- see startReading. + .option(LineReader.Option.ERASE_LINE_ON_FINISH, true) + // "!" is a shell history expansion in the reader's default configuration, which + // silently rewrites a request like: git commit -m "fixed!" + .option(LineReader.Option.DISABLE_EVENT_EXPANSION, true) + .build(); + // No WINCH handler here on purpose. The line reader installs its own for as long as it is + // reading -- which is the whole session -- and it already resizes the pinned region itself + // (LineReaderImpl.handleSignal calls Status.resize). Adding one of ours only put a second + // writer on the terminal, on the signal thread, at the exact moment the reader was redrawing: + // the row of "> > > > >" after a resize got worse, not better, when it was tried. + return new JLineTerminal(terminal, reader, Status.getStatus(terminal), Ansi.detect()); + } + + @Override + public void line(String text) { + if (text.indexOf('\n') >= 0 || text.indexOf('\r') >= 0) { + // One call must be one screen line: the status block is sized in lines, so a multi-line + // string handed over as "a line" desynchronises the reserved region. Callers fold their + // text themselves; this is the backstop for the ones that forget. + text.lines().forEach(this::line); + return; + } + synchronized (writing) { + if (input != null) { + // Once the reader thread exists it owns the screen, and nothing may write around it -- + // not even in the instant between two reads. A direct write there cuts into the escape + // sequence the next read is emitting and half of it lands in the scrollback as text: + // a stray "[?1h" above the prompt was exactly that. + reader.printAbove(text); + } else { + terminal.writer().println(text); + terminal.writer().flush(); + } + } + } + + /** + * Start the one thread that owns the keyboard, if it is not running yet. + * + *

Why a thread of its own. The prompt is supposed to be there at all times — while the + * agent is working, not only between turns — and only a thread that sits in {@code readLine} can + * offer that. Everything else then writes through {@link LineReader#printAbove}, which scrolls + * text in above the prompt and leaves it where it is. + * + *

Why exactly one. A terminal has one keyboard, and two threads reading it take turns at + * random. So every read in this class — a request, an approval answer — is served from the one + * queue this thread fills, and nothing else ever reads the terminal. + * + */ + private synchronized void startReading() { + if (input != null) { + return; + } + scrollToBottom(); + input = new Thread( + () -> { + while (!closed) { + try { + // One line, and it has to stay one line: the reader erases a single line + // when the input is submitted, so a two-line prompt -- the rule above the + // input, which is what was tried first -- leaves that rule behind on + // every Enter, a column of them after a few. + String line = reader.readLine("> "); + if (!line.isBlank()) { + // The input line is erased on Enter, so the conversation would lose + // what was asked. Echoing it above keeps the transcript readable. + line(ansi.bold("› " + line.strip())); + } + typed.put(line); + } catch (UserInterruptException e) { + // Ctrl-C: drop what was typed and ask again, as before. + } catch (EndOfFileException e) { + typed.offer(END_OF_INPUT); // Ctrl-D + return; + } catch (InterruptedException e) { + Thread.currentThread().interrupt(); + return; + } catch (RuntimeException e) { + typed.offer(END_OF_INPUT); // the terminal is gone; stop reading it + return; + } + } + }, + "agent-input"); + input.setDaemon(true); + input.start(); + } + + @Override + public void clearScreen() { + String capability = terminal.getStringCapability(InfoCmp.Capability.clear_screen); + if (capability == null) { + return; // a terminal that cannot clear: better nothing than a guessed escape sequence + } + // The capability is terminfo source, not the sequence itself: it reads "\E[H\E[2J", with the + // escape spelled out. Writing it as it comes prints that text on the screen, which is what a + // test caught. Curses expands it the way terminal.puts would, but into a string this class can + // hand to the reader instead of writing behind its back. + StringBuilder expanded = new StringBuilder(); + org.jline.utils.Curses.tputs(expanded, capability); + String clear = expanded.toString(); + synchronized (writing) { + if (input == null) { + terminal.writer().print(clear); + terminal.writer().flush(); + return; + } + // Through the reader, like every other write once it exists: printAbove leaves the prompt + // redrawn and the reader's idea of the cursor intact, which writing the escape sequence + // around it would not. The blank rows put the input back on the last row, where clearing + // to the top-left corner has just moved it away from. + reader.printAbove(clear + System.lineSeparator().repeat(blankRows())); + // reset() makes it forget what it believes is on screen; without that the update below is + // a no-op, because the content it would draw is the content it thinks is already there. + List lines = requested; + // Three steps, and all three were needed to make the block come back after a wipe: + // forget the drawing state, hand over an empty block so nothing is believed to be on + // screen, then render it again. With only the first two, Status drew the difference it + // computed against a belief the wipe had invalidated -- measured as a single character + // where a whole block was missing. + status.reset(); + block = List.of(); + status.update(List.of()); + if (!lines.isEmpty()) { + updateStatus(lines); + } + } + } + + /** + * How many rows to fill so the cursor ends up on the last usable one. + * + * @return the count, never negative + */ + private int blankRows() { + return Math.max(0, terminal.getSize().getRows() - 1); + } + + /** + * Push the cursor to the last usable row, once, before the first prompt is drawn. + * + *

The line reader draws its prompt wherever the cursor happens to be, which is directly after + * the last thing printed; only the status block is pinned to the bottom of the window. On a + * half-empty screen that leaves the input floating in the middle with the block far below it, and + * it only looks like one piece once enough output has scrolled the cursor down by itself — which + * is why it looked right after a few turns and wrong at the start. + * + *

Scrolling the screen once at startup makes that the state from the first prompt on: from + * then on every line printed scrolls, so the cursor stays on the last row for the rest of the + * session. The cost is a screen of blank lines above the session, which is what any program that + * wants its input at the bottom without taking over the whole screen has to pay. + */ + private void scrollToBottom() { + synchronized (writing) { + for (int row = 0; row < blankRows(); row++) { + terminal.writer().println(); + } + terminal.writer().flush(); + } + } + + /** + * The rule that frames the input box, as wide as the window. + * + * @return a line of {@code ─} + */ + private String rule() { + return "─".repeat(Math.max(10, terminal.getSize().getColumns() - 1)); + } + + @Override + public @Nullable String readLine(String prompt) { + // The prompt is the box this terminal draws, so the caller's is ignored: a "you> " in front of + // an input line that already sits in a frame is noise, and the frame cannot be handed in as a + // string because it is rebuilt on every window size. + startReading(); + try { + return take(typed.take()); + } catch (InterruptedException e) { + Thread.currentThread().interrupt(); + return null; + } + } + + /** + * Hand out a queued line, keeping the end-of-input marker in the queue. + * + * @param line what came off the queue + * @return the line, or {@code null} when input has ended + */ + private @Nullable String take(String line) { + if (END_OF_INPUT.equals(line)) { + typed.offer(END_OF_INPUT); + return null; + } + return line; + } + + @Override + public boolean hasPendingInput() { + // A blank line is not a request, so it must not count: the REPL skips it, and counting it + // meant that holding Enter cancelled one turn per keystroke and produced nothing. It stays in + // the queue, because an empty answer to an approval question means yes. + return typed.stream().anyMatch(line -> !line.isBlank()); + } + + @Override + public @Nullable String readKey(String prompt) { + // The prompt belongs to the input thread and cannot be changed while it is waiting, so the + // question is printed as an ordinary line above it and answered in the same input line as + // everything else. That costs an Enter, and buys the one thing worth more: a prompt that is + // there while the agent works. A single-key read here would need a second reader on the same + // terminal, and the two would take turns at random. + line(prompt); + startReading(); + try { + String answer = take(typed.take()); + return answer == null ? null : answer.trim().toLowerCase(Locale.ROOT); + } catch (InterruptedException e) { + Thread.currentThread().interrupt(); + return null; + } + } + + @Override + public void status(List lines) { + synchronized (writing) { + updateStatus(lines); + } + } + + private void updateStatus(List lines) { + requested = List.copyOf(lines); + if (lines.isEmpty()) { + if (block.isEmpty()) { + return; + } + block = List.of(); + status.update(List.of()); + return; + } + // The rule on top of this block is the one that separates the conversation from the input. + // It is the ONLY one drawn: a rule above the input line cannot be pinned (JLine's status + // region is below the prompt, never above it) and drawing it as output leaves one behind in + // the scrollback per turn, which is what "the line keeps travelling along" was. + // One row must never wrap: a wrapped row occupies two screen lines, the reserved region is + // sized in lines, and everything below it is then drawn in the wrong place -- which is how a + // long summary tore the block apart. + int width = Math.max(10, terminal.getSize().getColumns() - 1); + List rows = new java.util.ArrayList<>(); + rows.add(new AttributedString(rule(), AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT))); + for (String line : lines) { + rows.add( + new AttributedString(fit(line, width), AttributedStyle.DEFAULT.foreground(AttributedStyle.BRIGHT))); + } + block = List.copyOf(rows); + status.update(rows); + } + + /** + * Cut a status row to the window width. + * + * @param text the row + * @param width how many characters fit + * @return the row, ending in {@code …} when it had to be cut + */ + static String fit(String text, int width) { + // Counted in screen columns, not characters. An icon or an emoji occupies two columns and one + // character, so cutting by character length lets a row come out wider than the window, wrap + // onto a second screen line, and push everything below the reserved region out of place -- + // the same tearing a long summary caused, arriving through a different door. + AttributedString measured = new AttributedString(text); + if (measured.columnLength() <= width) { + return text; + } + return measured.columnSubSequence(0, Math.max(1, width - 1)).toString() + "…"; + } + + /** + * What a terminal sends for shift+tab: {@code ESC [ Z}, "backtab" (CSI Z). + * + *

It is bound literally as well as through terminfo, because JLine's Windows terminfo + * (windows-vtp.caps) declares no key_btab at all — so the capability + * lookup yields nothing there while the terminal itself, in virtual-terminal input mode, does + * send the sequence. + */ + private static final String BACKTAB = "\033[Z"; + + /** The name the cycle action is registered under; a widget is addressed by name, not by object. */ + private static final String CYCLE_MODE_WIDGET = "jllama-cycle-approval-mode"; + + @Override + public boolean onCycleMode(Runnable action) { + KeyMap keys = reader.getKeyMaps().get(LineReader.MAIN); + if (keys == null) { + return false; + } + reader.getWidgets().put(CYCLE_MODE_WIDGET, () -> { + action.run(); + return true; + }); + Reference widget = new Reference(CYCLE_MODE_WIDGET); + String fromTerminfo = KeyMap.key(terminal, InfoCmp.Capability.key_btab); + if (fromTerminfo != null) { + keys.bind(widget, fromTerminfo); + } + keys.bind(widget, BACKTAB); + return true; + } + + @Override + public boolean pinsStatus() { + return true; + } + + @Override + public Ansi ansi() { + return ansi; + } + + @Override + public void close() { + closed = true; + Thread reading = input; + if (reading != null) { + reading.interrupt(); + } + try { + status.update(List.of()); + status.close(); + terminal.close(); + } catch (IOException | RuntimeException e) { + // closing a terminal that is already gone must not fail the session + } + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java index e551b26af..76d39d070 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LocalAgent.java @@ -14,12 +14,14 @@ import java.time.Duration; import java.util.ArrayList; import java.util.List; +import java.util.Optional; +import java.util.concurrent.atomic.AtomicReference; import net.ladenthin.llama.LlamaModel; +import net.ladenthin.llama.args.LogFormat; import net.ladenthin.llama.parameters.ModelParameters; import net.ladenthin.llama.server.OpenAiCompatServer; import net.ladenthin.llama.server.OpenAiServerConfig; import org.atmosphere.ai.fs.AgentFileSystem; -import org.atmosphere.ai.fs.FileSystemTools; import org.atmosphere.ai.fs.WorkspaceAgentFileSystem; import org.atmosphere.ai.llm.ChatMessage; import org.atmosphere.ai.tool.ToolDefinition; @@ -57,7 +59,49 @@ public final class LocalAgent { /** The {@code {shell_section}} without {@code --allow-shell}. */ static final String NO_SHELL_PROMPT = "system-prompt-no-shell.txt"; + /** The {@code /help} overview. */ + static final String HELP_TEXT = "help.txt"; + + /** The per-step message of {@code /loop}; placeholders {@code {task}}, {@code {file}}, {@code {check_hint}}. */ + static final String LOOP_PROMPT = "loop-prompt.txt"; + + /** The skeleton written to {@link TaskLoop#LOOP_FILE}; placeholder {@code {task}}. */ + static final String LOOP_FILE_TEMPLATE = "loop-file-template.md"; + + /** The instructions {@code /compact} sends; placeholder {@code {focus}}. */ + static final String COMPACT_PROMPT = "compact-prompt.txt"; + + /** The system prompt of the summarizing turn: no tools, no agent role, just condense. */ + static final String COMPACT_SYSTEM_PROMPT = + "You summarize a conversation between a user and a coding assistant. Follow the user's" + + " instructions exactly and answer with the summary only."; + + /** How the user message of a compacted history begins; also how a repeat compaction is detected. */ + static final String SUMMARY_PREFIX = "Summary of the conversation so far:"; + + /** How much of a tool result is kept in the history of later turns. */ + private static final int HISTORY_RESULT_CHARS = 400; + + /** How often the activity line is refreshed while a turn runs. */ + private static final Duration ACTIVITY_INTERVAL = Duration.ofMillis(250); + + /** The usual rule of thumb, used wherever a token count has to be guessed from text. */ + private static final int CHARS_PER_TOKEN = 4; + + /** The spinner shown in the activity line. */ + private static final String ACTIVITY_FRAMES = "⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏"; + + /** The first line of the block while nothing is running. */ + static final String IDLE_LINE = "… waiting for input …"; + + /** The whimsical words the activity line picks from, one per turn. */ + static final String SPINNER_WORDS = "spinner-words.txt"; + private static final Duration SHELL_TIMEOUT = Duration.ofSeconds(120); + + /** A line that is exactly this opens and closes a multi-line message. */ + static final String BLOCK_FENCE = "\"\"\""; + private static final int SHELL_MAX_OUTPUT_CHARS = 20_000; private LocalAgent() {} @@ -103,6 +147,7 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream throws Exception { LlamaModel model = null; OpenAiCompatServer server = null; + AgentTerminal terminal = null; String baseUrl = options.getBaseUrl(); try { if (options.getModelPath() != null) { @@ -125,9 +170,17 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream } AgentFileSystem fileSystem = new WorkspaceAgentFileSystem(options.getWorkspace(), AgentFileSystem.Limits.defaults()); - List tools = new ArrayList<>(FileSystemTools.all()); + // Our own read_file/edit_file/grep replace the framework's (see WorkspaceTools); the + // read tracker is what lets an edit insist the file was read first. + List tools = new ArrayList<>(WorkspaceTools.all(new WorkspaceTools.ReadTracker())); if (options.isAllowShell()) { - tools.add(ShellTool.definition(options.getWorkspace(), SHELL_TIMEOUT, SHELL_MAX_OUTPUT_CHARS)); + // Live output: a two-minute build has to show that it is doing something. + AgentTerminal console = terminal; + tools.add(ShellTool.definition( + options.getWorkspace(), + SHELL_TIMEOUT, + SHELL_MAX_OUTPUT_CHARS, + line -> console.line(console.ansi().dim(" │ " + line)))); } AgentRunner runner = new AgentRunner( baseUrl, @@ -142,33 +195,188 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream + " tools=" + runner.toolNames()); List history = new ArrayList<>(); + ToolCallLog callLog = new ToolCallLog(); + // What was said, with the time, kept apart from the conversation the model is sent: that + // one is rewritten by /compact and has no timestamps at all. + Transcript transcript = new Transcript(options.getTranscript()); + AtomicReference mode = + new AtomicReference<>(options.isAuto() ? ApprovalMode.AUTO : ApprovalMode.MANUAL); + boolean interactive = options.getPrompt() == null && input != null; + BufferedReader reader = input == null ? null : new BufferedReader(input); + terminal = usesFullTerminal(options, interactive) ? JLineTerminal.open(commandNames()) : null; + if (terminal == null) { + terminal = new PlainTerminal(out, reader, Ansi.detect()); + } + captureNativeLog(terminal, options); + // One-shot runs have nobody at the keyboard, so the strategy gets no console and denies + // gated calls unless --auto was passed (see ConsoleApprovalStrategy). + TurnActivity activity = new TurnActivity(); + runner.approval( + new ConsoleApprovalStrategy(mode, terminal, interactive, activity), + ConsoleApprovalStrategy.policy()); + int contextSize = options.getModelPath() != null + ? options.getCtxSize() + : ServerProps.contextSize(baseUrl, options.getApiKey()); + if (options.getPrompt() != null) { - return turn(runner, fileSystem, options.getPrompt(), history, out) ? 0 : 1; + return turn( + runner, + fileSystem, + options.getPrompt(), + history, + terminal, + callLog, + 1, + ignored -> "", + activity) + .failure() + == null + ? 0 + : 1; } - if (input == null) { + if (reader == null) { err.println("No interactive input available; pass --prompt ."); return 2; } - BufferedReader reader = new BufferedReader(input); - err.println("Interactive mode: type a request, /clear to drop the history, /exit to quit."); + // Held rather than kept in locals so the shift+tab shortcut, which fires from inside the + // line reader, can re-render the status row with the numbers the last turn left behind. + java.util.concurrent.atomic.AtomicLong inputTokens = new java.util.concurrent.atomic.AtomicLong(); + java.util.concurrent.atomic.AtomicBoolean estimated = new java.util.concurrent.atomic.AtomicBoolean(); + AgentTerminal console = terminal; + java.util.function.Supplier idleStatus = () -> StatusLine.render( + options.getWorkspace(), + mode.get(), + inputTokens.get(), + estimated.get(), + contextSize, + tools.size(), + options.getModelId(), + options.getModelPath() == null); + boolean shortcut = console.onCycleMode(() -> { + mode.set(mode.get().next()); + if (console.pinsStatus()) { + console.status(List.of(IDLE_LINE, idleStatus.get())); + } else { + console.line(console.ansi().dim(idleStatus.get())); + } + }); + err.println("Interactive mode: type a request, /help for the commands." + + (shortcut ? " shift+tab switches the approval mode." : "")); + int turnNumber = 0; + String pendingNote = ""; + String lastMessage = ""; while (true) { - out.print("you> "); - out.flush(); - String line = reader.readLine(); - if (line == null || line.trim().equals("/exit") || line.trim().equals("/quit")) { + // Pinned to the bottom of the window on a real terminal; printed above the prompt on a + // plain stream, where there is nothing to pin and a repeatedly refreshed line would + // just fill a piped log. + String status = StatusLine.render( + options.getWorkspace(), + mode.get(), + inputTokens.get(), + estimated.get(), + contextSize, + tools.size(), + options.getModelId(), + options.getModelPath() == null); + if (terminal.pinsStatus()) { + terminal.status(List.of(IDLE_LINE, status)); + } else { + terminal.line(terminal.ansi().dim(status)); + } + String line = terminal.readLine("you> "); + if (line == null) { return 0; } + if (BLOCK_FENCE.equals(line.strip())) { + line = readBlock(terminal); + if (line == null) { + return 0; + } + } if (line.trim().isEmpty()) { continue; } - if (line.trim().equals("/clear")) { - history.clear(); - err.println("(history cleared)"); - continue; + Optional command = SlashCommands.parse(line); + if (command.isPresent()) { + if (command.get().command() == SlashCommands.Command.EXIT) { + return 0; + } + if (command.get().command() == SlashCommands.Command.RETRY) { + // Handled here rather than with the other commands, because it is not a command + // that answers something: it runs a turn, which only this loop can do. + if (lastMessage.isEmpty()) { + terminal.line("nothing to retry yet"); + continue; + } + dropLastExchange(history, lastMessage); + transcript.add(Transcript.Kind.NOTE, "retrying: " + lastMessage); + line = lastMessage; + } else { + inputTokens.set(handleCommand( + command.get(), + runner, + fileSystem, + history, + mode, + options, + contextSize, + inputTokens.get(), + estimated.get(), + terminal, + callLog, + transcript)); + continue; + } + } + if (!line.equals(lastMessage)) { + transcript.add(Transcript.Kind.USER, line); + } + lastMessage = line; + turnNumber++; + // Compact BEFORE the request that would overflow, not after it: afterwards the + // oversized request has already been sent, which is the one thing to avoid. + if (needsCompaction( + options, contextSize, estimateTokens(systemPrompt(options), history) + line.length() / 4)) { + terminal.line(terminal.ansi().yellow("(context nearly full — compacting first)")); + inputTokens.set(compact(runner, fileSystem, history, "", terminal)); + estimated.set(true); + } + // What the request carries before the model has answered anything; the turn adds to it. + long baseTokens = estimateTokens(systemPrompt(options), history) + line.length() / CHARS_PER_TOKEN; + ConsoleSession completed = turn( + runner, + fileSystem, + pendingNote.isEmpty() + ? line + : pendingNote + System.lineSeparator() + System.lineSeparator() + line, + history, + terminal, + callLog, + turnNumber, + running -> StatusLine.render( + options.getWorkspace(), + mode.get(), + liveTokens(baseTokens, running), + running.inputTokens() == 0, + contextSize, + tools.size(), + options.getModelId(), + options.getModelPath() == null), + activity); + for (ConsoleSession.ToolRound round : completed.rounds()) { + transcript.add( + Transcript.Kind.TOOL, round.name() + " " + round.argumentsJson() + " -> " + round.result()); } - turn(runner, fileSystem, line, history, out); + transcript.add(Transcript.Kind.AGENT, completed.text()); + pendingNote = toolNote(completed.rounds()); + estimated.set(completed.inputTokens() == 0); + inputTokens.set( + estimated.get() ? estimateTokens(systemPrompt(options), history) : completed.inputTokens()); } } finally { + if (terminal != null) { + terminal.close(); + } if (server != null) { server.close(); } @@ -178,17 +386,677 @@ static int run(AgentOptions options, java.io.@Nullable Reader input, PrintStream } } - private static boolean turn( - AgentRunner runner, AgentFileSystem fileSystem, String message, List history, PrintStream out) + /** + * Run {@code /loop}: work on one task step by step until it is done. + * + *

Two things are settled before the first step. The loop needs the {@link ApprovalMode#AUTO} + * mode — a run that asks before every write is not a loop, it is a conversation — so a manual + * session is asked once and left alone if the answer is no. And the history is not the + * loop's memory: {@link TaskLoop} drops it every step and keeps the state in a file, so nothing of + * the current conversation is used or changed here. + * + * @param runner the runner + * @param fileSystem the workspace filesystem + * @param terminal the console + * @param options the agent options, for the workspace + * @param mode the approval mode, possibly switched to auto here + * @param arguments everything after {@code /loop} + * @param callLog records the steps' tool calls for {@code /calls} + * @throws InterruptedException if interrupted while a step runs + */ + private static void loop( + AgentRunner runner, + AgentFileSystem fileSystem, + AgentTerminal terminal, + AgentOptions options, + AtomicReference mode, + String arguments, + ToolCallLog callLog) throws InterruptedException { - ConsoleSession session = new ConsoleSession(out, fileSystem); - runner.run(message, history, session); - boolean finished = session.await(TURN_TIMEOUT); + LoopOptions loopOptions; + try { + loopOptions = LoopOptions.parse(arguments); + } catch (IllegalArgumentException e) { + terminal.line(e.getMessage()); + return; + } + if (mode.get() != ApprovalMode.AUTO) { + String answer = + terminal.readKey("a loop cannot stop at every question — switch to auto for it? [y]es / [n]o: "); + if (answer == null || !(answer.startsWith("y") || answer.isEmpty())) { + terminal.line("loop: cancelled (use /mode auto to allow it)"); + return; + } + mode.set(ApprovalMode.AUTO); + } + if (!TaskLoop.canKeepNotes(runner.toolNames())) { + terminal.line("loop: needs the read_file and write_file tools to keep its notes"); + return; + } + TaskLoop.Outcome outcome = TaskLoop.run( + runner, + fileSystem, + terminal, + options.getWorkspace(), + loopOptions, + () -> false, + TaskLoop.DEFAULT_BUDGET, + callLog); + terminal.line( + outcome.completed() + ? terminal.ansi().green("loop: " + outcome.reason()) + : terminal.ansi().yellow("loop: " + outcome.reason())); + } + + /** + * Route llama.cpp's own log through the console instead of letting it write to stderr. + * + *

**This is what destroys a pinned block, and nothing on the Java side can defend against it.** + * With an in-process model the server logs to stderr, which is the same console; those writes go + * around the line reader, so the terminal scrolls lines JLine never sees and its reserved region + * ends up somewhere else than it believes. What that looks like: a warning printed into the middle + * of the input line (`> d1.19.029.542 W srv stop: cancel task`), and after a few of them the + * block is gone. Routing the log through {@link AgentTerminal#line} puts it under the same lock as + * everything else, so it scrolls in above the prompt like any other output. + * + *

Only for an in-process model: with {@code --base-url} the server is another process and its + * log is its own business, and calling this would load the native library for nothing. + * + * @param terminal where the log lines go + * @param options the parsed command line + */ + private static void captureNativeLog(AgentTerminal terminal, AgentOptions options) { + if (options.getModelPath() == null) { + return; + } + Ansi ansi = terminal.ansi(); + LlamaModel.setLogger(LogFormat.TEXT, (level, message) -> { + String text = message == null ? "" : message.strip(); + if (!text.isEmpty()) { + terminal.line(ansi.dim(text)); + } + }); + } + + /** + * Whether to drive the cursor-controlling console rather than the line-oriented one. + * + *

Both consoles are kept, and this is the only place that decides between them. The rich one + * needs someone at a terminal and permission to position the cursor; {@code --plain} + * withholds the second even when the first is true, which is what a session that is piped, + * logged, recorded, or run through something that only forwards lines needs. A run without an + * interactive input has no use for it either way. + * + * @param options the parsed command line + * @param interactive whether there is someone typing + * @return {@code true} to try the full terminal + */ + /** + * Take the last exchange out of the conversation, so a retry asks again instead of following on. + * + *

Leaving the failed answer in place would be the opposite of a retry: the model would see what + * it said last time and, being a model, would say it again. The question is removed with it, + * because the turn that follows adds it back. + * + *

Only a trailing exchange that really is the one being retried is touched — a history that was + * just replaced by a summary, or one that never got an answer, is left alone. + * + * @param history the conversation, modified in place + * @param message the question being asked again + */ + static void dropLastExchange(List history, String message) { + if (!history.isEmpty() + && "assistant".equals(history.get(history.size() - 1).role())) { + history.remove(history.size() - 1); + } + if (!history.isEmpty()) { + ChatMessage last = history.get(history.size() - 1); + if ("user".equals(last.role()) && message.equals(last.content())) { + history.remove(history.size() - 1); + } + } + } + + /** + * Read a block of lines, the way a fenced code block is written. + * + *

A console reads a line at a time, and Enter sends it — which makes pasting a stack trace or a + * function into the prompt impossible without it becoming several questions. A line that is + * exactly {@value #BLOCK_FENCE} starts a block and the next one closes it; everything between is + * one message, newlines and all. + * + *

Chosen over a key combination because it works in both consoles, survives a paste (the fence + * arrives as part of the pasted text), and needs nothing from the terminal. + * + * @param terminal where the lines come from + * @return the block, or {@code null} at end of input + */ + static @Nullable String readBlock(AgentTerminal terminal) { + StringBuilder block = new StringBuilder(); + while (true) { + String line = terminal.readLine("... "); + if (line == null) { + // End of input inside a block: what was collected is still a question worth asking, + // but there is nobody left to answer it, so the session ends as it would anyway. + return null; + } + if (BLOCK_FENCE.equals(line.strip())) { + return block.toString(); + } + if (block.length() > 0) { + block.append(System.lineSeparator()); + } + block.append(line); + } + } + + /** + * Read a saved transcript back in, as the conversation and as the record. + * + *

Only the questions and the answers become messages again. Tool calls and session notes are + * kept in the record but not replayed to the model: a tool result out of its round is not + * something any chat template has a place for, and inventing one would be worse than leaving the + * model to call the tool again if it needs to. + * + *

The file is resolved against the workspace when it is not an absolute path, so {@code /load + * session.txt} finds what {@code /save session.txt} wrote. + * + * @param name the file + * @param options the parsed command line, for the workspace + * @param history the conversation, replaced + * @param transcript the record, replaced + * @param terminal where to report + * @return the estimated input tokens of the loaded conversation + */ + private static long load( + String name, + AgentOptions options, + List history, + Transcript transcript, + AgentTerminal terminal) { + java.nio.file.Path file = java.nio.file.Path.of(name.strip()); + if (!file.isAbsolute()) { + file = options.getWorkspace().resolve(file); + } + List loaded; + try { + loaded = Transcript.parse(java.nio.file.Files.readString(file, java.nio.charset.StandardCharsets.UTF_8)); + } catch (java.io.IOException e) { + terminal.line("cannot read " + file + ": " + e.getMessage()); + return estimateTokens(systemPrompt(options), history); + } + if (loaded.isEmpty()) { + terminal.line(file + " holds no transcript entries"); + return estimateTokens(systemPrompt(options), history); + } + history.clear(); + int messages = 0; + for (Transcript.Entry entry : loaded) { + switch (entry.kind()) { + case USER -> { + history.add(ChatMessage.user(entry.text())); + messages++; + } + case AGENT -> { + history.add(ChatMessage.assistant(entry.text())); + messages++; + } + default -> { + // kept in the record, not replayed as a message + } + } + } + transcript.replaceWith(loaded); + transcript.add(Transcript.Kind.NOTE, "loaded " + file); + terminal.line("loaded " + loaded.size() + " entries from " + file + " (" + messages + + " of them replayed to the model)"); + return estimateTokens(systemPrompt(options), history); + } + + static boolean usesFullTerminal(AgentOptions options, boolean interactive) { + return interactive && !options.isPlain(); + } + + static ConsoleSession turn( + AgentRunner runner, + AgentFileSystem fileSystem, + String message, + List history, + AgentTerminal terminal, + ToolCallLog callLog, + int turnNumber, + java.util.function.Function stateLine, + TurnActivity activity) + throws InterruptedException { + ConsoleSession session = new ConsoleSession(terminal, fileSystem); + // The turn runs on a thread of Atmosphere's, not this one: execute() blocks until the whole + // turn including every tool round is done, which would leave nobody to drive the activity + // line -- the block sat on "… waiting for input …" for entire turns until this changed. + // start() adds the half that makes the always-present prompt worth having: the handle can + // close the stream the model is answering on, so a request typed mid-turn takes effect now. + org.atmosphere.ai.ExecutionHandle handle; + try { + handle = runner.start(message, history, session); + } catch (RuntimeException e) { + session.error(e); + handle = org.atmosphere.ai.ExecutionHandle.completed(); + } + TurnEnd end = awaitWithActivity(session, terminal, () -> stateLine.apply(session), activity, handle); history.add(ChatMessage.user(message)); + callLog.add(turnNumber, session.rounds()); if (!session.text().isEmpty()) { history.add(ChatMessage.assistant(session.text())); } - return finished && session.failure() == null; + if (end == TurnEnd.TIMED_OUT) { + session.error(new IllegalStateException("turn did not finish within " + TURN_TIMEOUT)); + } + return session; + } + + /** + * Put this turn's tool calls in front of its answer, so the next turn can see they happened. + * + *

Why this matters more than it looks. Without it the history holds the user's messages + * and the model's prose, and nothing else — so from the third or fourth turn on, a small model + * sees only its own paragraphs and no evidence that it ever used a tool. It then continues that + * pattern: it describes creating a file and running a build, reports an exit code, and + * writes nothing at all. That is not hypothetical; it happened on a real session, with the model + * inventing test results and a jar that never existed. + * + *

Where it goes, and why not somewhere more obvious. In front of the next user + * message. The first attempt put it in front of the assistant's own answer, and the model + * promptly copied it into its next reply — the user read "(tools I actually ran this turn: …)" as + * the first line of an answer, because text attributed to the assistant is text a model imitates. + * A system message mid-history would be cleaner still, but not every chat template accepts one: + * Mistral's requires strict user/assistant alternation and Gemma has no system role at all. + * Riding along with the next user message keeps the sequence template-safe everywhere. + * + *

Why a note rather than real {@code tool_calls} messages. Atmosphere's + * {@code AbstractAgentRuntime.assembleMessages} rebuilds every history entry as + * {@code new ChatMessage(role, content)} — the tool-call array and the tool-call id are dropped on + * the way out. Protocol-faithful replay is therefore impossible through the framework's history; + * what survives is the content, so the evidence goes there. It is also cheaper: one line per call + * instead of a message pair, with the result cut to {@value #HISTORY_RESULT_CHARS} characters. + * + * @param rounds the tool calls of the finished turn + * @return the note, or an empty string when no tool ran + */ + static String toolNote(List rounds) { + if (rounds.isEmpty()) { + return ""; + } + StringBuilder note = new StringBuilder("Record of the tools that actually ran in the previous turn." + + " This is a log for your reference; do not repeat it and do not mention it."); + for (ConsoleSession.ToolRound round : rounds) { + String result = round.result().replace("\r\n", " ").replace('\n', ' '); + note.append(System.lineSeparator()) + .append("- ") + .append(round.name()) + .append(" ") + .append(round.argumentsJson()) + .append(" -> ") + .append( + result.length() <= HISTORY_RESULT_CHARS + ? result + : result.substring(0, HISTORY_RESULT_CHARS) + " …[cut]"); + } + return note.toString(); + } + + /** + * Run one REPL command. + * + * @param command the parsed command + * @param runner the runner (used by {@code /compact}) + * @param fileSystem the workspace filesystem (used by {@code /compact}'s session) + * @param history the conversation history, modified in place by {@code /clear} and {@code /compact} + * @param mode the approval mode, modified in place by {@code /mode} + * @param options the options, for the status output + * @param contextSize the context window in tokens, or {@link StatusLine#UNKNOWN_CONTEXT} + * @param inputTokens the input tokens of the last turn + * @param estimated whether that number is an estimate + * @param terminal the console + * @param callLog every tool call of the session, for {@code /calls} + * @param transcript what was said, with the time, for {@code /save} + * @return the input tokens to show from now on (unchanged, or the summary's after {@code /compact}) + * @throws InterruptedException if interrupted while a summary is generated + */ + private static long handleCommand( + SlashCommands command, + AgentRunner runner, + AgentFileSystem fileSystem, + List history, + AtomicReference mode, + AgentOptions options, + int contextSize, + long inputTokens, + boolean estimated, + AgentTerminal terminal, + ToolCallLog callLog, + Transcript transcript) + throws InterruptedException { + switch (command.command()) { + case HELP -> prompt(HELP_TEXT).lines().forEach(terminal::line); + case CLEAR -> { + history.clear(); + // "Forget this session" has to mean the record too, or the word is not true. + transcript.clear(); + // The screen goes with it: what is still on it is a conversation the model no longer + // has, which reads as if it were still in play. + terminal.clearScreen(); + terminal.line("(history cleared)"); + } + case CLS -> terminal.clearScreen(); + case LOAD -> { + if (!command.hasArguments()) { + terminal.line("say which file: /load "); + return inputTokens; + } + return load(command.arguments(), options, history, transcript, terminal); + } + case SAVE -> { + try { + java.nio.file.Path written = transcript.save( + options.getWorkspace(), command.hasArguments() ? command.arguments() : null); + terminal.line("transcript: " + transcript.size() + " entries -> " + written); + } catch (java.io.IOException e) { + terminal.line("could not write the transcript: " + e.getMessage()); + } + } + case CALLS -> callLog.render().lines().forEach(terminal::line); + case TOOLS -> { + terminal.line("tools: " + String.join(", ", runner.toolNames())); + terminal.line("asks before running (manual mode): " + + String.join(", ", ConsoleApprovalStrategy.gated(runner.toolNames()))); + } + case MODE -> { + if (command.hasArguments()) { + try { + mode.set(ApprovalMode.parse(command.arguments())); + } catch (IllegalArgumentException e) { + terminal.line(e.getMessage()); + return inputTokens; + } + } + terminal.line("approval mode: " + mode.get().badge()); + } + case STATUS -> { + terminal.line(StatusLine.render( + options.getWorkspace(), + mode.get(), + inputTokens, + estimated, + contextSize, + runner.toolNames().size(), + options.getModelId(), + options.getModelPath() == null)); + terminal.line("workspace: " + options.getWorkspace()); + terminal.line("history: " + history.size() + " messages"); + } + case COMPACT -> { + // The record is deliberately untouched: compacting rewrites what the model is sent, + // not what happened. Only the fact that it happened is worth a line. + transcript.add(Transcript.Kind.NOTE, "compacted the conversation"); + return compact(runner, fileSystem, history, command.arguments(), terminal); + } + case LOOP -> loop(runner, fileSystem, terminal, options, mode, command.arguments(), callLog); + case EXIT -> { + // handled by the caller, which has to return from the loop + } + } + return inputTokens; + } + + /** + * Wait for a turn while the status line shows that something is happening. + * + *

A local model can think for a while before the first token arrives, and a silent console is + * indistinguishable from a hung one. The line is rewritten in place (it is the pinned status area, + * not the scrollback), so nothing the user has already read moves. + * + * @param session the running turn + * @param terminal the console + * @param stateLine the second row of the block, asked again on every redraw so the context figure + * moves while the turn runs rather than standing still until the next prompt + * @param activity paused while an approval question is open + * @param handle stops the running turn when the user types instead of waiting + * @return how the turn ended + * @throws InterruptedException if interrupted while waiting + */ + /** How a turn ended: on its own, because the user typed something, or because it ran too long. */ + enum TurnEnd { + /** The model produced its final answer (or errored). */ + FINISHED, + /** The user typed while it was working; the turn was cut short and that line is next. */ + INTERRUPTED, + /** Nothing arrived within {@link #TURN_TIMEOUT}. */ + TIMED_OUT + } + + static TurnEnd awaitWithActivity( + ConsoleSession session, + AgentTerminal terminal, + java.util.function.Supplier stateLine, + TurnActivity activity, + org.atmosphere.ai.ExecutionHandle handle) + throws InterruptedException { + long start = System.nanoTime(); + int frame = 0; + // one word per turn, not per frame: a word that changes ten times a second is noise + String word = spinnerWord(); + while (!session.await(ACTIVITY_INTERVAL)) { + long seconds = (System.nanoTime() - start) / 1_000_000_000L; + if (seconds > TURN_TIMEOUT.toSeconds()) { + return TurnEnd.TIMED_OUT; + } + if (activity.isPaused()) { + continue; // an approval question is waiting for its answer + } + if (terminal.hasPendingInput()) { + // Something was typed while the agent was working. Stop the turn rather than finish a + // request that has been overtaken: the handle closes the stream the model is answering + // on. The line itself stays queued and becomes the next message, so what the model + // produced so far is kept and the new instruction follows it. + handle.cancel(); + terminal.line(terminal.ansi().yellow("(interrupted — taking your message)")); + terminal.status(List.of(IDLE_LINE, stateLine.get())); + return TurnEnd.INTERRUPTED; + } + terminal.status(List.of( + activityLine( + ACTIVITY_FRAMES.charAt(frame++ % ACTIVITY_FRAMES.length()), + word, + seconds, + session.runningTool(), + session.runningSeconds(), + session.toolCalls()), + stateLine.get())); + } + terminal.status(List.of(IDLE_LINE, stateLine.get())); + return TurnEnd.FINISHED; + } + + /** + * What the pinned line says while a turn is running. + * + *

Naming the running tool is the point: a build or a test run can take minutes, and + * "working…" during a two-minute {@code mvn test} is indistinguishable from a hang. When no tool + * runs, the model is generating, which is its own kind of waiting. + * + * @param frame the spinner character + * @param seconds how long the whole turn has been running + * @param runningTool the tool executing right now, or {@code null} + * @param toolSeconds how long that tool has been running + * @param toolCalls how many tools ran in this turn so far + * @return the line + */ + static String activityLine( + char frame, String word, long seconds, @Nullable String runningTool, long toolSeconds, int toolCalls) { + String inside = runningTool == null ? seconds + "s" : runningTool + " " + toolSeconds + "s of " + seconds + "s"; + return frame + " " + word + "… (" + inside + (toolCalls == 0 ? "" : " · " + toolCalls + " tool calls") + ")"; + } + + /** + * A word for the activity line, drawn once per turn. + * + *

Our own list ({@value #SPINNER_WORDS}), not the one Claude Code ships: that one is extracted + * from a proprietary binary, and the public collections of it are either unlicensed or + * CC BY-NC-SA — neither is compatible with this project's MIT licence or with REUSE. Edit the + * resource to change them; no Java involved. + * + * @return one word, or {@code "Thinking"} when the list cannot be read + */ + static String spinnerWord() { + List words = prompt(SPINNER_WORDS) + .lines() + .map(String::strip) + .filter(word -> !word.isEmpty()) + .toList(); + return words.isEmpty() + ? "Thinking" + : words.get(java.util.concurrent.ThreadLocalRandom.current().nextInt(words.size())); + } + + /** + * Whether the history has to be summarized before the next request is sent. + * + *

The threshold is on the low side ({@link AgentOptions#DEFAULT_COMPACT_AT} %) because the + * number it is compared against is usually an estimate, and because the model's reply has to fit + * next to the prompt. With an unknown context size — a foreign endpoint whose {@code /props} says + * nothing — nothing is decided at all rather than guessed. + * + * @param options the options, for the switch and the threshold + * @param contextSize the context window, or {@link StatusLine#UNKNOWN_CONTEXT} + * @param estimatedTokens what the next request is expected to carry + * @return {@code true} when the history should be summarized first + */ + static boolean needsCompaction(AgentOptions options, int contextSize, long estimatedTokens) { + if (!options.isAutoCompact() || contextSize <= StatusLine.UNKNOWN_CONTEXT) { + return false; + } + return estimatedTokens * 100 >= (long) contextSize * options.getCompactAt(); + } + + /** + * Whether the history is the untouched result of a compaction. + * + * @param history the conversation + * @return {@code true} when it is exactly the summary and its acknowledgement + */ + static boolean isCompacted(List history) { + return history.size() == 2 + && history.get(0).content() != null + && history.get(0).content().startsWith(SUMMARY_PREFIX); + } + + /** + * A rough token count of what the next request will carry. + * + *

Used only for the status line, and only because llama.cpp sends its own count just to clients + * that ask for it ({@code stream_options.include_usage}), which Atmosphere's client does not. Four + * characters per token is the usual rule of thumb; the status line marks the number with a + * {@code ~} so nobody reads it as exact. + * + * @param systemPrompt the system prompt sent with every request + * @param history the conversation so far + * @return the estimated token count + */ + static long estimateTokens(String systemPrompt, List history) { + long characters = systemPrompt.length(); + for (ChatMessage message : history) { + characters += message.content() == null ? 0 : message.content().length(); + } + return characters / CHARS_PER_TOKEN; + } + + /** + * The context figure while a turn is running. + * + *

The count shown before the turn started is not the count during it: every tool round appends + * the call and its output to the conversation the next model call of the same turn is sent, so a + * turn that reads three files and runs a build can add thousands of tokens before the prompt comes + * back. The status row used to be rendered once and handed to the redraw loop as a fixed string, so + * it stood still for the whole turn and only moved at the next {@code you>} — which is precisely + * when it no longer matters. + * + * @param baseTokens what the request carried when it was sent + * @param session the running turn + * @return the server's own count once it reported one, else the base plus what the turn produced + */ + static long liveTokens(long baseTokens, ConsoleSession session) { + long reported = session.inputTokens(); + return reported > 0 ? reported : baseTokens + session.producedChars() / CHARS_PER_TOKEN; + } + + /** + * Summarize the history and continue from the summary. + * + *

The summary is generated by the same model with no tools, then replaces the history as + * a {@code user} message plus a short assistant acknowledgement — aider's shape, and the one that + * survives a strict chat template, because a conversation may not start with two assistant turns. + * The tool rounds of a turn are not in the history to begin with (only the user text and the final + * answer are), so what is condensed here is what the next turn would have replayed anyway. + * + * @param runner the runner + * @param fileSystem the workspace filesystem for the summary session + * @param history the history, replaced in place + * @param focus optional extra instructions from {@code /compact } + * @param terminal the console + * @return the input tokens the summarizing call reported + * @throws InterruptedException if interrupted while the summary is generated + */ + private static long compact( + AgentRunner runner, + AgentFileSystem fileSystem, + List history, + String focus, + AgentTerminal terminal) + throws InterruptedException { + if (history.isEmpty()) { + terminal.line("(nothing to compact)"); + return 0; + } + if (isCompacted(history)) { + // After a compaction the history IS the summary plus its acknowledgement. Summarizing that + // again returns the same text for another model call -- and re-sends a byte-identical + // prompt, which is what makes llama.cpp log "need to evaluate at least 1 token". + terminal.line("(the history is already a summary — nothing to compact)"); + return estimateTokens("", history); + } + String instructions = prompt(COMPACT_PROMPT) + .replace("{focus}", focus.isEmpty() ? "" : System.lineSeparator() + "Focus on: " + focus); + int before = history.size(); + terminal.line("(compacting " + before + " messages …)"); + ConsoleSession session = new ConsoleSession(terminal, fileSystem); + Thread worker = new Thread( + () -> { + try { + runner.runWithoutTools(instructions, List.copyOf(history), session, COMPACT_SYSTEM_PROMPT); + } catch (RuntimeException e) { + session.error(e); + } + }, + "agent-compact"); + worker.setDaemon(true); + worker.start(); + if (awaitWithActivity( + session, + terminal, + () -> "/compact", + new TurnActivity(), + org.atmosphere.ai.ExecutionHandle.completed()) + != TurnEnd.FINISHED + || session.text().isBlank()) { + terminal.line("(compact failed; history kept)"); + return 0; + } + history.clear(); + history.add(ChatMessage.user( + SUMMARY_PREFIX + System.lineSeparator() + session.text().strip())); + history.add(ChatMessage.assistant("Understood, I will continue from that summary.")); + terminal.line("(compacted " + before + " messages into a summary of " + + session.text().strip().length() + " characters)"); + return session.inputTokens(); } /** @@ -221,6 +1089,17 @@ static ModelParameters modelParameters(AgentOptions options) { return parameters; } + /** + * Every command name and alias, for tab completion. + * + * @return the names, each with its leading slash + */ + static List commandNames() { + return java.util.Arrays.stream(SlashCommands.Command.values()) + .flatMap(command -> command.names().stream()) + .toList(); + } + /** * The default system prompt, or the {@code --system} override. * diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LoopOptions.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LoopOptions.java new file mode 100644 index 000000000..e9210ad09 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/LoopOptions.java @@ -0,0 +1,147 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.time.Duration; +import java.util.Locale; +import org.jspecify.annotations.Nullable; + +/** + * What {@code /loop} was asked to do: the task, and the limits that stop it. + * + *

Syntax: {@code /loop [--every ] [--max ] [--check ] }. The flags come + * first, the rest of the line is the task verbatim — so a task may contain anything, including words + * that look like flags, as long as they are not at the front. + * + * @param task what the agent should work on, restated unchanged every step + * @param interval the pause between steps, or {@code null} to run them back to back + * @param maxSteps the hard cap on steps + * @param check a command that must succeed before {@code <>} is accepted, or + * {@code null} to trust the model + */ +public record LoopOptions( + String task, + @Nullable Duration interval, + int maxSteps, + @Nullable String check) { + + /** Steps before the loop gives up on its own. */ + public static final int DEFAULT_MAX_STEPS = 20; + + /** + * Parse the argument of {@code /loop}. + * + * @param arguments everything after the command name + * @return the parsed options + * @throws IllegalArgumentException when a flag has no value, a duration or number is malformed, or + * no task is left + */ + public static LoopOptions parse(String arguments) { + String rest = arguments == null ? "" : arguments.trim(); + Duration interval = null; + int maxSteps = DEFAULT_MAX_STEPS; + String check = null; + while (rest.startsWith("--")) { + String[] flag = split(rest); + switch (flag[0]) { + case "--every" -> { + String[] value = split(flag[1]); + interval = parseDuration(require(value[0], "--every")); + rest = value[1]; + } + case "--max" -> { + String[] value = split(flag[1]); + maxSteps = parseSteps(require(value[0], "--max")); + rest = value[1]; + } + case "--check" -> { + String[] value = splitQuoted(flag[1]); + check = require(value[0], "--check"); + rest = value[1]; + } + default -> throw new IllegalArgumentException("Unknown /loop flag: " + flag[0]); + } + } + if (rest.isEmpty()) { + throw new IllegalArgumentException("Usage: /loop [--every 5m] [--max 20] [--check ''] "); + } + return new LoopOptions(rest, interval, maxSteps, check); + } + + /** + * A duration written the way a human writes it: {@code 30s}, {@code 5m}, {@code 2h}, or a bare + * number of minutes. + * + * @param text the duration + * @return the parsed duration + * @throws IllegalArgumentException when it is not a duration or is not positive + */ + static Duration parseDuration(String text) { + String value = text.trim().toLowerCase(Locale.ROOT); + char unit = value.charAt(value.length() - 1); + String number = Character.isDigit(unit) ? value : value.substring(0, value.length() - 1); + long amount; + try { + amount = Long.parseLong(number); + } catch (NumberFormatException e) { + throw new IllegalArgumentException("Not a duration: " + text + " (try 30s, 5m or 2h)", e); + } + Duration duration = + switch (Character.isDigit(unit) ? 'm' : unit) { + case 's' -> Duration.ofSeconds(amount); + case 'm' -> Duration.ofMinutes(amount); + case 'h' -> Duration.ofHours(amount); + default -> throw new IllegalArgumentException("Unknown time unit in " + text + " (use s, m or h)"); + }; + if (duration.isZero() || duration.isNegative()) { + throw new IllegalArgumentException("The interval must be positive: " + text); + } + return duration; + } + + private static int parseSteps(String text) { + try { + int steps = Integer.parseInt(text.trim()); + if (steps <= 0) { + throw new IllegalArgumentException("--max must be positive: " + text); + } + return steps; + } catch (NumberFormatException e) { + throw new IllegalArgumentException("--max expects a number, got: " + text, e); + } + } + + private static String require(String value, String flag) { + if (value.isEmpty()) { + throw new IllegalArgumentException("Missing value for " + flag); + } + return value; + } + + /** Split off the first whitespace-separated word: {@code [word, rest]}. */ + private static String[] split(String text) { + int space = text.indexOf(' '); + return space < 0 + ? new String[] {text.trim(), ""} + : new String[] { + text.substring(0, space).trim(), text.substring(space + 1).trim() + }; + } + + /** Like {@link #split}, but a leading {@code '...'} or {@code "..."} keeps its spaces. */ + private static String[] splitQuoted(String text) { + String trimmed = text.trim(); + if (trimmed.startsWith("'") || trimmed.startsWith("\"")) { + char quote = trimmed.charAt(0); + int end = trimmed.indexOf(quote, 1); + if (end > 0) { + return new String[] { + trimmed.substring(1, end), trimmed.substring(end + 1).trim() + }; + } + } + return split(trimmed); + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/MarkdownConsole.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/MarkdownConsole.java new file mode 100644 index 000000000..6e15f06fc --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/MarkdownConsole.java @@ -0,0 +1,125 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.util.function.Consumer; + +/** + * Renders the streamed answer as it arrives: just enough Markdown to make it readable. + * + *

Append-only, one line at a time. Tokens are buffered until a line is complete, then that + * line is written and never touched again. Redrawing the answer on every token — what the Ink and + * Bubble Tea based clients do — is what produces their overdraw and truncation bugs, and it needs + * cursor control that breaks as soon as the output is piped into a file. The price here is that a + * line appears only once it ends. + * + *

Handled: fenced code blocks (one bit of state), ATX headings, bullet markers, and inline + * {@code **bold**} / {@code `code`}. Italics are deliberately not handled — a lone {@code *} is more + * often a glob or a multiplication than emphasis, and getting that wrong garbles ordinary text. + * Everything else is passed through unchanged, which is what a terminal wants anyway. + * + *

Only the console sees this; {@link ConsoleSession} keeps the raw text for the history, so + * nothing that goes back to the model is affected. + */ +public final class MarkdownConsole { + + private final Consumer sink; + private final Ansi ansi; + private final StringBuilder pending = new StringBuilder(); + private boolean inFence; + + /** + * Create a renderer. + * + * @param sink receives each rendered line, without its terminator + * @param ansi the styles (a plain instance writes the text unchanged) + */ + public MarkdownConsole(Consumer sink, Ansi ansi) { + this.sink = sink; + this.ansi = ansi; + } + + /** + * Take the next streamed chunk; every complete line in it is rendered and written. + * + * @param chunk the chunk as it arrived, possibly a fragment of a line + */ + public void append(String chunk) { + pending.append(chunk); + int newline; + while ((newline = pending.indexOf("\n")) >= 0) { + String line = pending.substring(0, newline); + pending.delete(0, newline + 1); + sink.accept(render(line.endsWith("\r") ? line.substring(0, line.length() - 1) : line)); + } + } + + /** Write what is left of an unfinished line, e.g. an answer that does not end with a newline. */ + public void flush() { + if (pending.length() > 0) { + sink.accept(render(pending.toString())); + pending.setLength(0); + } + inFence = false; + } + + /** + * Render one complete line. + * + * @param line the line without its terminator + * @return the line with escape sequences, or unchanged when styling is off + */ + String render(String line) { + String content = line.stripLeading(); + String indent = line.substring(0, line.length() - content.length()); + if (content.startsWith("```") || content.startsWith("~~~")) { + inFence = !inFence; + return ansi.dim(line); + } + if (inFence) { + return ansi.cyan(line); + } + int heading = 0; + while (heading < content.length() && content.charAt(heading) == '#') { + heading++; + } + if (heading > 0 && heading <= 6 && content.startsWith("# ", heading - 1)) { + return indent + ansi.bold(inline(content.substring(heading + 1).strip())); + } + if (content.length() > 2 && "-*+".indexOf(content.charAt(0)) >= 0 && content.charAt(1) == ' ') { + return indent + ansi.cyan("•") + " " + inline(content.substring(2)); + } + return indent + inline(content); + } + + /** + * Style {@code **bold**} and {@code `code`} inside one line. + * + * @param text the line content + * @return the styled content + */ + String inline(String text) { + StringBuilder result = new StringBuilder(text.length()); + int index = 0; + while (index < text.length()) { + int code = text.indexOf('`', index); + int bold = text.indexOf("**", index); + boolean codeFirst = code >= 0 && (bold < 0 || code < bold); + int start = codeFirst ? code : bold; + if (start < 0) { + break; + } + String marker = codeFirst ? "`" : "**"; + int end = text.indexOf(marker, start + marker.length()); + if (end < 0) { + break; // an unclosed marker: leave the rest as typed + } + String span = text.substring(start + marker.length(), end); + result.append(text, index, start).append(codeFirst ? ansi.cyan(span) : ansi.bold(span)); + index = end + marker.length(); + } + return result.append(text.substring(index)).toString(); + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java new file mode 100644 index 000000000..1312f39ef --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/PlainTerminal.java @@ -0,0 +1,89 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.io.BufferedReader; +import java.io.IOException; +import java.io.PrintStream; +import java.util.Locale; +import org.jspecify.annotations.Nullable; + +/** + * An {@link AgentTerminal} over a plain stream pair: what a one-shot run, a piped session and the + * tests get. + * + *

No cursor control at all — the status line is printed as an ordinary line before the prompt and + * scrolls away like everything else, and an answer needs Enter. That is the point: this + * implementation stays correct when the output is a file. + */ +public final class PlainTerminal implements AgentTerminal { + + private final PrintStream out; + private final @Nullable BufferedReader in; + private final Ansi ansi; + + /** + * Create a plain console. + * + * @param out where output goes + * @param in where input is read, or {@code null} when there is none (one-shot runs) + * @param ansi the styles + */ + public PlainTerminal(PrintStream out, @Nullable BufferedReader in, Ansi ansi) { + this.out = out; + this.in = in; + this.ansi = ansi; + } + + @Override + public void line(String text) { + out.println(text); + out.flush(); + } + + @Override + public @Nullable String readLine(String prompt) { + out.print(prompt); + out.flush(); + return read(); + } + + @Override + public @Nullable String readKey(String prompt) { + out.print(prompt); + out.flush(); + String answer = read(); + return answer == null ? null : answer.trim().toLowerCase(Locale.ROOT); + } + + @Override + public void status(java.util.List lines) { + // Nothing can be pinned on a plain stream, and the caller already prints the status line above + // the prompt. Dropping it here is what keeps a piped session free of half-drawn spinner lines. + } + + @Override + public boolean pinsStatus() { + return false; + } + + @Override + public Ansi ansi() { + return ansi; + } + + @Override + public void close() { + out.flush(); + } + + private @Nullable String read() { + try { + return in == null ? null : in.readLine(); + } catch (IOException e) { + return null; + } + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ServerProps.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ServerProps.java new file mode 100644 index 000000000..32e21bf5b --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ServerProps.java @@ -0,0 +1,101 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.net.URI; +import java.net.http.HttpClient; +import java.net.http.HttpRequest; +import java.net.http.HttpResponse; +import java.time.Duration; +import java.util.List; +import java.util.regex.Matcher; +import java.util.regex.Pattern; + +/** + * Reads the context size of a running server from llama.cpp's {@code GET /props}, so the status line + * can show a percentage in {@code --base-url} mode too. + * + *

Both java-llama.cpp's {@code OpenAiCompatServer} and upstream {@code llama-server} answer with + * {@code {"default_generation_settings":{"n_ctx":…}}}. The endpoint sits next to the OpenAI routes + * rather than under {@code /v1}, and this project's server serves it under both, so both are tried. + * Every failure — a server without the route, a foreign OpenAI-compatible endpoint, a timeout — is + * reported as "unknown"; the status line then shows the token count without a percentage rather than + * a made-up denominator. + */ +public final class ServerProps { + + private static final Pattern N_CTX = Pattern.compile("\"n_ctx\"\\s*:\\s*(\\d+)"); + private static final Duration TIMEOUT = Duration.ofSeconds(3); + + private ServerProps() {} + + /** + * Look the context size up. + * + * @param baseUrl the OpenAI-compatible base URL, e.g. {@code http://127.0.0.1:8080/v1} + * @param apiKey the bearer token to send + * @return the context size in tokens, or {@link StatusLine#UNKNOWN_CONTEXT} when it cannot be read + */ + public static int contextSize(String baseUrl, String apiKey) { + HttpClient client = HttpClient.newBuilder().connectTimeout(TIMEOUT).build(); + for (String url : candidates(baseUrl)) { + int size = read(client, url, apiKey); + if (size != StatusLine.UNKNOWN_CONTEXT) { + return size; + } + } + return StatusLine.UNKNOWN_CONTEXT; + } + + /** + * The {@code /props} URLs tried, in order. + * + * @param baseUrl the base URL + * @return the candidate URLs + */ + static List candidates(String baseUrl) { + String trimmed = baseUrl.endsWith("/") ? baseUrl.substring(0, baseUrl.length() - 1) : baseUrl; + if (trimmed.endsWith("/v1")) { + String root = trimmed.substring(0, trimmed.length() - "/v1".length()); + return List.of(root + "/props", trimmed + "/props"); + } + return List.of(trimmed + "/props"); + } + + /** + * Extract {@code n_ctx} from a {@code /props} body. + * + * @param body the response body + * @return the context size, or {@link StatusLine#UNKNOWN_CONTEXT} when the field is absent + */ + static int parseContextSize(String body) { + Matcher matcher = N_CTX.matcher(body); + if (!matcher.find()) { + return StatusLine.UNKNOWN_CONTEXT; + } + try { + return Integer.parseInt(matcher.group(1)); + } catch (NumberFormatException e) { + return StatusLine.UNKNOWN_CONTEXT; + } + } + + private static int read(HttpClient client, String url, String apiKey) { + try { + HttpRequest request = HttpRequest.newBuilder(URI.create(url)) + .timeout(TIMEOUT) + .header("Authorization", "Bearer " + apiKey) + .GET() + .build(); + HttpResponse response = client.send(request, HttpResponse.BodyHandlers.ofString()); + return response.statusCode() == 200 ? parseContextSize(response.body()) : StatusLine.UNKNOWN_CONTEXT; + } catch (InterruptedException e) { + Thread.currentThread().interrupt(); + return StatusLine.UNKNOWN_CONTEXT; + } catch (RuntimeException | java.io.IOException e) { + return StatusLine.UNKNOWN_CONTEXT; + } + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ShellTool.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ShellTool.java index 7e2604d27..a3493ac4d 100644 --- a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ShellTool.java +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ShellTool.java @@ -4,7 +4,9 @@ package net.ladenthin.llama.atmosphere; +import java.io.BufferedReader; import java.io.IOException; +import java.io.InputStreamReader; import java.nio.charset.StandardCharsets; import java.nio.file.Path; import java.time.Duration; @@ -12,6 +14,7 @@ import java.util.concurrent.ExecutionException; import java.util.concurrent.TimeUnit; import java.util.concurrent.TimeoutException; +import java.util.function.Consumer; import org.atmosphere.ai.tool.ToolDefinition; /** @@ -41,6 +44,20 @@ private ShellTool() {} * @return the definition */ public static ToolDefinition definition(Path workspace, Duration defaultTimeout, int maxOutputChars) { + return definition(workspace, defaultTimeout, maxOutputChars, line -> {}); + } + + /** + * Build the tool definition with live output. + * + * @param workspace the working directory of every command + * @param defaultTimeout the timeout applied when the model does not pass {@code timeout_seconds} + * @param maxOutputChars output is truncated to this many characters (tail kept, head marked) + * @param liveOutput receives each output line while the command is still running + * @return the definition + */ + public static ToolDefinition definition( + Path workspace, Duration defaultTimeout, int maxOutputChars, Consumer liveOutput) { return ToolDefinition.builder( TOOL_NAME, LocalAgent.prompt(DESCRIPTION_RESOURCE).replace("{shell}", shellName())) .parameter(PARAM_COMMAND, "The command line to run through " + shellName(), "string", true) @@ -51,7 +68,7 @@ public static ToolDefinition definition(Path workspace, Duration defaultTimeout, return "Error: '" + PARAM_COMMAND + "' is required"; } Duration timeout = timeoutOf(args.get(PARAM_TIMEOUT), defaultTimeout); - return run(workspace, command.toString(), timeout, maxOutputChars); + return run(workspace, command.toString(), timeout, maxOutputChars, liveOutput); }) .build(); } @@ -106,18 +123,47 @@ private static Duration timeoutOf(Object raw, Duration fallback) { */ static String run(Path workspace, String command, Duration timeout, int maxOutputChars) throws IOException, InterruptedException { + return run(workspace, command, timeout, maxOutputChars, line -> {}); + } + + /** + * Run one command through the platform shell, reporting its output as it arrives. + * + *

The output is read line by line rather than in one go at the end. That is what lets the + * console show a long build while it runs — a silent minute is indistinguishable from a hang — and + * it is also what keeps the pipe drained: a process whose output nobody reads blocks once the + * pipe buffer is full, which on Windows is roughly 4 KB. + * + * @param workspace the working directory + * @param command the command line + * @param timeout kill the process after this long + * @param maxOutputChars truncate the captured output to this many characters + * @param liveOutput receives each line as it is read + * @return a text block starting with {@code exit code: N}, followed by the output + * @throws IOException if the process cannot be started + * @throws InterruptedException if interrupted while waiting + */ + static String run(Path workspace, String command, Duration timeout, int maxOutputChars, Consumer liveOutput) + throws IOException, InterruptedException { ProcessBuilder builder = isWindows() ? new ProcessBuilder("cmd.exe", "/c", command) : new ProcessBuilder("sh", "-c", command); builder.directory(workspace.toFile()); builder.redirectErrorStream(true); Process process = builder.start(); process.getOutputStream().close(); - CompletableFuture output = CompletableFuture.supplyAsync(() -> { - try { - return process.getInputStream().readAllBytes(); + CompletableFuture output = CompletableFuture.supplyAsync(() -> { + StringBuilder collected = new StringBuilder(); + try (BufferedReader reader = + new BufferedReader(new InputStreamReader(process.getInputStream(), StandardCharsets.UTF_8))) { + String line; + while ((line = reader.readLine()) != null) { + collected.append(line).append(System.lineSeparator()); + liveOutput.accept(line); + } } catch (IOException e) { - return ("[output unreadable: " + e.getMessage() + "]").getBytes(StandardCharsets.UTF_8); + collected.append("[output unreadable: ").append(e.getMessage()).append("]"); } + return collected.toString(); }); boolean finished = process.waitFor(timeout.toMillis(), TimeUnit.MILLISECONDS); if (!finished) { @@ -129,7 +175,7 @@ static String run(Path workspace, String command, Duration timeout, int maxOutpu } String text; try { - text = new String(output.get(5, TimeUnit.SECONDS), StandardCharsets.UTF_8); + text = output.get(5, TimeUnit.SECONDS); } catch (ExecutionException | TimeoutException e) { text = "[output unavailable: " + e.getMessage() + "]"; } diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java new file mode 100644 index 000000000..bc8e2949c --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/SlashCommands.java @@ -0,0 +1,121 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.util.List; +import java.util.Locale; +import java.util.Optional; +import org.jspecify.annotations.Nullable; + +/** + * The REPL's client-side commands: a line the agent answers itself instead of sending it to the model. + * + *

Dispatch is deliberately conservative. A line is a command only when it starts with {@code /} + * and its first word names a known command; everything else — including an unknown + * {@code /foo} — goes to the model verbatim. That is what makes {@code /usr/bin/env} or a line of + * Markdown work with no escape syntax, at the price of a typo being answered by the model rather than + * rejected. (aider rejects unknown commands outright and needs no escape because its REPL is + * line-oriented; this agent is asked prose far more often than it is asked commands.) + * + * @param command the command + * @param arguments the rest of the line, trimmed; empty when the line was just the command + */ +public record SlashCommands(Command command, String arguments) { + + /** The commands the REPL answers itself. */ + public enum Command { + /** Print the command overview. */ + HELP("/help", "/?", "/commands"), + /** Drop the conversation history, and wipe the screen with it. */ + CLEAR("/clear", "/reset", "/new"), + /** Wipe the screen, keeping the conversation. */ + CLS("/cls", "/clear-screen"), + /** Write what was said, with the time, to a file in the workspace. */ + SAVE("/save", "/transcript"), + /** Ask the last question again, without the answer that came back. */ + RETRY("/retry", "/again"), + /** Read a saved transcript back in as the conversation. */ + LOAD("/load", "/resume"), + /** Summarize the history and continue with the summary; the argument steers the summary. */ + COMPACT("/compact"), + /** Keep working on one task until it is done; see {@link TaskLoop}. */ + LOOP("/loop"), + /** Show or set the approval mode; the argument is {@code manual} or {@code auto}. */ + MODE("/mode", "/approve"), + /** Print endpoint, model, tools, approval mode and context usage. */ + STATUS("/status"), + /** List the tools offered to the model. */ + TOOLS("/tools"), + /** Show every tool call of this session — the receipt for what really happened. */ + CALLS("/calls", "/log"), + /** Leave the REPL. */ + EXIT("/exit", "/quit"); + + private final List names; + + Command(String... names) { + this.names = List.of(names); + } + + /** + * The names that select this command, the canonical one first. + * + * @return the names, each including the leading slash + */ + public List names() { + return names; + } + + /** + * The canonical name. + * + * @return e.g. {@code "/help"} + */ + public String canonicalName() { + return names.get(0); + } + } + + /** + * Parse one REPL line. + * + * @param line the raw line as typed + * @return the command and its arguments, or empty when the line is a message for the model + */ + public static Optional parse(@Nullable String line) { + // A BOM at the start of piped input would otherwise hide the slash and send /help to the model. + String trimmed = line == null ? "" : line.replace("", "").trim(); + if (!trimmed.startsWith("/")) { + return Optional.empty(); + } + int space = indexOfWhitespace(trimmed); + String name = (space < 0 ? trimmed : trimmed.substring(0, space)).toLowerCase(Locale.ROOT); + String arguments = space < 0 ? "" : trimmed.substring(space + 1).trim(); + for (Command command : Command.values()) { + if (command.names().contains(name)) { + return Optional.of(new SlashCommands(command, arguments)); + } + } + return Optional.empty(); + } + + private static int indexOfWhitespace(String text) { + for (int i = 0; i < text.length(); i++) { + if (Character.isWhitespace(text.charAt(i))) { + return i; + } + } + return -1; + } + + /** + * Whether an argument was given. + * + * @return {@code true} when the line carried more than the command name + */ + public boolean hasArguments() { + return !arguments.isEmpty(); + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java new file mode 100644 index 000000000..0f18ba717 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/StatusLine.java @@ -0,0 +1,134 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.util.Locale; + +/** + * The one-line status pinned below the answer: workspace, approval mode, context usage, tool count + * and model id — the four things whose answer changes what the next request does. + * + *

Context usage is the input side of the last completed turn — the prompt the server had to + * process, which is what fills the context window; the generated tokens of that turn are already part + * of the next request's input. Claude Code's status line computes its percentage the same way. + * The count comes from the server when it reports usage. llama.cpp only sends the trailing usage + * chunk when the client asks for it ({@code stream_options.include_usage}) and Atmosphere's client + * does not, so in practice the number is an estimate from the text length (four characters per + * token, the usual rule of thumb) and is then marked with a {@code ~}. It is meant to answer "am I + * close to the limit, should I /compact", not to be exact. + * + *

The size is known with {@code --model} (it is the + * {@code --ctx-size} the agent loaded the model with) and with {@code --base-url} it comes from the + * server's {@code /props}; when that lookup fails the line shows the token count alone instead of + * inventing a denominator. + */ +public final class StatusLine { + + /** In front of the workspace path. */ + private static final String WORKSPACE_ICON = "📁"; + + /** In front of the context figure. */ + private static final String CONTEXT_ICON = "📊"; + + /** In front of the tool count. */ + private static final String TOOLS_ICON = "🔧"; + + /** In front of a model this process loaded itself. */ + private static final String LOCAL_MODEL_ICON = "🤖"; + + /** In front of a model served by something else over the network. */ + private static final String REMOTE_MODEL_ICON = "🌐"; + + /** Above this many characters the workspace path is shortened to its last two segments. */ + private static final int MAX_PATH_CHARS = 40; + + /** The context size is unknown (no {@code /props}, and no in-process model). */ + public static final int UNKNOWN_CONTEXT = 0; + + private StatusLine() {} + + /** + * Render the status line. + * + * @param workspace the directory the tools work in + * @param mode the approval mode + * @param inputTokens the input tokens of the last turn, or {@code 0} before the first one + * @param estimated whether that number is an estimate rather than the server's own count + * @param contextSize the context window in tokens, or {@link #UNKNOWN_CONTEXT} + * @param tools how many tools are offered to the model + * @param modelId the model id sent in every request + * @param remote whether the model is served by another process rather than loaded here + * @return one line, without a trailing newline + */ + public static String render( + java.nio.file.Path workspace, + ApprovalMode mode, + long inputTokens, + boolean estimated, + int contextSize, + int tools, + String modelId, + boolean remote) { + // Every part is an icon, a space, and its value. The icons carry what the words used to, so + // the line stays short enough to survive a narrow window; the space is what keeps an icon from + // running into its value, which is easy to lose when the glyph is wide. + return "[" + WORKSPACE_ICON + " " + shorten(workspace) + + " · " + mode.badge() + + " · " + CONTEXT_ICON + " " + context(inputTokens, estimated, contextSize) + + " · " + TOOLS_ICON + " " + tools + + " · " + (remote ? REMOTE_MODEL_ICON : LOCAL_MODEL_ICON) + " " + modelId + "]"; + } + + /** + * The workspace path, shortened from the left when it would take over the line. + * + *

The tools work relative to this directory and the shell starts in it, so it belongs on the + * line that is always visible — but a deep path would push everything else off the screen, so only + * the last two segments survive, marked with a leading ellipsis. + * + * @param workspace the workspace directory + * @return the path, or its tail + */ + static String shorten(java.nio.file.Path workspace) { + String full = workspace.toString(); + if (full.length() <= MAX_PATH_CHARS || workspace.getNameCount() < 2) { + return full; + } + return "…" + workspace.getFileSystem().getSeparator() + + workspace.subpath(workspace.getNameCount() - 2, workspace.getNameCount()); + } + + /** + * The context part of the line on its own. + * + * @param inputTokens the input tokens of the last turn + * @param estimated whether that number is an estimate (rendered with a leading {@code ~}) + * @param contextSize the context window in tokens, or {@link #UNKNOWN_CONTEXT} + * @return e.g. {@code "1.2k/16k"}, or {@code "1.2k"} when the size is unknown + */ + static String context(long inputTokens, boolean estimated, int contextSize) { + // Two numbers in k, no percentage: everyone reads 12k/16k at a glance, and a percentage of a + // number that is itself an estimate suggests a precision this does not have. + String used = (estimated ? "~" : "") + abbreviate(inputTokens); + return contextSize <= UNKNOWN_CONTEXT ? used : used + "/" + abbreviate(contextSize); + } + + /** + * A short token count: {@code 812}, {@code 1.2k}, {@code 16k}. + * + * @param tokens the count + * @return the abbreviated form + */ + static String abbreviate(long tokens) { + if (tokens < 1000) { + return Long.toString(tokens); + } + double thousands = tokens / 1000.0; + // one decimal below 10k (1.2k), none above (16k) -- the decimal carries no information there + return thousands < 10 + ? String.format(Locale.ROOT, "%.1fk", thousands) + : String.format(Locale.ROOT, "%.0fk", thousands); + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java new file mode 100644 index 000000000..099a32a02 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TaskLoop.java @@ -0,0 +1,262 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.io.IOException; +import java.io.UncheckedIOException; +import java.nio.charset.StandardCharsets; +import java.nio.file.Files; +import java.nio.file.Path; +import java.time.Duration; +import java.util.List; +import java.util.regex.Pattern; +import org.jspecify.annotations.Nullable; + +/** + * {@code /loop}: keep working on one task, step by step, until it is done — the state lives in a file, + * not in the conversation. + * + *

Every step sends the same message: the task verbatim plus an instruction to read + * {@value #LOOP_FILE}, do one concrete step, and write down what happened. The conversation history is + * dropped between steps, which is the point of the file: the context never grows, so the loop + * can run for hours, and what the model knows is exactly what it wrote down. (This is the shape + * Claude Code's own ralph-wiggum plugin uses — re-inject the original prompt and let the filesystem be + * the memory.) + * + *

Why a text marker and not a "done" tool. A small local model produces a well-formed tool + * call far less reliably than a line of text — below 7B, malformed calls are the norm, and + * mini-SWE-agent reaches its SWE-bench results with a plain-text sentinel and no tool-call API at all. + * So the loop ends when a line is exactly {@value #SENTINEL}; a substring never counts, which + * is what keeps "I will answer <<TASK_COMPLETE>> when I am done" from ending the run. + * + *

The model saying "done" is not proof. With {@code --check } the sentinel is only + * accepted when that command succeeds; otherwise its output goes back into the next step. A 4B model + * declares victory early, and this is the cheapest defence against it. + * + *

Limits, all of them enforced here rather than trusted to the model: a step cap + * ({@link LoopOptions#DEFAULT_MAX_STEPS} by default), a wall-clock budget, and a stall detector — if + * neither the file nor any tool call changed anything for {@value #STALL_LIMIT} steps in a row, the + * loop stops instead of burning tokens on a model that repeats itself. + */ +public final class TaskLoop { + + /** The file the loop keeps its plan and its notes in, inside the workspace. */ + public static final String LOOP_FILE = "AGENT-LOOP.md"; + + /** The line that ends the loop. Matched as a whole line, never as a substring. */ + public static final String SENTINEL = "<>"; + + /** Steps without any change before the loop gives up. */ + public static final int STALL_LIMIT = 3; + + /** How long a loop may run before it stops on its own. */ + public static final Duration DEFAULT_BUDGET = Duration.ofHours(2); + + private static final Pattern SENTINEL_LINE = Pattern.compile("^\\s*" + Pattern.quote(SENTINEL) + "\\s*$"); + + private TaskLoop() {} + + /** + * Whether an answer ends the loop. + * + * @param answer the model's complete answer for one step + * @return {@code true} when one of its lines is exactly the sentinel + */ + public static boolean isComplete(String answer) { + return answer.lines().anyMatch(line -> SENTINEL_LINE.matcher(line).matches()); + } + + /** + * The message sent for every step. + * + * @param options the loop options + * @return the prompt, with the task, the file name and the check command filled in + */ + public static String stepPrompt(LoopOptions options) { + String checkHint = options.check() == null + ? "" + : System.lineSeparator() + "Before you declare the task done, run this and make sure it succeeds: " + + options.check(); + return LocalAgent.prompt(LocalAgent.LOOP_PROMPT) + .replace("{task}", options.task()) + .replace("{file}", LOOP_FILE) + .replace("{check_hint}", checkHint); + } + + /** + * Create the loop file when it does not exist yet. + * + * @param workspace the workspace directory + * @param task the task, written into the file + * @return the path of the loop file + */ + public static Path ensureLoopFile(Path workspace, String task) { + Path file = workspace.resolve(LOOP_FILE); + try { + if (!Files.exists(file)) { + Files.writeString( + file, + LocalAgent.prompt(LocalAgent.LOOP_FILE_TEMPLATE).replace("{task}", task) + + System.lineSeparator(), + StandardCharsets.UTF_8); + } + return file; + } catch (IOException e) { + throw new UncheckedIOException("Cannot create " + file, e); + } + } + + /** + * A fingerprint of the loop file, to tell a step that changed something from one that did not. + * + * @param file the loop file + * @return size and content hash, or {@code -1} when the file cannot be read + */ + public static long fingerprint(Path file) { + try { + return Files.exists(file) + ? Files.readString(file, StandardCharsets.UTF_8).hashCode() + : -1; + } catch (IOException e) { + return -1; + } + } + + /** + * Why a loop ended. + * + * @param reason the wording shown to the user + * @param completed whether the model declared the task finished + */ + public record Outcome(String reason, boolean completed) {} + + /** + * Run the loop until it is done, stopped, or out of budget. + * + * @param runner the runner + * @param fileSystem the workspace filesystem for the sessions + * @param terminal the console + * @param workspace the workspace directory + * @param options what to work on and for how long + * @param stopped polled between steps; {@code true} ends the loop (Ctrl-C) + * @param budget the wall-clock limit + * @param callLog records every tool call of every step, so /calls shows what the loop did + * @return why it ended + * @throws InterruptedException if interrupted while waiting for a step or an interval + */ + public static Outcome run( + AgentRunner runner, + org.atmosphere.ai.fs.AgentFileSystem fileSystem, + AgentTerminal terminal, + Path workspace, + LoopOptions options, + java.util.function.BooleanSupplier stopped, + Duration budget, + ToolCallLog callLog) + throws InterruptedException { + Path file = ensureLoopFile(workspace, options.task()); + terminal.line("loop: " + options.task()); + terminal.line("loop: state in " + file + ", max " + options.maxSteps() + " steps, budget " + + budget.toMinutes() + " min" + + (options.interval() == null + ? "" + : ", every " + options.interval().toSeconds() + "s") + + (options.check() == null ? "" : ", check: " + options.check())); + + long deadline = System.nanoTime() + budget.toNanos(); + long lastFingerprint = fingerprint(file); + int stalled = 0; + String extra = ""; + + for (int step = 1; step <= options.maxSteps(); step++) { + if (stopped.getAsBoolean()) { + return new Outcome("stopped after " + (step - 1) + " steps", false); + } + if (System.nanoTime() > deadline) { + return new Outcome( + "budget of " + budget.toMinutes() + " min used up after " + (step - 1) + " steps", false); + } + terminal.line(terminal.ansi().dim("── loop step " + step + "/" + options.maxSteps() + " ──")); + // the status row is rendered on every redraw, and a lambda may not close over the counter + String stepLabel = "loop step " + step + "/" + options.maxSteps() + " · " + options.task(); + + // A fresh history every step: the file is the memory, so the context cannot grow. + ConsoleSession session = LocalAgent.turn( + runner, + fileSystem, + stepPrompt(options) + extra, + new java.util.ArrayList<>(), + terminal, + callLog, + step, + ignored -> stepLabel, + new TurnActivity()); + extra = ""; + if (session.failure() != null) { + return new Outcome("step " + step + " failed: " + session.failure(), false); + } + + // The marker is checked BEFORE the stall detector: a step that only answers "done" changes + // no file and calls no tool, so the other order would report "no progress" on the very + // step that finished the task. + if (isComplete(session.text())) { + String failure = checkFailure(options, workspace); + if (failure == null) { + return new Outcome("done after " + step + " steps", true); + } + terminal.line(terminal.ansi().yellow("loop: the check failed, continuing")); + extra = System.lineSeparator() + "You answered " + SENTINEL + ", but the check (" + + options.check() + ") failed:" + System.lineSeparator() + failure + + System.lineSeparator() + "Fix that first."; + } + + long fingerprint = fingerprint(file); + boolean changed = fingerprint != lastFingerprint || session.toolCalls() > 0; + lastFingerprint = fingerprint; + stalled = changed ? 0 : stalled + 1; + if (stalled >= STALL_LIMIT) { + return new Outcome("no progress for " + STALL_LIMIT + " steps (nothing written, no tools used)", false); + } + + if (options.interval() != null && step < options.maxSteps()) { + Thread.sleep(options.interval().toMillis()); + } + } + return new Outcome("step limit of " + options.maxSteps() + " reached", false); + } + + /** + * Run the check command, if there is one. + * + * @param options the loop options + * @param workspace the directory the command runs in + * @return the command output when it failed, or {@code null} when it succeeded or there is none + */ + private static @Nullable String checkFailure(LoopOptions options, Path workspace) { + String command = options.check(); + if (command == null) { + return null; + } + try { + String result = ShellTool.run(workspace, command, Duration.ofMinutes(10), 4000); + return result.startsWith("exit code: 0") ? null : result; + } catch (IOException e) { + return "the check could not be started: " + e.getMessage(); + } catch (InterruptedException e) { + Thread.currentThread().interrupt(); + return "the check was interrupted"; + } + } + + /** + * The tool names a loop needs; used only for the console hint. + * + * @param toolNames the offered tools + * @return {@code true} when the file tools that the loop file needs are present + */ + public static boolean canKeepNotes(List toolNames) { + return toolNames.contains("read_file") && toolNames.contains("write_file"); + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TextEdits.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TextEdits.java new file mode 100644 index 000000000..b7888630a --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TextEdits.java @@ -0,0 +1,304 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.util.ArrayList; +import java.util.List; +import org.jspecify.annotations.Nullable; + +/** + * The text side of editing a file: what to match, what to write back, and what to say when it does + * not match. No file access, no tools — so every rule here is testable on plain strings. + * + *

Line endings are the reason this class exists. A model answers in LF, always; a file on + * Windows is usually CRLF. Comparing the two directly never matches, which is why the naive + * implementation silently fails on every Windows file. So the file is normalized to LF before + * matching and written back in its own ending, and a byte-order mark is taken off first and + * put back afterwards — it is invisible, and it sits exactly where the first match would be. + * + *

A failed edit is not a free retry. Measured on SWE-agent trajectories: any edit attempt + * eventually succeeds in 90.5 % of cases, but only 57.2 % once a single edit has failed — a failure + * derails the rest of the run. That is why a miss does not answer "not found" but shows the closest + * lines in the file, and an ambiguous match names the line numbers instead of asking for "more + * context". aider's threshold is used for the closest-line search: similarity ≥ {@value #CANDIDATE_MIN_SIMILARITY}. + * + *

Several edits are all-or-nothing, which deviates from every shipping agent — they apply + * sequentially and leave a half-edited file behind. Here the whole batch works on one in-memory + * string and is written only if every edit matched, because at that point atomicity costs nothing and + * a half-applied batch is the state a model reasons about worst. + */ +public final class TextEdits { + + /** How similar a line must be to be worth showing after a failed match (aider's threshold). */ + public static final double CANDIDATE_MIN_SIMILARITY = 0.6; + + /** How many near misses are shown. */ + public static final int MAX_CANDIDATES = 3; + + /** The byte-order mark, as the character it decodes to. */ + private static final char BOM = ''; + + private TextEdits() {} + + /** + * One replacement. + * + * @param oldString the text to find, in LF form + * @param newString what replaces it + * @param replaceAll whether every occurrence is replaced instead of requiring exactly one + */ + public record Edit(String oldString, String newString, boolean replaceAll) {} + + /** Thrown when an edit cannot be applied; the message is what the model gets to read. */ + public static final class EditException extends RuntimeException { + private static final long serialVersionUID = 1L; + + /** + * Create the exception. + * + * @param message the explanation shown to the model + */ + public EditException(String message) { + super(message); + } + } + + /** + * The line ending a text uses. + * + * @param text the file content as read + * @return {@code "\r\n"} when the first line ending is a carriage-return pair, else {@code "\n"} + */ + public static String lineEnding(String text) { + int newline = text.indexOf('\n'); + return newline > 0 && text.charAt(newline - 1) == '\r' ? "\r\n" : "\n"; + } + + /** + * Strip a byte-order mark and convert every line ending to LF. + * + * @param text the file content as read + * @return the normalized content + */ + public static String normalize(String text) { + String withoutBom = text.isEmpty() || text.charAt(0) != BOM ? text : text.substring(1); + return withoutBom.replace("\r\n", "\n").replace("\r", "\n"); + } + + /** + * Put the file's own line ending and byte-order mark back. + * + * @param normalized the edited content in LF form + * @param original the content as it was read, for its ending and mark + * @return the content to write + */ + public static String denormalize(String normalized, String original) { + String ending = lineEnding(original); + String restored = "\n".equals(ending) ? normalized : normalized.replace("\n", ending); + return !original.isEmpty() && original.charAt(0) == BOM ? BOM + restored : restored; + } + + /** + * Apply every edit to {@code original}, or none of them. + * + * @param original the file content as read + * @param edits the edits, applied in order to the result of the previous one + * @return the content to write back, with the original line ending and mark + * @throws EditException when any edit does not match, naming what went wrong + */ + public static String apply(String original, List edits) { + if (edits.isEmpty()) { + throw new EditException("No edits were given."); + } + String content = normalize(original); + for (int i = 0; i < edits.size(); i++) { + Edit edit = edits.get(i); + String where = edits.size() == 1 ? "" : " (edit " + (i + 1) + " of " + edits.size() + ")"; + content = applyOne(content, edit, where); + } + return denormalize(content, original); + } + + private static String applyOne(String content, Edit edit, String where) { + String oldString = normalize(edit.oldString()); + if (oldString.isEmpty()) { + throw new EditException("old_string must not be empty" + where + "."); + } + if (oldString.equals(edit.newString())) { + throw new EditException("old_string and new_string are identical" + where + "."); + } + List lines = matchLines(content, oldString); + if (lines.isEmpty()) { + throw new EditException(notFoundMessage(content, oldString, where)); + } + if (lines.size() > 1 && !edit.replaceAll()) { + throw new EditException("old_string occurs " + lines.size() + " times" + where + ", on lines " + join(lines) + + ". Add surrounding lines to make it unique, or set replace_all."); + } + return edit.replaceAll() + ? content.replace(oldString, edit.newString()) + : replaceFirst(content, oldString, edit.newString()); + } + + private static String replaceFirst(String content, String oldString, String newString) { + int index = content.indexOf(oldString); + return content.substring(0, index) + newString + content.substring(index + oldString.length()); + } + + /** + * The 1-based line numbers where {@code oldString} starts. + * + * @param content the normalized content + * @param oldString the normalized text to find + * @return every match position, as line numbers + */ + static List matchLines(String content, String oldString) { + List lines = new ArrayList<>(); + int index = content.indexOf(oldString); + while (index >= 0) { + lines.add(lineOf(content, index)); + index = content.indexOf(oldString, index + 1); + } + return lines; + } + + private static int lineOf(String content, int index) { + int line = 1; + for (int i = 0; i < index; i++) { + if (content.charAt(i) == '\n') { + line++; + } + } + return line; + } + + /** + * The message for a miss: the closest lines in the file, so the next attempt can be corrected + * rather than guessed. + * + * @param content the normalized content + * @param oldString the normalized text that was not found + * @param where which edit of the batch failed + * @return the message + */ + static String notFoundMessage(String content, String oldString, String where) { + StringBuilder message = new StringBuilder("old_string was not found" + where + " — nothing was changed."); + List candidates = candidates(content, oldString); + if (candidates.isEmpty()) { + message.append(" No similar line exists; read the file again and copy the text from it" + + " (without the line numbers the read tool prints)."); + } else { + message.append(" The closest lines in the file are:"); + for (String candidate : candidates) { + message.append(System.lineSeparator()).append(" ").append(candidate); + } + } + return message.toString(); + } + + /** + * The lines most similar to the first line of {@code oldString}. + * + * @param content the normalized content + * @param oldString the text that was not found + * @return up to {@value #MAX_CANDIDATES} lines as {@code " 12: text"}, best first + */ + static List candidates(String content, String oldString) { + String needle = oldString.lines().findFirst().orElse(oldString).strip(); + if (needle.isEmpty()) { + return List.of(); + } + record Candidate(int line, String text, double score) {} + List scored = new ArrayList<>(); + String[] lines = content.split("\n", -1); + for (int i = 0; i < lines.length; i++) { + double score = similarity(needle, lines[i].strip()); + if (score >= CANDIDATE_MIN_SIMILARITY) { + scored.add(new Candidate(i + 1, lines[i], score)); + } + } + return scored.stream() + .sorted((a, b) -> Double.compare(b.score(), a.score())) + .limit(MAX_CANDIDATES) + .map(candidate -> candidate.line() + ": " + candidate.text()) + .toList(); + } + + /** + * How similar two lines are, as 1 minus the edit distance over the longer length. + * + * @param a one line + * @param b the other + * @return a value between 0 and 1 + */ + static double similarity(String a, String b) { + if (a.equals(b)) { + return 1; + } + if (a.isEmpty() || b.isEmpty()) { + return 0; + } + int distance = editDistance(a, b); + return 1.0 - (double) distance / Math.max(a.length(), b.length()); + } + + private static int editDistance(String a, String b) { + int[] previous = new int[b.length() + 1]; + int[] current = new int[b.length() + 1]; + for (int j = 0; j <= b.length(); j++) { + previous[j] = j; + } + for (int i = 1; i <= a.length(); i++) { + current[0] = i; + for (int j = 1; j <= b.length(); j++) { + int substitution = previous[j - 1] + (a.charAt(i - 1) == b.charAt(j - 1) ? 0 : 1); + current[j] = Math.min(substitution, Math.min(previous[j] + 1, current[j - 1] + 1)); + } + int[] swap = previous; + previous = current; + current = swap; + } + return previous[b.length()]; + } + + private static String join(List lines) { + StringBuilder text = new StringBuilder(); + for (int i = 0; i < lines.size(); i++) { + text.append(i == 0 ? "" : ", ").append(lines.get(i)); + } + return text.toString(); + } + + /** + * The lines around an edit, so the answer shows what the file now looks like instead of the whole + * file or nothing at all. + * + * @param content the content after the edit, in LF form + * @param oldString the text that was replaced, to locate the region + * @param newString what it was replaced with + * @param context how many lines above and below are included + * @return the snippet as {@code " 12: text"} lines, or {@code null} when the region is gone + */ + static @Nullable String snippet(String content, String oldString, String newString, int context) { + int index = newString.isEmpty() ? content.indexOf(normalize(oldString)) : content.indexOf(newString); + if (index < 0) { + return null; + } + int match = lineOf(content, index); + int start = Math.max(1, match - context); + // the replacement may be several lines; the window ends after it, not after a fixed size + int end = match + Math.max(0, (int) newString.lines().count() - 1) + context; + String[] lines = content.split("\n", -1); + StringBuilder text = new StringBuilder(); + for (int i = start; i <= Math.min(end, lines.length); i++) { + text.append(i == start ? "" : System.lineSeparator()) + .append(" ") + .append(i) + .append(": ") + .append(lines[i - 1]); + } + return text.toString(); + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ToolCallLog.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ToolCallLog.java new file mode 100644 index 000000000..9530719a1 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/ToolCallLog.java @@ -0,0 +1,100 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.time.LocalTime; +import java.time.format.DateTimeFormatter; +import java.util.ArrayList; +import java.util.List; + +/** + * Every tool call of the session, in order, for {@code /calls}. + * + *

It answers one question the transcript cannot: did that actually happen? A model that + * runs out of context, or simply drifts, starts describing work instead of doing it — reporting an + * exit code, a created file, a passing test, all invented. The scrollback looks convincing, because + * the prose is the same either way. This log only ever grows when a tool really ran, so an empty or + * short list is the proof. + * + *

Kept small on purpose: one line per call, arguments and result cut hard. It is a receipt, not a + * second transcript. + */ +public final class ToolCallLog { + + /** Characters kept of the arguments and of the result. */ + private static final int PREVIEW_CHARS = 120; + + private static final DateTimeFormatter TIME = DateTimeFormatter.ofPattern("HH:mm:ss"); + + private final List entries = new ArrayList<>(); + + /** + * One recorded call. + * + * @param time when it ran + * @param turn the user turn it belonged to, counting from 1 + * @param name the tool + * @param arguments the arguments, shortened + * @param result what came back, shortened + */ + public record Entry(LocalTime time, int turn, String name, String arguments, String result) {} + + /** + * Record the calls of one finished turn. + * + * @param turn the turn number + * @param rounds the calls, in order + */ + public void add(int turn, List rounds) { + for (ConsoleSession.ToolRound round : rounds) { + entries.add(new Entry( + LocalTime.now(), turn, round.name(), cut(round.argumentsJson()), cut(oneLine(round.result())))); + } + } + + /** + * How many calls were made in this session. + * + * @return the count + */ + public int size() { + return entries.size(); + } + + /** + * The log as the console shows it. + * + * @return one line per call, or a sentence saying there were none + */ + public String render() { + if (entries.isEmpty()) { + return "No tool has been called in this session — everything so far was text only."; + } + StringBuilder text = new StringBuilder(); + for (Entry entry : entries) { + text.append(entry.time().format(TIME)) + .append(" turn ") + .append(entry.turn()) + .append(" ") + .append(entry.name()) + .append(" ") + .append(entry.arguments()) + .append(System.lineSeparator()) + .append(" ↳ ") + .append(entry.result()) + .append(System.lineSeparator()); + } + text.append(entries.size()).append(" calls"); + return text.toString(); + } + + private static String oneLine(String text) { + return text.replace("\r\n", " ").replace('\n', ' ').strip(); + } + + private static String cut(String text) { + return text.length() <= PREVIEW_CHARS ? text : text.substring(0, PREVIEW_CHARS) + " …"; + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Transcript.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Transcript.java new file mode 100644 index 000000000..c03cc9603 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/Transcript.java @@ -0,0 +1,262 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.io.IOException; +import java.nio.charset.StandardCharsets; +import java.nio.file.Files; +import java.nio.file.Path; +import java.nio.file.StandardOpenOption; +import java.time.LocalDateTime; +import java.time.format.DateTimeFormatter; +import java.util.List; +import java.util.concurrent.CopyOnWriteArrayList; +import org.jspecify.annotations.Nullable; + +/** + * What was said in this session, in the order it was said, with the time. + * + *

It is not the conversation the model is sent. That one is rewritten by {@code /compact} — a + * summary replaces the turns it summarises — and it never carried a time at all, so it cannot answer + * "what did I ask before lunch" or "what did it actually reply". This record only ever grows. + * {@code /compact} adds a note to it and changes nothing else; {@code /clear} empties it, because + * that command means "forget this session" and leaving the text behind would make that untrue. + * + *

A list, not a map keyed by the time. Two entries can share a millisecond — a tool result + * and the answer that follows it regularly do — and a map would keep one of them and silently drop + * the other. Insertion order already is time order, which is the only ordering anyone wants here. + * + *

With a file configured, every entry is also appended as it happens, so a session that is killed + * still leaves what it had. Without one, {@code /save} writes the whole thing on request. + */ +public final class Transcript { + + private static final DateTimeFormatter STAMP = DateTimeFormatter.ofPattern("yyyy-MM-dd HH:mm:ss"); + + private static final DateTimeFormatter FILE_STAMP = DateTimeFormatter.ofPattern("yyyy-MM-dd_HH-mm-ss"); + + /** The shape {@link #format} writes: {@code [stamp] kind: text}. */ + private static final java.util.regex.Pattern HEAD = + java.util.regex.Pattern.compile("\\[(\\d{4}-\\d{2}-\\d{2} \\d{2}:\\d{2}:\\d{2})\\] (\\w+): (.*)"); + + /** Who said it. */ + public enum Kind { + /** What the user typed. */ + USER("you"), + /** What the model answered. */ + AGENT("agent"), + /** A tool call and what it returned. */ + TOOL("tool"), + /** Something the session did: compacted, interrupted, mode changed. */ + NOTE("note"); + + private final String label; + + Kind(String label) { + this.label = label; + } + + /** + * The word written in the file. + * + * @return the label + */ + public String label() { + return label; + } + } + + /** + * One thing that was said. + * + * @param at when + * @param kind who + * @param text what, verbatim and uncut + */ + public record Entry(LocalDateTime at, Kind kind, String text) {} + + private final List entries = new CopyOnWriteArrayList<>(); + private final @Nullable Path liveFile; + + /** + * Keep a session transcript in memory only. + */ + public Transcript() { + this(null); + } + + /** + * Keep a session transcript, and append every entry to a file as it happens. + * + * @param liveFile the file to append to, or {@code null} to keep it in memory + */ + public Transcript(@Nullable Path liveFile) { + this.liveFile = liveFile; + } + + /** + * Record something. + * + *

Blank text is dropped: an empty answer is not worth a line, and the file would fill with + * them on a session of interrupted turns. + * + * @param kind who said it + * @param text what was said + */ + public void add(Kind kind, String text) { + if (text == null || text.isBlank()) { + return; + } + Entry entry = new Entry(LocalDateTime.now(), kind, text.strip()); + entries.add(entry); + appendLive(entry); + } + + /** + * Everything recorded, oldest first. + * + * @return the entries + */ + public List entries() { + return List.copyOf(entries); + } + + /** + * How many entries were recorded. + * + * @return the count + */ + public int size() { + return entries.size(); + } + + /** Forget the session, as {@code /clear} means it. */ + public void clear() { + entries.clear(); + } + + /** + * The whole transcript as text. + * + * @return one block per entry, oldest first + */ + public String render() { + StringBuilder text = new StringBuilder(); + for (Entry entry : entries) { + text.append(format(entry)); + } + return text.toString(); + } + + /** + * Write the transcript to a file. + * + * @param directory where it goes, normally the workspace + * @param name the file name, or {@code null} for one named after the time it was written + * @return the file that was written + * @throws IOException if it cannot be written + */ + public Path save(Path directory, @Nullable String name) throws IOException { + String fileName = name == null || name.isBlank() + ? "transcript-" + LocalDateTime.now().format(FILE_STAMP) + ".txt" + : name.strip(); + Path file = directory.resolve(fileName); + Files.createDirectories(file.toAbsolutePath().getParent()); + Files.writeString(file, render(), StandardCharsets.UTF_8); + return file; + } + + /** + * Read back a transcript that was written by {@link #save} or by {@code --transcript}. + * + *

A line that does not start with a stamp belongs to the entry above it: an answer keeps its + * newlines when it is written, so an entry is not the same thing as a line. Reading line by line + * would turn one answer into several, each of them nonsense on its own. + * + *

Anything before the first stamped line is ignored rather than guessed at — a file that is not + * a transcript yields no entries instead of one wrong one. + * + * @param text the file content + * @return the entries, oldest first + */ + public static List parse(String text) { + List parsed = new java.util.ArrayList<>(); + StringBuilder pending = new StringBuilder(); + LocalDateTime at = null; + Kind kind = null; + for (String line : text.split("\\r?\\n", -1)) { + java.util.regex.Matcher head = HEAD.matcher(line); + if (head.matches()) { + flush(parsed, at, kind, pending); + at = LocalDateTime.parse(head.group(1), STAMP); + kind = kindOf(head.group(2)); + pending.setLength(0); + pending.append(head.group(3)); + } else if (kind != null) { + pending.append(System.lineSeparator()).append(line); + } + } + flush(parsed, at, kind, pending); + return List.copyOf(parsed); + } + + private static void flush( + List parsed, @Nullable LocalDateTime at, @Nullable Kind kind, StringBuilder pending) { + if (at != null && kind != null && !pending.toString().isBlank()) { + parsed.add(new Entry(at, kind, pending.toString().strip())); + } + } + + private static Kind kindOf(String label) { + for (Kind candidate : Kind.values()) { + if (candidate.label().equals(label)) { + return candidate; + } + } + return Kind.NOTE; + } + + /** + * Replace everything recorded with what was read from a file. + * + * @param loaded the entries to keep + */ + public void replaceWith(List loaded) { + entries.clear(); + entries.addAll(loaded); + } + + private void appendLive(Entry entry) { + if (liveFile == null) { + return; + } + try { + Path parent = liveFile.toAbsolutePath().getParent(); + if (parent != null) { + Files.createDirectories(parent); + } + Files.writeString( + liveFile, + format(entry), + StandardCharsets.UTF_8, + StandardOpenOption.CREATE, + StandardOpenOption.APPEND); + } catch (IOException e) { + // A transcript that cannot be written must not end the session: the point of it is to + // survive a bad ending, so it may not cause one. + } + } + + /** + * One entry as it appears in the file. + * + * @param entry the entry + * @return the block, ending in a newline + */ + private static String format(Entry entry) { + String separator = System.lineSeparator(); + return "[" + entry.at().format(STAMP) + "] " + entry.kind().label() + ": " + entry.text() + separator; + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TurnActivity.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TurnActivity.java new file mode 100644 index 000000000..18b18d1a7 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/TurnActivity.java @@ -0,0 +1,41 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.util.concurrent.atomic.AtomicBoolean; + +/** + * A switch that stops the activity line from redrawing while something else owns the terminal. + * + *

It exists because a turn runs on its own thread. Atmosphere's {@code execute} is synchronous — + * it returns only when the whole turn including every tool round is done — so the console thread has + * to drive the spinner while a second thread runs the turn. That is fine until the approval prompt + * appears: it reads a single key in raw mode on the turn's thread, and a status redraw arriving from + * the console thread in the middle of that writes escape sequences across the question. So the prompt + * pauses the redraw for as long as it is waiting for an answer. + */ +public final class TurnActivity { + + private final AtomicBoolean paused = new AtomicBoolean(); + + /** Stop redrawing the activity line. */ + public void pause() { + paused.set(true); + } + + /** Redraw it again. */ + public void resume() { + paused.set(false); + } + + /** + * Whether redrawing is currently suspended. + * + * @return {@code true} while something else owns the terminal + */ + public boolean isPaused() { + return paused.get(); + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/WorkspaceSearch.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/WorkspaceSearch.java new file mode 100644 index 000000000..7bca82cf8 --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/WorkspaceSearch.java @@ -0,0 +1,243 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.io.IOException; +import java.nio.charset.StandardCharsets; +import java.nio.file.FileVisitResult; +import java.nio.file.Files; +import java.nio.file.Path; +import java.nio.file.SimpleFileVisitor; +import java.nio.file.attribute.BasicFileAttributes; +import java.util.ArrayList; +import java.util.LinkedHashMap; +import java.util.List; +import java.util.Map; +import java.util.Set; +import java.util.regex.Matcher; +import java.util.regex.Pattern; +import java.util.regex.PatternSyntaxException; + +/** + * Searching the workspace for text. + * + *

Why this exists instead of the framework's grep. Atmosphere walks the workspace + * alphabetically and spends one global 2-second deadline and one global 500-hit budget on + * that walk. In any real project {@code .git} sorts before {@code src}, so both budgets are consumed + * by the repository's own object store, the build output and {@code node_modules} before the source + * is reached — the tool then reports a truncated result that does not contain the code at all. + * Excluding those directories is the entire fix, and it is a filter in the walk. + * + *

The second reason is what the model reads. Results are grouped by file with line numbers (so a + * ranged read can follow), every line is cut at {@value #MAX_LINE_CHARS} characters (a minified file + * is otherwise a single 200 KB "line"), and truncation is always stated: a search that + * silently returns part of the matches is read as a complete answer, which is the documented way this + * class of tool misleads an agent. + */ +public final class WorkspaceSearch { + + /** Directories never searched, whatever the pattern. */ + public static final Set EXCLUDED_DIRECTORIES = Set.of( + ".git", + ".hg", + ".svn", + "node_modules", + "target", + "build", + "dist", + "out", + ".gradle", + ".idea", + ".mvn", + ".venv", + "venv", + "__pycache__", + ".cache", + ".next", + "vendor"); + + /** Files larger than this are skipped: they are generated, not written. */ + public static final long MAX_FILE_BYTES = 2_000_000; + + /** A matching line is cut here. */ + public static final int MAX_LINE_CHARS = 200; + + /** Matches shown per file before the rest is summarized. */ + public static final int MAX_MATCHES_PER_FILE = 20; + + /** The whole search stops here. */ + public static final int MAX_TOTAL_MATCHES = 100; + + /** A pattern that backtracks is cut off after this. */ + public static final long DEADLINE_MILLIS = 5_000; + + private WorkspaceSearch() {} + + /** + * One matching line. + * + * @param path the path relative to the workspace root, with {@code /} separators + * @param line the 1-based line number + * @param text the line, already cut to {@value #MAX_LINE_CHARS} characters + */ + public record Hit(String path, int line, String text) {} + + /** + * The result of a search. + * + * @param hits the matches, in walk order + * @param filesWithMatches how many files matched, including those not shown + * @param truncated whether the search stopped before the end + */ + public record Result(List hits, int filesWithMatches, boolean truncated) {} + + /** + * Search below {@code root}. + * + * @param root the workspace root + * @param pattern the regular expression + * @param subdirectory the directory to search, relative to the root, or {@code null} for all + * @param glob an optional file-name glob such as {@code *.java}, or {@code null} for all files + * @return the matches + * @throws IllegalArgumentException when the pattern is not a valid regular expression + */ + public static Result search(Path root, String pattern, String subdirectory, String glob) { + Pattern regex; + try { + regex = Pattern.compile(pattern); + } catch (PatternSyntaxException e) { + throw new IllegalArgumentException("Invalid regular expression: " + e.getMessage(), e); + } + Path start = subdirectory == null || subdirectory.isBlank() + ? root + : root.resolve(subdirectory).normalize(); + if (!start.startsWith(root) || !Files.isDirectory(start)) { + return new Result(List.of(), 0, false); + } + java.nio.file.PathMatcher nameMatcher = + glob == null || glob.isBlank() ? null : start.getFileSystem().getPathMatcher("glob:" + glob); + + List hits = new ArrayList<>(); + Set matchedFiles = new java.util.LinkedHashSet<>(); + long deadline = System.currentTimeMillis() + DEADLINE_MILLIS; + boolean[] truncated = {false}; + try { + Files.walkFileTree(start, Set.of(), Integer.MAX_VALUE, new SimpleFileVisitor<>() { + @Override + public FileVisitResult preVisitDirectory(Path directory, BasicFileAttributes attributes) { + String name = directory.getFileName() == null + ? "" + : directory.getFileName().toString(); + return EXCLUDED_DIRECTORIES.contains(name) || (!directory.equals(start) && name.startsWith(".")) + ? FileVisitResult.SKIP_SUBTREE + : FileVisitResult.CONTINUE; + } + + @Override + public FileVisitResult visitFile(Path file, BasicFileAttributes attributes) { + if (hits.size() >= MAX_TOTAL_MATCHES || System.currentTimeMillis() > deadline) { + truncated[0] = true; + return FileVisitResult.TERMINATE; + } + // Use the attributes the walk already has: a separate isRegularFile/size call per + // file goes through CreateFileW on Windows and dominates the walk. + if (!attributes.isRegularFile() || attributes.size() > MAX_FILE_BYTES) { + return FileVisitResult.CONTINUE; + } + if (nameMatcher != null && !nameMatcher.matches(file.getFileName())) { + return FileVisitResult.CONTINUE; + } + searchFile(root, file, regex, hits, matchedFiles, truncated); + return FileVisitResult.CONTINUE; + } + + @Override + public FileVisitResult visitFileFailed(Path file, IOException e) { + return FileVisitResult.CONTINUE; + } + }); + } catch (IOException e) { + throw new IllegalArgumentException("Search failed: " + e.getMessage(), e); + } + return new Result(List.copyOf(hits), matchedFiles.size(), truncated[0]); + } + + private static void searchFile( + Path root, Path file, Pattern regex, List hits, Set matchedFiles, boolean[] truncated) { + List lines; + try { + // Binary content fails to decode as UTF-8 and is skipped -- the cheap equivalent of + // ripgrep's "a file with a NUL byte is binary". + lines = Files.readAllLines(file, StandardCharsets.UTF_8); + } catch (IOException | RuntimeException e) { + return; + } + String relative = root.relativize(file).toString().replace('\\', '/'); + int inThisFile = 0; + for (int i = 0; i < lines.size(); i++) { + if (hits.size() >= MAX_TOTAL_MATCHES) { + truncated[0] = true; + return; + } + Matcher matcher = regex.matcher(lines.get(i)); + if (!matcher.find()) { + continue; + } + matchedFiles.add(relative); + inThisFile++; + if (inThisFile > MAX_MATCHES_PER_FILE) { + truncated[0] = true; + return; + } + String text = lines.get(i); + hits.add(new Hit( + relative, + i + 1, + text.length() <= MAX_LINE_CHARS ? text : text.substring(0, MAX_LINE_CHARS) + " …[cut]")); + } + } + + /** + * Render a result for the model: grouped by file, line-numbered, and honest about truncation. + * + * @param result the search result + * @param filesOnly whether to list only the file names + * @return the text the tool returns + */ + public static String format(Result result, boolean filesOnly) { + if (result.hits().isEmpty()) { + return "(no matches)"; + } + Map> byFile = new LinkedHashMap<>(); + for (Hit hit : result.hits()) { + byFile.computeIfAbsent(hit.path(), key -> new ArrayList<>()).add(hit); + } + StringBuilder text = new StringBuilder(); + if (filesOnly) { + byFile.keySet().forEach(path -> text.append(path).append(System.lineSeparator())); + } else { + for (Map.Entry> file : byFile.entrySet()) { + text.append(file.getKey()).append(System.lineSeparator()); + for (Hit hit : file.getValue()) { + text.append(" ") + .append(hit.line()) + .append(": ") + .append(hit.text()) + .append(System.lineSeparator()); + } + text.append(System.lineSeparator()); + } + } + text.append(result.hits().size()) + .append(" matches in ") + .append(result.filesWithMatches()) + .append(" files"); + if (result.truncated()) { + text.append(" (TRUNCATED — there are more; narrow the search with `glob`, `dir`" + + " or a more specific pattern)"); + } + return text.toString(); + } +} diff --git a/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/WorkspaceTools.java b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/WorkspaceTools.java new file mode 100644 index 000000000..36970ff7d --- /dev/null +++ b/llama-atmosphere-agent/src/main/java/net/ladenthin/llama/atmosphere/WorkspaceTools.java @@ -0,0 +1,398 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import java.nio.file.Path; +import java.util.ArrayList; +import java.util.LinkedHashSet; +import java.util.List; +import java.util.Map; +import java.util.Set; +import org.atmosphere.ai.fs.AgentFileSystem; +import org.atmosphere.ai.fs.FileSystemTools; +import org.atmosphere.ai.fs.WorkspaceAgentFileSystem; +import org.atmosphere.ai.tool.ToolDefinition; +import org.atmosphere.ai.tool.ToolExecutor; +import org.atmosphere.ai.tool.ToolKind; +import org.atmosphere.ai.tool.ToolParameter; +import org.jspecify.annotations.Nullable; + +/** + * The three tools this agent provides itself, replacing the framework's versions of the same names: + * {@code read_file}, {@code edit_file} and {@code grep}. Everything else — {@code ls}, + * {@code write_file}, {@code glob}, {@code delete}, {@code rename} — stays Atmosphere's. + * + *

They are replacements, not additions, on purpose: two tools that both claim to "read a file" is + * the worst outcome for tool selection. Underneath they use the same {@link AgentFileSystem} the + * framework binds to the session, so path validation, the workspace confinement and the size limits + * are unchanged; only the parameters the model sees and the text it gets back are ours. + * + *

Why each one differs from the framework's: + * + *

    + *
  • read_file takes {@code offset}/{@code limit} and prints line numbers. Reading a whole + * file costs context for nothing and measurably lowers task success (SWE-agent: 12.7 % with + * whole files against 18.0 % with a 100-line window). + *
  • edit_file normalizes line endings (the framework compares raw content, so no edit ever + * matches in a CRLF file), explains a miss with the nearest lines, names the line numbers of an + * ambiguous match, and can apply several edits at once — see {@link TextEdits}. + *
  • grep skips {@code .git}, build output and dependency directories — the framework's walk + * is alphabetical with global budgets, so those directories eat the result before the source is + * reached — and states when it truncated; see {@link WorkspaceSearch}. + *
+ */ +public final class WorkspaceTools { + + /** Lines returned by {@code read_file} when the model gives no limit. */ + public static final int DEFAULT_READ_LIMIT = 400; + + /** Lines of context shown around a completed edit. */ + private static final int EDIT_SNIPPET_CONTEXT = 3; + + private WorkspaceTools() {} + + /** + * Remembers which files were read, so an edit can insist on it. + * + *

Not a staleness check: an {@code old_string} that matches the current content exactly and + * unambiguously is safe whether or not the file changed meanwhile. What this prevents is the + * other case — a model inventing the text it wants to replace. One instance per session. + */ + public static final class ReadTracker { + private final Set read = new LinkedHashSet<>(); + + /** + * Note that a file was read. + * + * @param path the path as the model wrote it + */ + public void markRead(String path) { + read.add(normalizePath(path)); + } + + /** + * Whether a file was read in this session. + * + * @param path the path as the model wrote it + * @return {@code true} when it was read before + */ + public boolean wasRead(String path) { + return read.contains(normalizePath(path)); + } + + private static String normalizePath(String path) { + return path.replace('\\', '/').replaceAll("^\\./", ""); + } + } + + /** + * The tool set: Atmosphere's tools where they are fine, ours where they are not. + * + * @param tracker the read tracker shared by {@code read_file} and {@code edit_file} + * @return the tools to offer the model + */ + public static List all(ReadTracker tracker) { + return List.of( + FileSystemTools.ls(), + readFile(tracker), + FileSystemTools.writeFile(), + editFile(tracker), + FileSystemTools.glob(), + grep(), + FileSystemTools.delete(), + FileSystemTools.rename()); + } + + /** + * {@code read_file}: a window of a file, with line numbers. + * + * @param tracker records the read so an edit is allowed afterwards + * @return the tool + */ + public static ToolDefinition readFile(ReadTracker tracker) { + return ToolDefinition.builder( + FileSystemTools.READ_FILE, + "Read a file from the workspace. Returns the lines numbered as ` 12: text`." + + " Reads at most " + DEFAULT_READ_LIMIT + " lines at a time; use offset and limit to" + + " page through a longer file. The line numbers are display only — never include them" + + " in old_string when you edit.") + .parameter("file_path", "File path relative to the workspace root", "string", true) + .parameter("offset", "First line to show, 1-based (default 1)", "integer", false) + .parameter("limit", "How many lines to show (default " + DEFAULT_READ_LIMIT + ")", "integer", false) + .returnType("string") + .executor(withFileSystem((arguments, injectables) -> { + AgentFileSystem fs = fileSystem(injectables); + if (fs == null) { + return "File tools unavailable: no agent filesystem is bound to this session."; + } + String path = string(arguments, "file_path"); + if (path == null) { + return "Error: file_path is required"; + } + try { + List lines = fs.read(path).lines().toList(); + int offset = Math.max(1, integer(arguments, "offset", 1)); + int limit = Math.max(1, integer(arguments, "limit", DEFAULT_READ_LIMIT)); + tracker.markRead(path); + return window(lines, offset, limit); + } catch (IllegalArgumentException e) { + return "Error: " + e.getMessage(); + } + })) + .kind(ToolKind.READ) + .build(); + } + + /** + * Render the requested window with line numbers and say what was left out. + * + * @param lines the file's lines + * @param offset the first line, 1-based + * @param limit how many lines + * @return the text for the model + */ + static String window(List lines, int offset, int limit) { + if (lines.isEmpty()) { + return "(empty file)"; + } + if (offset > lines.size()) { + return "(offset " + offset + " is past the end; the file has " + lines.size() + " lines)"; + } + int last = Math.min(lines.size(), offset + limit - 1); + StringBuilder text = new StringBuilder(); + for (int i = offset; i <= last; i++) { + text.append(" ").append(i).append(": ").append(lines.get(i - 1)).append(System.lineSeparator()); + } + if (offset > 1 || last < lines.size()) { + text.append("(showing lines ") + .append(offset) + .append("–") + .append(last) + .append(" of ") + .append(lines.size()) + .append("; use offset/limit for the rest)"); + } + return text.toString(); + } + + /** + * {@code edit_file}: exact replacement with line-ending handling, corrective errors, and several + * edits in one call. + * + * @param tracker enforces that the file was read first + * @return the tool + */ + public static ToolDefinition editFile(ReadTracker tracker) { + ToolParameter edits = ToolParameter.ofArray( + "edits", + "Several replacements applied to the same file, in order. Use instead of" + + " old_string/new_string. All of them must match, or the file is left untouched.", + false, + new ToolParameter( + "edit", + "One replacement", + "object", + false, + List.of(), + null, + List.of( + new ToolParameter("old_string", "The exact text to replace", "string", true), + new ToolParameter("new_string", "The replacement", "string", true), + new ToolParameter("replace_all", "Replace every occurrence", "boolean", false)))); + return ToolDefinition.builder( + FileSystemTools.EDIT_FILE, + "Edit a file by replacing exact text. Read the file first and copy old_string from it" + + " (without the line numbers the read tool prints). old_string must match exactly" + + " once unless replace_all is set. Line endings are handled for you.") + .parameter("file_path", "File path relative to the workspace root", "string", true) + .parameter("old_string", "The exact text to replace", "string", false) + .parameter("new_string", "The replacement text", "string", false) + .parameter("replace_all", "Replace every occurrence instead of requiring exactly one", "boolean", false) + .parameter(edits) + .returnType("string") + .executor(withFileSystem((arguments, injectables) -> { + AgentFileSystem fs = fileSystem(injectables); + if (fs == null) { + return "File tools unavailable: no agent filesystem is bound to this session."; + } + String path = string(arguments, "file_path"); + if (path == null) { + return "Error: file_path is required"; + } + List list; + try { + list = parseEdits(arguments); + } catch (IllegalArgumentException e) { + return "Error: " + e.getMessage(); + } + if (!tracker.wasRead(path)) { + return "Error: read " + path + " before editing it, so old_string comes from the file" + + " rather than from memory."; + } + try { + String original = fs.read(path); + String edited = TextEdits.apply(original, list); + fs.write(path, edited); + return "Edited " + path + " (" + list.size() + (list.size() == 1 ? " edit)" : " edits)") + + describe(edited, list); + } catch (TextEdits.EditException e) { + return "Error: " + e.getMessage(); + } catch (IllegalArgumentException e) { + return "Error: " + e.getMessage(); + } + })) + .kind(ToolKind.EDIT) + .build(); + } + + private static String describe(String edited, List edits) { + TextEdits.Edit last = edits.get(edits.size() - 1); + String snippet = TextEdits.snippet( + TextEdits.normalize(edited), last.oldString(), last.newString(), EDIT_SNIPPET_CONTEXT); + return snippet == null ? "" : System.lineSeparator() + snippet; + } + + /** + * Read the edits from the call, in either shape. + * + * @param arguments the tool arguments + * @return the edits, in order + * @throws IllegalArgumentException when neither shape is present or an entry is incomplete + */ + static List parseEdits(Map arguments) { + Object many = arguments == null ? null : arguments.get("edits"); + if (many instanceof Iterable entries) { + List edits = new ArrayList<>(); + for (Object entry : entries) { + if (!(entry instanceof Map map)) { + throw new IllegalArgumentException( + "every entry of edits must be an object with old_string" + " and new_string"); + } + Object oldString = map.get("old_string"); + Object newString = map.get("new_string"); + if (oldString == null) { + throw new IllegalArgumentException("every entry of edits needs old_string"); + } + edits.add(new TextEdits.Edit( + oldString.toString(), + newString == null ? "" : newString.toString(), + Boolean.TRUE.equals(map.get("replace_all")) + || "true".equalsIgnoreCase(String.valueOf(map.get("replace_all"))))); + } + if (edits.isEmpty()) { + throw new IllegalArgumentException("edits was empty"); + } + return edits; + } + String oldString = string(arguments, "old_string"); + if (oldString == null) { + throw new IllegalArgumentException("old_string is required (or pass edits)"); + } + String newString = string(arguments, "new_string"); + return List.of( + new TextEdits.Edit(oldString, newString == null ? "" : newString, bool(arguments, "replace_all"))); + } + + /** + * {@code grep}: search the workspace, skipping what is not source. + * + * @return the tool + */ + public static ToolDefinition grep() { + return ToolDefinition.builder( + FileSystemTools.GREP, + "Search the workspace with a regular expression. Returns matching lines grouped by file" + + " as ` 12: text`, capped at " + WorkspaceSearch.MAX_TOTAL_MATCHES + " matches;" + + " .git, build output and dependency directories are skipped. Says so when the" + + " result is truncated.") + .parameter("pattern", "The regular expression to search for", "string", true) + .parameter("dir", "Directory to search under, relative to the workspace root", "string", false) + .parameter("glob", "Only search files matching this name pattern, e.g. *.java", "string", false) + .parameter("files_only", "Return just the file names instead of the matching lines", "boolean", false) + .returnType("string") + .executor(withFileSystem((arguments, injectables) -> { + AgentFileSystem fs = fileSystem(injectables); + if (!(fs instanceof WorkspaceAgentFileSystem workspace)) { + return "Search unavailable: this session has no workspace directory."; + } + String pattern = string(arguments, "pattern"); + if (pattern == null) { + return "Error: pattern is required"; + } + try { + Path root = workspace.root(); + WorkspaceSearch.Result result = WorkspaceSearch.search( + root, pattern, string(arguments, "dir"), string(arguments, "glob")); + return WorkspaceSearch.format(result, bool(arguments, "files_only")); + } catch (IllegalArgumentException e) { + return "Error: " + e.getMessage(); + } + })) + .kind(ToolKind.READ) + .build(); + } + + /** + * Adapt a two-argument function to {@link ToolExecutor}, whose single abstract method takes only + * the arguments — the injectables arrive through its default overload, and that is where the + * session's filesystem lives. + * + * @param body what the tool does + * @return the executor + */ + private static ToolExecutor withFileSystem(ToolBody body) { + return new ToolExecutor() { + @Override + public Object execute(Map arguments) { + return execute(arguments, Map.of()); + } + + @Override + public Object execute(Map arguments, Map, Object> injectables) { + return body.run(arguments, injectables); + } + }; + } + + /** The body of a tool: arguments plus the session's injectables. */ + @FunctionalInterface + private interface ToolBody { + /** + * Run the tool. + * + * @param arguments the call's arguments + * @param injectables the session scope, carrying the filesystem + * @return what the model sees + */ + Object run(Map arguments, Map, Object> injectables); + } + + private static @Nullable AgentFileSystem fileSystem(@Nullable Map, Object> injectables) { + return FileSystemTools.resolveFileSystem(injectables == null ? Map.of() : injectables) + .orElse(null); + } + + private static @Nullable String string(@Nullable Map arguments, String name) { + Object value = arguments == null ? null : arguments.get(name); + return value == null || value.toString().isEmpty() ? null : value.toString(); + } + + private static boolean bool(@Nullable Map arguments, String name) { + Object value = arguments == null ? null : arguments.get(name); + return Boolean.TRUE.equals(value) || "true".equalsIgnoreCase(String.valueOf(value)); + } + + private static int integer(@Nullable Map arguments, String name, int fallback) { + Object value = arguments == null ? null : arguments.get(name); + if (value instanceof Number number) { + return number.intValue(); + } + try { + return value == null ? fallback : Integer.parseInt(value.toString().trim()); + } catch (NumberFormatException e) { + return fallback; + } + } +} diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/compact-prompt.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/compact-prompt.txt new file mode 100644 index 000000000..2e025534a --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/compact-prompt.txt @@ -0,0 +1,11 @@ +Summarize the conversation so far so that it can continue in a fresh context with nothing important lost. Write the summary as plain text with these sections, and leave out a section only when it would be empty: + +1. Goal — what the user wants, in their own words where possible. +2. Facts — what was established about the project: paths, file names, commands, versions, decisions and their reasons. +3. Work done — what was changed or produced, file by file. +4. Problems — what failed, and what fixed it or is still open. +5. State — where things stand right now. +6. Next step — the single most obvious continuation, or "none" when the task is finished. + +Quote exact names, paths, commands and error messages verbatim; they are the part that cannot be reconstructed. Do not add advice, do not repeat the instructions, and do not use any tool — answer with the summary only. +{focus} diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/compact-prompt.txt.license b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/compact-prompt.txt.license new file mode 100644 index 000000000..b918686f7 --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/compact-prompt.txt.license @@ -0,0 +1,3 @@ +SPDX-FileCopyrightText: 2026 Bernard Ladenthin + +SPDX-License-Identifier: MIT diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt new file mode 100644 index 000000000..690c4c2b7 --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt @@ -0,0 +1,29 @@ +Commands (everything else is sent to the model): + + /help this overview (/?, /commands) + /status endpoint, model, tools, mode, context use + /tools the tools offered to the model + /calls every tool call of this session, with its result (/log) + /mode [manual|auto] show or set the approval mode: manual asks first (shown as + the mode symbol on the status line), auto runs through (/approve) + /compact [focus] summarize the history and continue with the summary + (happens on its own once the context is ~70% full; --auto-compact false + turns that off, --compact-at moves the threshold) + /loop [--every 5m] [--max 20] [--check ''] + keep working on until the model answers <>; + the state lives in AGENT-LOOP.md, not in the conversation + /clear drop the history, and wipe the screen with it (/reset, /new) + /cls wipe the screen, keep the conversation (/clear-screen, or Ctrl-L) + /save [name] write what was said, with timestamps, into the workspace (/transcript) + /retry ask the last question again, dropping the answer (/again) + /load read a saved transcript back in as the conversation (/resume) + /exit leave (/quit) + +The input line sits in the box at the bottom and is there while the agent works: you can type +at any time. A line typed during a turn stops it and is sent as the next message. +Shift+Tab switches the approval mode. + +In manual mode every tool that writes or runs a command asks first: [y]es runs it once, +[n]o tells the model the user cancelled it, [a]uto stops asking for the rest of the session. +Reading tools (ls, read_file, glob, grep) never ask — they cannot change anything. +An unknown /command is sent to the model, not rejected. diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt.license b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt.license new file mode 100644 index 000000000..b918686f7 --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/help.txt.license @@ -0,0 +1,3 @@ +SPDX-FileCopyrightText: 2026 Bernard Ladenthin + +SPDX-License-Identifier: MIT diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-file-template.md b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-file-template.md new file mode 100644 index 000000000..8642b4ffe --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-file-template.md @@ -0,0 +1,17 @@ +# Agent loop + +## Task + +{task} + +## Plan + +- [ ] first step (replace this with the real plan) + +## Notes + +What was tried, what worked, what failed. Keep it short; this file is re-read every step. + +## Next step + +One sentence: what to do next. diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-file-template.md.license b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-file-template.md.license new file mode 100644 index 000000000..b918686f7 --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-file-template.md.license @@ -0,0 +1,3 @@ +SPDX-FileCopyrightText: 2026 Bernard Ladenthin + +SPDX-License-Identifier: MIT diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-prompt.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-prompt.txt new file mode 100644 index 000000000..34a7b7f09 --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-prompt.txt @@ -0,0 +1,12 @@ +Work on this task, one step at a time: + +{task} + +The file {file} is your memory between steps — the conversation itself is not kept, so anything you do not write down is lost. Read that file first. Then do exactly ONE concrete step: make the change, run the command, check the result. Afterwards update {file}: tick off what you finished, note what you learned or what failed, and write down what the next step is. + +When the whole task is finished and you have verified it, answer with exactly this line and nothing else: + +<> + +Never write that line for any other reason, never write it while planning, and never explain it. +{check_hint} diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-prompt.txt.license b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-prompt.txt.license new file mode 100644 index 000000000..b918686f7 --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/loop-prompt.txt.license @@ -0,0 +1,3 @@ +SPDX-FileCopyrightText: 2026 Bernard Ladenthin + +SPDX-License-Identifier: MIT diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/spinner-words.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/spinner-words.txt new file mode 100644 index 000000000..e16bfe437 --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/spinner-words.txt @@ -0,0 +1,26 @@ +Fettling +Whittling +Rummaging +Squirrelling +Beavering +Trundling +Ambling +Pootling +Cobbling +Kerfuffling +Whirligigging +Bustling +Beetling +Scuttling +Lumbering +Chugging +Puffing +Clattering +Burbling +Purring +Cranking +Rejigging +Tootling +Fossicking +Wombling +Head-scratching diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/spinner-words.txt.license b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/spinner-words.txt.license new file mode 100644 index 000000000..b918686f7 --- /dev/null +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/spinner-words.txt.license @@ -0,0 +1,3 @@ +SPDX-FileCopyrightText: 2026 Bernard Ladenthin + +SPDX-License-Identifier: MIT diff --git a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/system-prompt.txt b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/system-prompt.txt index 1cc6d722c..1e2f59396 100644 --- a/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/system-prompt.txt +++ b/llama-atmosphere-agent/src/main/resources/net/ladenthin/llama/atmosphere/system-prompt.txt @@ -4,4 +4,4 @@ The file tools ls, read_file, write_file, edit_file, glob, grep, delete and rena {shell_section} -Work step by step: read a file before you edit it, check the result after a change, and finish with a short summary. Answer in the user's language. +Work step by step. read_file shows a numbered window of a file — those numbers are display only, never copy them into old_string. Edit a file only after reading it (edit_file refuses otherwise), copy old_string from what you read, and check the result after a change. Finish with a short summary. Never claim that you ran a command, created a file or saw a result unless you actually called the tool in this turn and read what it returned; if you did not, say so plainly instead of describing what would happen. Answer in the user's language. diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AgentOptionsTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AgentOptionsTest.java index 44117a05f..6d039ff4e 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AgentOptionsTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AgentOptionsTest.java @@ -17,6 +17,24 @@ class AgentOptionsTest { + @Test + void plainChoosesTheLineOrientedConsoleEvenWithATerminal() { + // Both consoles stay; this is the only switch between them. A run with no one typing never + // uses the rich one anyway, which is why the flag is not the whole answer. + AgentOptions rich = AgentOptions.parse(new String[] {"--base-url", "http://localhost:1/v1"}); + AgentOptions plain = AgentOptions.parse(new String[] {"--base-url", "http://localhost:1/v1", "--plain"}); + + assertThat(rich.isPlain(), is(false)); + assertThat(plain.isPlain(), is(true)); + assertThat(LocalAgent.usesFullTerminal(rich, true), is(true)); + assertThat("asked for plain, so not even with a terminal", LocalAgent.usesFullTerminal(plain, true), is(false)); + assertThat( + "nobody typing, so there is nothing to pin either way", + LocalAgent.usesFullTerminal(rich, false), + is(false)); + assertThat(LocalAgent.usesFullTerminal(plain, false), is(false)); + } + @Test void baseUrlModeWithDefaults() { AgentOptions options = AgentOptions.parse(new String[] {"--base-url", "http://127.0.0.1:8080/v1/"}); diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ApprovalWireTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ApprovalWireTest.java new file mode 100644 index 000000000..10df5ee0e --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ApprovalWireTest.java @@ -0,0 +1,173 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.contains; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.empty; +import static org.hamcrest.Matchers.hasSize; +import static org.hamcrest.Matchers.is; + +import com.fasterxml.jackson.databind.JsonNode; +import java.io.BufferedReader; +import java.io.ByteArrayOutputStream; +import java.io.PrintStream; +import java.io.StringReader; +import java.nio.charset.StandardCharsets; +import java.nio.file.Path; +import java.time.Duration; +import java.util.List; +import java.util.concurrent.CopyOnWriteArrayList; +import java.util.concurrent.atomic.AtomicReference; +import net.ladenthin.llama.server.OpenAiCompatServer; +import net.ladenthin.llama.server.OpenAiServerConfig; +import org.atmosphere.ai.RetryPolicy; +import org.atmosphere.ai.fs.AgentFileSystem; +import org.atmosphere.ai.fs.WorkspaceAgentFileSystem; +import org.atmosphere.ai.tool.ToolDefinition; +import org.junit.jupiter.api.Test; +import org.junit.jupiter.api.io.TempDir; + +/** + * The approval gate over the real wire: a denial must stop the tool and reach the model as a + * tool result, so it replans instead of believing the command ran. + * + *

Both halves matter and neither is ours: Atmosphere decides whether to ask + * ({@link ConsoleApprovalStrategy#policy()}), and Atmosphere turns the answer into the {@code role: + * "tool"} message. This test drives the whole path — scripted llama.cpp chunks through the real + * {@link OpenAiCompatServer}, the real tool loop, a console answer of {@code n} or {@code y} — and + * asserts what ends up on the wire. + */ +class ApprovalWireTest { + + private static final String MODEL_ID = "local-model"; + private static final String API_KEY = "k"; + private static final Duration TIMEOUT = Duration.ofSeconds(30); + + @TempDir + Path workspace; + + private final ByteArrayOutputStream console = new ByteArrayOutputStream(); + + private AgentTerminal terminal(String typed) { + return new PlainTerminal( + new PrintStream(console, true, StandardCharsets.UTF_8), + new BufferedReader(new StringReader(typed)), + Ansi.PLAIN); + } + + private AgentRunner runner(OpenAiCompatServer server, List tools, String typed) { + AtomicReference mode = new AtomicReference<>(ApprovalMode.MANUAL); + return new AgentRunner( + "http://127.0.0.1:" + server.getPort() + "/v1", + API_KEY, + MODEL_ID, + tools, + "You are a test agent.", + 0.0, + 64, + 10) + .retryPolicy(RetryPolicy.NONE) + .approval( + new ConsoleApprovalStrategy(mode, terminal(typed), true, new TurnActivity()), + ConsoleApprovalStrategy.policy()); + } + + private ConsoleSession session() { + AgentFileSystem fs = new WorkspaceAgentFileSystem(workspace, AgentFileSystem.Limits.defaults()); + return new ConsoleSession( + new PlainTerminal(new PrintStream(console, true, StandardCharsets.UTF_8), null, Ansi.PLAIN), fs); + } + + private static OpenAiServerConfig config() { + return OpenAiServerConfig.builder() + .host("127.0.0.1") + .port(0) + .apiKey(API_KEY) + .modelId(MODEL_ID) + .build(); + } + + private static ToolDefinition shellLike(String name, List invocations) { + return ToolDefinition.builder(name, "Test tool " + name) + .parameter("command", "The command line", "string", true) + .executor(args -> { + invocations.add(String.valueOf(args.get("command"))); + return "ran"; + }) + .build(); + } + + private static ScriptedBackend backend(String toolName) { + return new ScriptedBackend((call, request) -> call == 1 + ? ScriptedBackend.toolCallTurn("call_1", toolName, "{\"command\":\"rm -rf build\"}") + : ScriptedBackend.textTurn("Understood.")); + } + + private static String toolResult(JsonNode request) { + for (JsonNode message : request.path("messages")) { + if ("tool".equals(message.path("role").asText())) { + return message.path("content").asText(); + } + } + return ""; + } + + @Test + void aDeniedShellCallNeverRunsAndTheModelIsToldItWasCancelled() throws Exception { + List invocations = new CopyOnWriteArrayList<>(); + ScriptedBackend backend = backend(ShellTool.TOOL_NAME); + try (OpenAiCompatServer server = new OpenAiCompatServer(backend, config()).start()) { + AgentRunner runner = + runner(server, List.of(shellLike(ShellTool.TOOL_NAME, invocations)), "n" + System.lineSeparator()); + ConsoleSession session = session(); + + runner.run("Delete the build directory.", List.of(), session); + + assertThat(session.await(TIMEOUT), is(true)); + assertThat("the executor must not have run", invocations, is(empty())); + List requests = backend.requests(); + assertThat(requests, hasSize(2)); + // Atmosphere's own wording; the point is that the model is told, not what it says. + assertThat(toolResult(requests.get(1)), containsString("cancelled")); + assertThat(console.toString(StandardCharsets.UTF_8), containsString("rm -rf build")); + } + } + + @Test + void anApprovedShellCallRunsAndItsRealOutputIsSentBack() throws Exception { + List invocations = new CopyOnWriteArrayList<>(); + ScriptedBackend backend = backend(ShellTool.TOOL_NAME); + try (OpenAiCompatServer server = new OpenAiCompatServer(backend, config()).start()) { + AgentRunner runner = + runner(server, List.of(shellLike(ShellTool.TOOL_NAME, invocations)), "y" + System.lineSeparator()); + ConsoleSession session = session(); + + runner.run("Delete the build directory.", List.of(), session); + + assertThat(session.await(TIMEOUT), is(true)); + assertThat(invocations, contains("rm -rf build")); + assertThat(toolResult(backend.requests().get(1)), is("ran")); + } + } + + @Test + void aReadingToolIsNotGatedAndRunsWithoutAnyAnswer() throws Exception { + List invocations = new CopyOnWriteArrayList<>(); + ScriptedBackend backend = backend("ls"); + try (OpenAiCompatServer server = new OpenAiCompatServer(backend, config()).start()) { + // empty console input: if this tool asked, the strategy would read EOF and deny + AgentRunner runner = runner(server, List.of(shellLike("ls", invocations)), ""); + ConsoleSession session = session(); + + runner.run("List the files.", List.of(), session); + + assertThat(session.await(TIMEOUT), is(true)); + assertThat(invocations, contains("rm -rf build")); + assertThat(toolResult(backend.requests().get(1)), is("ran")); + } + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AtmosphereToolLoopIntegrationTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AtmosphereToolLoopIntegrationTest.java index 198f3bcde..b4f423a1a 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AtmosphereToolLoopIntegrationTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AtmosphereToolLoopIntegrationTest.java @@ -113,7 +113,10 @@ private AgentRunner runner(List tools) { private ConsoleSession session() { AgentFileSystem fs = new WorkspaceAgentFileSystem(workspace, AgentFileSystem.Limits.defaults()); - return new ConsoleSession(new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), fs); + return new ConsoleSession( + new PlainTerminal( + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), null, Ansi.PLAIN), + fs); } /** diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AtmosphereWireContractTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AtmosphereWireContractTest.java index 02b8153cf..cfb14b82e 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AtmosphereWireContractTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/AtmosphereWireContractTest.java @@ -74,7 +74,10 @@ private AgentRunner runner(OpenAiCompatServer server, String apiKey, List invocations, String result) { diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java new file mode 100644 index 000000000..e3264a4ab --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleApprovalStrategyTest.java @@ -0,0 +1,165 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; + +import java.io.BufferedReader; +import java.io.ByteArrayOutputStream; +import java.io.PrintStream; +import java.io.StringReader; +import java.nio.charset.StandardCharsets; +import java.time.Duration; +import java.time.Instant; +import java.util.Map; +import java.util.concurrent.atomic.AtomicReference; +import org.atmosphere.ai.approval.ApprovalStrategy.ApprovalOutcome; +import org.atmosphere.ai.approval.PendingApproval; +import org.atmosphere.ai.approval.ToolApprovalPolicy; +import org.atmosphere.ai.tool.ToolDefinition; +import org.junit.jupiter.api.Test; + +class ConsoleApprovalStrategyTest { + + @org.junit.jupiter.api.Test + void everyOfferedToolIsEitherGatedOrDeclaredReadOnly() { + // The gate is a list of names, so a tool that upstream adds or renames drops out of it and + // then runs without asking -- in manual mode, silently. This is the check that turns that into + // a red build: every tool the model is offered must be classified, one way or the other. + java.util.List offered = + new java.util.ArrayList<>(WorkspaceTools.all(new WorkspaceTools.ReadTracker()).stream() + .map(ToolDefinition::name) + .toList()); + offered.add(ShellTool.TOOL_NAME); + + for (String tool : offered) { + boolean gated = ConsoleApprovalStrategy.GATED_TOOLS.contains(tool); + boolean readOnly = ConsoleApprovalStrategy.READ_ONLY_TOOLS.contains(tool); + assertThat( + tool + " is in neither set: decide whether it has to ask before it runs", + gated || readOnly, + is(true)); + assertThat(tool + " cannot be both", gated && readOnly, is(false)); + } + assertThat( + "a set that names tools nobody offers is stale", + offered.containsAll(ConsoleApprovalStrategy.GATED_TOOLS), + is(true)); + assertThat(offered.containsAll(ConsoleApprovalStrategy.READ_ONLY_TOOLS), is(true)); + } + + private final ByteArrayOutputStream console = new ByteArrayOutputStream(); + private final TurnActivity activity = new TurnActivity(); + + private PendingApproval approval() { + return new PendingApproval( + "id-1", + ShellTool.TOOL_NAME, + Map.of("command", "rm -rf build"), + null, + "console", + Instant.now().plus(Duration.ofMinutes(5))); + } + + private ApprovalOutcome ask(AtomicReference mode, String typed) { + BufferedReader reader = typed == null ? null : new BufferedReader(new StringReader(typed)); + AgentTerminal terminal = + new PlainTerminal(new PrintStream(console, true, StandardCharsets.UTF_8), reader, Ansi.PLAIN); + // The strategy never touches the session; Atmosphere passes it only so a UI can emit events. + return new ConsoleApprovalStrategy(mode, terminal, typed != null, activity).awaitApproval(approval(), null); + } + + private String consoleText() { + return console.toString(StandardCharsets.UTF_8); + } + + @Test + void yesRunsTheToolOnceAndKeepsAsking() { + AtomicReference mode = new AtomicReference<>(ApprovalMode.MANUAL); + + assertThat(ask(mode, "y\n"), is(ApprovalOutcome.APPROVED)); + assertThat(mode.get(), is(ApprovalMode.MANUAL)); + assertThat(consoleText(), containsString(ShellTool.TOOL_NAME)); + assertThat(consoleText(), containsString("rm -rf build")); + } + + @Test + void anEmptyAnswerMeansYes() { + assertThat(ask(new AtomicReference<>(ApprovalMode.MANUAL), "\n"), is(ApprovalOutcome.APPROVED)); + } + + @Test + void noDeniesAndAtmosphereTellsTheModel() { + // The cancellation text itself is Atmosphere's ("Action cancelled by user"), which is why this + // side only has to return DENIED -- see ApprovalWireTest for the message reaching the model. + assertThat(ask(new AtomicReference<>(ApprovalMode.MANUAL), "n\n"), is(ApprovalOutcome.DENIED)); + } + + @Test + void autoApprovesAndStopsAskingForTheRestOfTheSession() { + AtomicReference mode = new AtomicReference<>(ApprovalMode.MANUAL); + + assertThat(ask(mode, "a\n"), is(ApprovalOutcome.APPROVED)); + assertThat(mode.get(), is(ApprovalMode.AUTO)); + // the next call must not read anything: an empty reader would otherwise mean "input closed" + assertThat(ask(mode, ""), is(ApprovalOutcome.APPROVED)); + } + + @Test + void anUnreadableAnswerIsAskedAgain() { + assertThat(ask(new AtomicReference<>(ApprovalMode.MANUAL), "maybe\ny\n"), is(ApprovalOutcome.APPROVED)); + assertThat(consoleText(), containsString("please answer y, n or a")); + } + + @Test + void withoutAConsoleTheAnswerIsNo() { + // One-shot mode: nobody can answer, so the safe outcome is a denial with a printed reason -- + // never a silent auto-approval, which would make an unattended run the most permissive one. + assertThat(ask(new AtomicReference<>(ApprovalMode.MANUAL), null), is(ApprovalOutcome.DENIED)); + assertThat(consoleText(), containsString("--auto")); + + assertThat(ask(new AtomicReference<>(ApprovalMode.AUTO), null), is(ApprovalOutcome.APPROVED)); + } + + @Test + void closedInputDeniesRatherThanBlocking() { + assertThat(ask(new AtomicReference<>(ApprovalMode.MANUAL), ""), is(ApprovalOutcome.DENIED)); + assertThat(consoleText(), containsString("input closed")); + } + + @Test + void writingToolsAndTheShellAreGatedReadingToolsAreNot() { + ToolApprovalPolicy policy = ConsoleApprovalStrategy.policy(); + + for (String gated : new String[] {ShellTool.TOOL_NAME, "write_file", "edit_file", "delete", "rename"}) { + assertThat(gated, policy.requiresApproval(stub(gated)), is(true)); + } + for (String free : new String[] {"ls", "read_file", "glob", "grep"}) { + assertThat(free, policy.requiresApproval(stub(free)), is(false)); + } + assertThat( + ConsoleApprovalStrategy.gated(java.util.List.of("ls", "write_file", ShellTool.TOOL_NAME)), + is(java.util.List.of("write_file", ShellTool.TOOL_NAME))); + } + + @Test + void theSpinnerIsPausedWhileTheQuestionIsOpen() { + // The turn runs on its own thread while the console thread redraws the status block four times + // a second. A redraw arriving in the middle of a raw-mode key read writes escape sequences + // across the question, so the prompt owns the terminal until it has an answer. + assertThat(activity.isPaused(), is(false)); + + ask(new AtomicReference<>(ApprovalMode.MANUAL), "y" + System.lineSeparator()); + + assertThat("and hands it back afterwards", activity.isPaused(), is(false)); + assertThat(consoleText(), containsString("allow?")); + } + + private static ToolDefinition stub(String name) { + return ToolDefinition.builder(name, "test").executor(args -> "").build(); + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java new file mode 100644 index 000000000..c91f08bb3 --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleFormattingTest.java @@ -0,0 +1,191 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.not; + +import java.util.Map; +import org.junit.jupiter.api.Test; + +/** The console cosmetics: colour decision, status line, and the streaming Markdown renderer. */ +class ConsoleFormattingTest { + + // ----- Ansi ----- + + private static Ansi detect(Map env, boolean terminal) { + return Ansi.detect(env::get, () -> terminal); + } + + @Test + void colourIsOnOnlyOnATerminal() { + assertThat(detect(Map.of(), true).isEnabled(), is(true)); + assertThat(detect(Map.of(), false).isEnabled(), is(false)); + } + + @Test + void theEnvironmentCanForceColourOnOrOff() { + // NO_COLOR: "when present and not an empty string ... prevents the addition of ANSI color" + assertThat(detect(Map.of("NO_COLOR", "1"), true).isEnabled(), is(false)); + assertThat(detect(Map.of("NO_COLOR", ""), true).isEnabled(), is(true)); + assertThat(detect(Map.of("TERM", "dumb"), true).isEnabled(), is(false)); + assertThat(detect(Map.of("CLICOLOR", "0"), true).isEnabled(), is(false)); + // forcing wins over "not a terminal", e.g. when piping into a pager + assertThat(detect(Map.of("CLICOLOR_FORCE", "1"), false).isEnabled(), is(true)); + // ... but NO_COLOR is checked after CLICOLOR_FORCE, which is the documented precedence + assertThat(detect(Map.of("CLICOLOR_FORCE", "1", "NO_COLOR", "1"), false).isEnabled(), is(true)); + } + + @Test + void aPlainInstanceReturnsTheTextUnchanged() { + assertThat(Ansi.PLAIN.bold("x") + Ansi.PLAIN.dim("y") + Ansi.PLAIN.red("z"), is("xyz")); + assertThat(detect(Map.of(), true).bold("x"), containsString("\u001b[")); + } + + // ----- StatusLine ----- + + @Test + void theStatusLineShowsWorkspaceModeContextToolsAndModel() { + String line = StatusLine.render( + java.nio.file.Path.of("/tmp/ws"), ApprovalMode.MANUAL, 1234, false, 16384, 9, "local-model", false); + + // an icon, a space, its value -- the same shape for every part of the line + assertThat(line, containsString("📁 ")); + assertThat(line, containsString("ws · ⏸ manual · 📊 1.2k/16k · 🔧 9 · 🤖 local-model]")); + } + + @Test + void shiftTabCyclesTheModeAndAPlainStreamDeclinesTheShortcut() { + // the key can only be seen by a console that owns the keyboard; everything else keeps /mode + assertThat(ApprovalMode.MANUAL.next(), is(ApprovalMode.AUTO)); + assertThat(ApprovalMode.AUTO.next(), is(ApprovalMode.MANUAL)); + assertThat( + "a cycle, so it always returns to where it started", + ApprovalMode.MANUAL.next().next(), + is(ApprovalMode.MANUAL)); + + AgentTerminal plain = + new PlainTerminal(new java.io.PrintStream(new java.io.ByteArrayOutputStream()), null, Ansi.PLAIN); + assertThat( + plain.onCycleMode(() -> { + throw new AssertionError("must not run"); + }), + is(false)); + } + + @Test + void eachModeCarriesItsOwnSymbol() { + // the glyph is what makes the mode findable at a glance; the word stays next to it + assertThat(ApprovalMode.MANUAL.badge(), is("⏸ manual")); + assertThat(ApprovalMode.AUTO.badge(), is("⏵⏵ auto")); + assertThat( + StatusLine.render(java.nio.file.Path.of("/tmp/ws"), ApprovalMode.AUTO, 0, true, 0, 1, "m", false), + containsString("⏵⏵ auto")); + } + + @Test + void aRemoteEndpointIsMarkedDifferentlyFromAModelLoadedHere() { + java.nio.file.Path ws = java.nio.file.Path.of("/tmp/ws"); + assertThat(StatusLine.render(ws, ApprovalMode.MANUAL, 0, true, 0, 1, "m", false), containsString("🤖 m")); + assertThat(StatusLine.render(ws, ApprovalMode.MANUAL, 0, true, 0, 1, "m", true), containsString("🌐 m")); + } + + @Test + void aLongWorkspacePathIsShortenedToItsLastTwoSegments() { + // the path is on every line of the session, so it must not push the rest off the screen + java.nio.file.Path deep = java.nio.file.Path.of("/home/someone/projects/customer/service/backend/module"); + assertThat(StatusLine.shorten(deep), containsString("backend")); + assertThat(StatusLine.shorten(deep), containsString("module")); + assertThat(StatusLine.shorten(deep).startsWith("…"), is(true)); + assertThat( + StatusLine.shorten(java.nio.file.Path.of("/tmp/ws")), + is(java.nio.file.Path.of("/tmp/ws").toString())); + } + + @Test + void contextIsShownInThousandsAndWithoutASizeWhenItIsUnknown() { + assertThat(StatusLine.context(812, false, 32768), is("812/33k")); + assertThat(StatusLine.context(16000, false, 32768), is("16k/33k")); + assertThat(StatusLine.context(0, false, StatusLine.UNKNOWN_CONTEXT), is("0")); + assertThat(StatusLine.context(2500, false, StatusLine.UNKNOWN_CONTEXT), is("2.5k")); + } + + @Test + void anEstimatedCountIsMarkedWithATilde() { + // llama.cpp reports usage only to clients that ask for it, and Atmosphere does not, so the + // number normally comes from LocalAgent.estimateTokens -- the tilde says so. + assertThat(StatusLine.context(2500, true, 16384), is("~2.5k/16k")); + assertThat( + LocalAgent.estimateTokens( + "0123456789", java.util.List.of(org.atmosphere.ai.llm.ChatMessage.user("0123456789"))), + is(5L)); + } + + // ----- MarkdownConsole ----- + + private static String render(String text, Ansi ansi) { + StringBuilder buffer = new StringBuilder(); + MarkdownConsole console = + new MarkdownConsole(line -> buffer.append(line).append(System.lineSeparator()), ansi); + // one character at a time: the renderer must not depend on where the stream splits + for (int i = 0; i < text.length(); i++) { + console.append(text.substring(i, i + 1)); + } + console.flush(); + return buffer.toString(); + } + + @Test + void withoutColourTheTextIsPassedThroughUnchanged() { + String markdown = "# Title\n\nSome **bold** and `code`.\n- one\n- two\n"; + + assertThat( + render(markdown, Ansi.PLAIN), + is("Title\n\nSome bold and code.\n• one\n• two\n".replace("\n", System.lineSeparator()))); + } + + @Test + void headingsBulletsAndInlineSpansAreStyled() { + Ansi ansi = detect(Map.of(), true); + String out = render("## Heading\n- item with **bold**\ntext with `code` inside\n", ansi); + + assertThat(out, containsString(ansi.bold("Heading"))); + assertThat(out, containsString(ansi.cyan("•"))); + assertThat(out, containsString(ansi.bold("bold"))); + assertThat(out, containsString(ansi.cyan("code"))); + // the markers themselves are gone, that is the point of rendering + assertThat(out, not(containsString("**"))); + assertThat(out, not(containsString("##"))); + } + + @Test + void aFencedBlockIsStyledAsAWholeAndNotParsedInside() { + Ansi ansi = detect(Map.of(), true); + String out = render("```java\nint a = b * c; // **not bold**\n```\n", ansi); + + assertThat(out, containsString(ansi.cyan("int a = b * c; // **not bold**"))); + } + + @Test + void anUnclosedMarkerIsLeftAsTyped() { + // A half-streamed "**bold" must never swallow the rest of the line. + assertThat(render("a **b\n", Ansi.PLAIN), is("a **b" + System.lineSeparator())); + assertThat(render("2 * 3 * 4\n", Ansi.PLAIN), is("2 * 3 * 4" + System.lineSeparator())); + } + + // ----- the pinned block ----- + + @Test + void aStatusRowIsCutToTheWindowWidthBecauseAWrappedRowBreaksTheBlock() { + // A wrapped row takes two screen lines while the reserved region is sized in lines, so + // everything below it lands in the wrong place -- a long /compact summary tore the block apart. + assertThat(JLineTerminal.fit("short", 20), is("short")); + assertThat(JLineTerminal.fit("0123456789", 10), is("0123456789")); + assertThat(JLineTerminal.fit("0123456789x", 10), is("012345678…")); + assertThat(JLineTerminal.fit("0123456789x", 10).length(), is(10)); + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleSessionTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleSessionTest.java new file mode 100644 index 000000000..c687dc5e2 --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ConsoleSessionTest.java @@ -0,0 +1,210 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.not; + +import java.nio.file.Path; +import java.util.ArrayList; +import java.util.List; +import java.util.Map; +import org.atmosphere.ai.AiEvent; +import org.atmosphere.ai.fs.AgentFileSystem; +import org.atmosphere.ai.fs.WorkspaceAgentFileSystem; +import org.junit.jupiter.api.Test; +import org.junit.jupiter.api.io.TempDir; + +/** + * What the console is handed for one turn, with a terminal that only records. + * + *

The rule every test here defends: one call to the terminal is one screen line. The pinned + * block at the bottom is sized in lines, so a single "line" carrying ten newlines pushes the screen + * ten rows further than the terminal accounted for and the block is drawn across the output — which is + * what a {@code write_file} call with a whole file in its arguments did. + */ +class ConsoleSessionTest { + + @TempDir + Path workspace; + + /** A terminal that records what it was told to print. */ + private static final class RecordingTerminal implements AgentTerminal { + private final List lines = new ArrayList<>(); + private volatile boolean pending; + + @Override + public boolean hasPendingInput() { + return pending; + } + + @Override + public void line(String text) { + lines.add(text); + } + + @Override + public String readLine(String prompt) { + return null; + } + + @Override + public String readKey(String prompt) { + return null; + } + + @Override + public void status(List statusLines) {} + + @Override + public boolean pinsStatus() { + return false; + } + + @Override + public Ansi ansi() { + return Ansi.PLAIN; + } + + @Override + public void close() {} + } + + private final RecordingTerminal terminal = new RecordingTerminal(); + + private ConsoleSession session() { + AgentFileSystem fs = new WorkspaceAgentFileSystem(workspace, AgentFileSystem.Limits.defaults()); + return new ConsoleSession(terminal, fs); + } + + private void assertEveryLineIsOneLine() { + for (String line : terminal.lines) { + assertThat("a printed line must not contain a newline: " + line, line.contains("\n"), is(false)); + assertThat(line.contains("\r"), is(false)); + } + } + + @Test + void aToolCallWithAWholeFileInItsArgumentsStaysOnOneLine() { + String fileContent = "# Project Summary\n\n## Build\n\nmvn package\n".repeat(20); + ConsoleSession session = session(); + + session.emit(new AiEvent.ToolStart("write_file", Map.of("path", "project_summary.md", "content", fileContent))); + + assertEveryLineIsOneLine(); + assertThat(terminal.lines, is(not(List.of()))); + assertThat(terminal.lines.get(0), containsString("write_file")); + assertThat(terminal.lines.get(0), containsString("project_summary.md")); + assertThat("and it is cut, not merely joined", terminal.lines.get(0).length() < 400, is(true)); + } + + @Test + void aMultiLineToolResultStaysOnOneLineToo() { + ConsoleSession session = session(); + + session.emit(new AiEvent.ToolStart("run_command", Map.of("command", "mvn -version"))); + session.emit(new AiEvent.ToolResult("run_command", "exit code: 0\nline one\r\nline two\n")); + session.emit(new AiEvent.ToolError("run_command", "boom\nand more")); + + assertEveryLineIsOneLine(); + } + + /** A cancellable turn that only records whether it was stopped. */ + private static final class RecordingHandle implements org.atmosphere.ai.ExecutionHandle { + private final java.util.concurrent.CompletableFuture done = + new java.util.concurrent.CompletableFuture<>(); + private volatile boolean cancelled; + + @Override + public void cancel() { + cancelled = true; + done.complete(null); + } + + @Override + public boolean isDone() { + return done.isDone(); + } + + @Override + public java.util.concurrent.CompletableFuture whenDone() { + return done; + } + } + + @Test + void typingWhileTheAgentWorksStopsTheTurnAndTheLineIsNotConsumed() throws InterruptedException { + // the whole point of a prompt that is there during a turn: a request that has been overtaken + // must not keep running, and what was typed stays queued to become the next message + ConsoleSession session = session(); // never completed: only the typing can end this wait + RecordingHandle handle = new RecordingHandle(); + terminal.pending = true; + + LocalAgent.TurnEnd end = + LocalAgent.awaitWithActivity(session, terminal, () -> "state", new TurnActivity(), handle); + + assertThat(end, is(LocalAgent.TurnEnd.INTERRUPTED)); + assertThat("the stream the model is answering on is closed", handle.cancelled, is(true)); + assertThat( + "and it says so rather than looking like a finished answer", + terminal.lines.stream().anyMatch(l -> l.contains("interrupted")), + is(true)); + } + + @Test + void aTurnThatFinishesOnItsOwnIsNotCancelled() throws InterruptedException { + ConsoleSession session = session(); + RecordingHandle handle = new RecordingHandle(); + session.complete(); + + LocalAgent.TurnEnd end = + LocalAgent.awaitWithActivity(session, terminal, () -> "state", new TurnActivity(), handle); + + assertThat(end, is(LocalAgent.TurnEnd.FINISHED)); + assertThat(handle.cancelled, is(false)); + } + + @Test + void theContextFigureGrowsWhileTheTurnRuns() { + // it used to be rendered once before the turn and handed over as a fixed string, so it stood + // still through every tool round and only moved at the next prompt + ConsoleSession session = session(); + assertThat(LocalAgent.liveTokens(1000, session), is(1000L)); + + session.send("a".repeat(400)); + session.emit(new AiEvent.ToolStart("read_file", Map.of("file_path", "x"))); + session.emit(new AiEvent.ToolResult("read_file", "b".repeat(4000))); + + assertThat( + "the tool output counts too, it is in the next call's prompt", + LocalAgent.liveTokens(1000, session) > 2000L, + is(true)); + } + + @Test + void aReportedCountWinsOverTheEstimate() { + ConsoleSession session = session(); + session.send("a".repeat(4000)); + session.usage(new org.atmosphere.ai.TokenUsage(7777, 10, 0, 7787, "m")); + + assertThat(LocalAgent.liveTokens(1000, session), is(7777L)); + } + + @Test + void theRecordedRoundKeepsTheFullArgumentsEvenThoughTheConsoleShowsLess() { + String fileContent = "a".repeat(5000); + ConsoleSession session = session(); + + session.emit(new AiEvent.ToolStart("write_file", Map.of("content", fileContent))); + + assertThat("the console is cut …", terminal.lines.get(0).length() < 400, is(true)); + assertThat( + "… but what the model is told later is not cut here", + session.rounds().get(0).argumentsJson().length() > 4000, + is(true)); + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/InterruptedTurnTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/InterruptedTurnTest.java new file mode 100644 index 000000000..98fb4b0d3 --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/InterruptedTurnTest.java @@ -0,0 +1,211 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.hasSize; +import static org.hamcrest.Matchers.is; + +import com.fasterxml.jackson.databind.JsonNode; +import java.nio.file.Path; +import java.util.ArrayList; +import java.util.List; +import net.ladenthin.llama.server.OpenAiCompatServer; +import net.ladenthin.llama.server.OpenAiServerConfig; +import org.atmosphere.ai.RetryPolicy; +import org.atmosphere.ai.fs.AgentFileSystem; +import org.atmosphere.ai.fs.WorkspaceAgentFileSystem; +import org.atmosphere.ai.llm.ChatMessage; +import org.junit.jupiter.api.Test; +import org.junit.jupiter.api.io.TempDir; + +/** + * Typing while the agent works stops that turn — and the turn after it has to run. + * + *

Reported as "it does not carry on by itself": the interruption printed its line and then nothing + * followed. The sequence is driven here end to end against the real {@link OpenAiCompatServer} with + * scripted llama.cpp chunks, because every part of it is a different piece of machinery — the + * cancellable entry point of the runtime, the queue the console keeps, and the history the next turn + * is sent with — and a defect in any of them looks identical from the outside. + */ +class InterruptedTurnTest { + + private static final String MODEL_ID = "local-model"; + + /** Longer than one activity tick, so the wait actually looks for typed input. */ + private static final java.time.Duration SLOW_ENOUGH_TO_INTERRUPT = java.time.Duration.ofMillis(700); + + @TempDir + Path workspace; + + /** A console that answers "yes, something was typed" whenever the test says so. */ + private final class Typing implements AgentTerminal { + private volatile boolean pending; + private final List lines = new ArrayList<>(); + + @Override + public void line(String text) { + lines.add(text); + } + + @Override + public String readLine(String prompt) { + return null; + } + + @Override + public String readKey(String prompt) { + return null; + } + + @Override + public boolean hasPendingInput() { + return pending; + } + + @Override + public void status(List statusLines) {} + + @Override + public boolean pinsStatus() { + return false; + } + + @Override + public Ansi ansi() { + return Ansi.PLAIN; + } + + @Override + public void close() {} + } + + private final Typing terminal = new Typing(); + + private OpenAiCompatServer server(ScriptedBackend backend) throws Exception { + return new OpenAiCompatServer( + backend, + OpenAiServerConfig.builder() + .host("127.0.0.1") + .port(0) + .modelId(MODEL_ID) + .build()) + .start(); + } + + private AgentRunner runner(OpenAiCompatServer server) { + return new AgentRunner( + "http://127.0.0.1:" + server.getPort() + "/v1", + "k", + MODEL_ID, + List.of(), + "You are a test agent.", + 0.0, + 64, + 4) + .retryPolicy(RetryPolicy.NONE); + } + + /** + * A turn that takes long enough to be interrupted. + * + *

Without this the test proves nothing: the interruption is only looked for while waiting for + * the turn, and a scripted answer arrives before the first look. That is also the behaviour in a + * real session — a turn that is already finished is not cut short — so the delay is what makes + * this the reported situation rather than a different one. + * + * @param call which model call this is + * @return the scripted answer, after a pause on the first call + */ + private static List slowFirstTurn(int call) { + if (call == 1) { + try { + Thread.sleep(SLOW_ENOUGH_TO_INTERRUPT.toMillis()); + } catch (InterruptedException e) { + Thread.currentThread().interrupt(); + } + } + return ScriptedBackend.textTurn("answer " + call); + } + + @Test + void theTurnAfterAnInterruptedOneRunsAndCarriesTheHistory() throws Exception { + ScriptedBackend backend = new ScriptedBackend((call, request) -> slowFirstTurn(call)); + try (OpenAiCompatServer server = server(backend)) { + AgentRunner runner = runner(server); + AgentFileSystem files = new WorkspaceAgentFileSystem(workspace, AgentFileSystem.Limits.defaults()); + List history = new ArrayList<>(); + ToolCallLog log = new ToolCallLog(); + + // The user types while the first turn is still running. + terminal.pending = true; + ConsoleSession first = LocalAgent.turn( + runner, files, "a poem please", history, terminal, log, 1, ignored -> "", new TurnActivity()); + + assertThat( + "the interruption is announced", + terminal.lines.stream().anyMatch(line -> line.contains("interrupted")), + is(true)); + + // What was typed is now the next message, and nothing is pending any more. + terminal.pending = false; + ConsoleSession second = LocalAgent.turn( + runner, files, "make it longer", history, terminal, log, 2, ignored -> "", new TurnActivity()); + + assertThat("the turn after the interruption produced an answer", second.text(), containsString("answer")); + assertThat( + "and it was not itself reported as failed", + second.failure(), + is(org.hamcrest.Matchers.nullValue())); + assertThat( + "the interrupted turn is in the history as asked", + history.get(0).content(), + is("a poem please")); + + List requests = backend.requests(); + assertThat("both turns reached the server", requests.size() >= 2, is(true)); + JsonNode last = requests.get(requests.size() - 1).path("messages"); + assertThat( + "the second request carries the typed line", + last.get(last.size() - 1).path("content").asText(), + is("make it longer")); + assertThat("and the interrupted question before it", last.toString(), containsString("a poem please")); + assertThat(first.rounds(), hasSize(0)); + } + } + + @Test + void aSecondLineTypedDuringTheReplacementTurnStopsThatOneToo() throws Exception { + // Not a defect: each typed line overtakes the turn it arrived in. It is pinned because the + // symptom -- an interruption that seems to lead nowhere -- looks the same as a turn that never + // starts, and telling the two apart afterwards is what took the longest. + ScriptedBackend backend = new ScriptedBackend((call, request) -> { + try { + Thread.sleep(SLOW_ENOUGH_TO_INTERRUPT.toMillis()); + } catch (InterruptedException e) { + Thread.currentThread().interrupt(); + } + return ScriptedBackend.textTurn("answer " + call); + }); + try (OpenAiCompatServer server = server(backend)) { + AgentRunner runner = runner(server); + AgentFileSystem files = new WorkspaceAgentFileSystem(workspace, AgentFileSystem.Limits.defaults()); + List history = new ArrayList<>(); + ToolCallLog log = new ToolCallLog(); + + terminal.pending = true; + LocalAgent.turn(runner, files, "one", history, terminal, log, 1, ignored -> "", new TurnActivity()); + LocalAgent.turn(runner, files, "two", history, terminal, log, 2, ignored -> "", new TurnActivity()); + + assertThat( + "both were cut short, and both said so", + terminal.lines.stream() + .filter(line -> line.contains("interrupted")) + .count(), + is(2L)); + } + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java new file mode 100644 index 000000000..be58cc79c --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/JLineTerminalTest.java @@ -0,0 +1,323 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.notNullValue; +import static org.hamcrest.Matchers.nullValue; + +import java.io.ByteArrayInputStream; +import java.io.ByteArrayOutputStream; +import java.nio.charset.StandardCharsets; +import java.util.List; +import org.jline.terminal.Size; +import org.jline.terminal.Terminal; +import org.jline.terminal.TerminalBuilder; +import org.junit.jupiter.api.Test; + +/** + * What the screen would look like, asserted on the bytes the terminal emits. + * + *

A JLine terminal built over two streams renders exactly like one on a TTY — same escape + * sequences, same line reader — so a drawing bug is visible in the output without a console. That is + * worth having, because this class is the one place where a mistake is invisible to every other test + * and obvious to whoever is using the agent. + * + *

The bug these tests were written for: the rule above the input used to be the first line of a + * two-line prompt, while {@code ERASE_LINE_ON_FINISH} erases exactly one line — so every Enter + * left a rule behind, and holding Enter drew a column of them. Drawing it as ordinary output instead + * only moved the problem: then one rule stayed in the scrollback per turn and travelled up with it. + * There is now exactly one rule, the first line of the pinned block, and these tests hold the console + * to writing none at all. + */ +class JLineTerminalTest { + + /** As many columns as a narrow window, so a wrapped line would be obvious. */ + private static final Size SIZE = new Size(60, 10); + + /** What a cleared screen looks like on the wire: erase the whole display. */ + private static final String ERASE_DISPLAY = "\u001b[2J"; + + private final ByteArrayOutputStream emitted = new ByteArrayOutputStream(); + + private Terminal terminal(String keystrokes) throws Exception { + return terminal(new ByteArrayInputStream(keystrokes.getBytes(StandardCharsets.UTF_8))); + } + + private Terminal terminal(java.io.InputStream keystrokes) throws Exception { + return TerminalBuilder.builder() + .streams(keystrokes, emitted) + .type("xterm-256color") + // The rule is drawn with U+2500. Both are needed: the writer encodes through the + // stdout charset, which is not the one .encoding() sets, and without it every rule + // arrives as a row of "?" and the assertions compare against bytes nobody wrote. + .encoding(StandardCharsets.UTF_8) + .stdoutEncoding(StandardCharsets.UTF_8) + .size(SIZE) + .provider("exec") + .build(); + } + + private String screen() { + return emitted.toString(StandardCharsets.UTF_8); + } + + /** + * How many times a full-width rule was written, in either of the two forms JLine uses. + * + *

Counting only one of them is how the first version of this test passed against the very bug + * it was written for: a box character goes out as UTF-8 {@code U+2500} from {@code printAbove}, + * but inside a prompt JLine switches to the DEC line-drawing character set and sends + * {@code ESC(0} + a row of {@code q} + {@code ESC(B}. A rule carried in the prompt is therefore + * invisible to a search for {@code ─}. + * + * @return how many rules were written + */ + private int rules() { + int width = SIZE.getColumns() - 1; + return occurrences("─".repeat(width)) + occurrences("\u001b(0" + "q".repeat(width)); + } + + private int occurrences(String needle) { + int count = 0; + for (int at = screen().indexOf(needle); at >= 0; at = screen().indexOf(needle, at + 1)) { + count++; + } + return count; + } + + @Test + void pressingEnterSeveralTimesLeavesNoRuleInTheScrollback() throws Exception { + try (Terminal terminal = terminal("\n\n\n\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + for (int i = 0; i < 4; i++) { + assertThat("an empty line is still a line", console.readLine("ignored"), is("")); + } + + assertThat("the only rule is the pinned one, and it is never written as output", rules(), is(0)); + } + } + + @Test + void whatWasTypedSurvivesAboveTheInputLine() throws Exception { + try (Terminal terminal = terminal("hello\nworld\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + assertThat(console.readLine("ignored"), is("hello")); + assertThat(console.readLine("ignored"), is("world")); + + // the input line itself is erased on Enter, so the echo is what keeps the transcript + assertThat(screen(), containsString("hello")); + assertThat(screen(), containsString("world")); + assertThat("and still no rule travels along with it", rules(), is(0)); + } + } + + @Test + void anEmptyLineIsNotAPendingRequest() throws Exception { + try (Terminal terminal = terminal("\n \nreal\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + console.readLine("ignored"); // starts the reader; the rest queues up behind it + + waitFor(console::hasPendingInput); + // Two blank lines are queued in front of it, and it is still the real one that counts: + // otherwise holding Enter cancels one turn per keystroke and produces nothing. + assertThat(console.hasPendingInput(), is(true)); + assertThat(console.readLine("ignored").isBlank(), is(true)); + assertThat(console.readLine("ignored"), is("real")); + // End of input is pending too, and has to be: a session whose console has closed must + // stop waiting rather than keep a turn running for nobody. + assertThat(console.hasPendingInput(), is(true)); + assertThat("and it reads as no line at all", console.readLine("ignored"), is(nullValue())); + } + } + + /** + * Wait for the reader thread to have caught up. + * + * @param condition what to wait for + * @throws InterruptedException if interrupted while waiting + */ + private void waitFor(java.util.function.BooleanSupplier condition) throws InterruptedException { + for (int attempt = 0; attempt < 200 && !condition.getAsBoolean(); attempt++) { + Thread.sleep(10); + } + } + + @Test + void theScreenIsScrolledSoTheInputStartsAtTheBottom() throws Exception { + // The reader draws its prompt where the cursor is, and only the block below is pinned to the + // window. Without this the input floats after the output with the block far below it, and the + // two only meet once enough output has scrolled the cursor down on its own. + try (Terminal terminal = terminal("\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + console.readLine("ignored"); + + long blankLines = + screen().chars().filter(character -> character == '\n').count(); + assertThat( + "the cursor is pushed to the last row before the first prompt", + blankLines >= SIZE.getRows() - 1, + is(true)); + } + } + + @Test + void clearingTheScreenWipesItAndLeavesTheReaderWorking() throws Exception { + try (Terminal terminal = terminal("first\nsecond\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + assertThat(console.readLine("ignored"), is("first")); + int before = screen().length(); + + console.clearScreen(); + + // The capability is terminfo source ("\E[H\E[2J"), so what must reach the screen is the + // expanded form. Writing the capability as it comes prints it as text, which is what this + // assertion caught the first time it ran. + assertThat( + terminal.getStringCapability(org.jline.utils.InfoCmp.Capability.clear_screen), is(notNullValue())); + assertThat("erase display reached the screen", screen().substring(before), containsString(ERASE_DISPLAY)); + assertThat("and the prompt still reads afterwards", console.readLine("ignored"), is("second")); + } + } + + @Test + void makingTheWindowNarrowerRedrawsTheBlockAtTheNewWidth() throws Exception { + // This pins an assumption about JLine rather than logic of ours: it re-cuts the rows it holds + // when the window shrinks, so nothing here has to. That is worth a test because the whole + // bottom block depends on it -- a row wider than the window wraps onto a second screen line, + // and the reserved region cannot survive that. An upgrade that changed it would show up here + // instead of on somebody's screen. + try (Terminal terminal = terminal("\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + console.status(List.of("state")); + int before = screen().length(); + + terminal.setSize(new Size(30, 10)); + console.status(List.of("state")); + + String afterResize = screen().substring(before); + assertThat("the block was drawn again", afterResize.isEmpty(), is(false)); + assertThat( + "and never again at the width of the window that is gone", + afterResize.contains("─".repeat(SIZE.getColumns() - 1)), + is(false)); + } + } + + @Test + void aRowTooWideForTheNewWindowIsCutRatherThanWrapped() throws Exception { + try (Terminal terminal = terminal("\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + console.status(List.of("x".repeat(50))); + + terminal.setSize(new Size(20, 10)); + console.status(List.of("x".repeat(50))); + + assertThat( + "the row that was rendered for the wide window is not reused", occurrences("x".repeat(50)), is(1)); + assertThat("it is cut for the narrow one", screen(), containsString("…")); + } + } + + @Test + void everyResizeDrawsExactlyOnePrompt() throws Exception { + // The reported artefact is a second, stale "> " left on screen after dragging the window. + // This drives the path that redraws it -- a real size change plus the signal, with the reader + // sitting in readLine as it does all session -- and pins that shrinking, growing and changing + // the row count each produce one prompt and not two. It holds for every size tried, which is + // what says the remaining artefact is not in this path. + java.io.PipedOutputStream keys = new java.io.PipedOutputStream(); + try (Terminal terminal = terminal(new java.io.PipedInputStream(keys)); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + Thread reader = new Thread(() -> console.readLine("ignored")); + reader.setDaemon(true); + reader.start(); + Thread.sleep(200); + console.status(List.of("state row")); + + int[][] sizes = {{30, 10}, {90, 10}, {45, 10}, {120, 24}}; + for (int[] size : sizes) { + int before = screen().length(); + + terminal.setSize(new Size(size[0], size[1])); + terminal.raise(Terminal.Signal.WINCH); + Thread.sleep(150); + + String drawn = screen().substring(before); + long prompts = + drawn.chars().filter(character -> character == '>').count(); + assertThat("one prompt after resizing to " + size[0] + "x" + size[1], prompts, is(1L)); + } + } + } + + @Test + void theBlockIsBackOnScreenAfterAClear() throws Exception { + // Clearing erases the block along with everything else, and the pinned region is redrawn only + // when its content changes -- so after a clear it believes it is still on screen and draws + // nothing, leaving the bottom of the window empty. + try (Terminal terminal = terminal("go\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + // As in a session: the reader owns the screen before anything is drawn into the block. The + // pause is not decoration -- the reader thread starts the next prompt as soon as one + // returns, and clearing while that is in flight makes what JLine emits depend on which of + // the two got there first. This test is about the block coming back, not about that race. + console.readLine("ignored"); + Thread.sleep(200); + console.status(List.of("a distinctive state row")); + int before = screen().length(); + + console.clearScreen(); + + assertThat( + "the block is drawn again after the wipe", + screen().substring(before), + containsString("a distinctive state row")); + } + } + + @Test + void controlLIsBoundToTheReadersOwnClearScreen() throws Exception { + // 0x0C is Ctrl-L. It is bound by JLine itself, so /cls is the second way to do this rather + // than the only one -- worth pinning, because a keymap option could silently take it away. + try (Terminal terminal = terminal("\u000cstill here\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + assertThat(console.readLine("ignored"), is("still here")); + assertThat("Ctrl-L cleared rather than being typed into the line", screen(), containsString(ERASE_DISPLAY)); + } + } + + @Test + void aRowIsCutByScreenColumnsNotByCharacters() { + // An icon takes two columns and one character. Cutting by character length lets the row come + // out wider than the window, wrap onto a second screen line, and push everything below the + // reserved region out of place -- the tearing that a long summary caused, through another door. + String icons = "📁".repeat(20); + + String cut = JLineTerminal.fit(icons, 10); + + // What matters is the width on screen, not how many characters that took. + assertThat("the row fits the window", new org.jline.utils.AttributedString(cut).columnLength() <= 10, is(true)); + assertThat( + "a cut by characters would have kept nine icons, which is eighteen columns", + cut.codePointCount(0, cut.length()) < 9, + is(true)); + assertThat(cut.endsWith("…"), is(true)); + assertThat("plain text is untouched when it fits", JLineTerminal.fit("short", 10), is("short")); + } + + @Test + void aMultiLineStringIsStillPrintedAsSeveralLines() throws Exception { + try (Terminal terminal = terminal("\n"); + JLineTerminal console = JLineTerminal.over(terminal, List.of())) { + console.line("first" + System.lineSeparator() + "second"); + + assertThat(screen(), containsString("first")); + assertThat(screen(), containsString("second")); + } + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java index dd1566307..b6ff7c634 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/LocalAgentTest.java @@ -5,6 +5,7 @@ package net.ladenthin.llama.atmosphere; import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.contains; import static org.hamcrest.Matchers.containsString; import static org.hamcrest.Matchers.hasItem; import static org.hamcrest.Matchers.hasSize; @@ -19,6 +20,7 @@ import java.nio.charset.StandardCharsets; import java.nio.file.Files; import java.nio.file.Path; +import java.util.ArrayList; import java.util.List; import net.ladenthin.llama.server.OpenAiCompatServer; import net.ladenthin.llama.server.OpenAiServerConfig; @@ -34,6 +36,8 @@ */ class LocalAgentTest { + private static final String NL = System.lineSeparator(); + @TempDir Path workspace; @@ -53,7 +57,7 @@ private static OpenAiCompatServer server(ScriptedBackend backend) throws Excepti void oneShotTurnReadsAWorkspaceFileThroughTheBuiltInFileTools() throws Exception { Files.writeString(workspace.resolve("hello.txt"), "VALUE=42\n"); ScriptedBackend backend = new ScriptedBackend((call, request) -> call == 1 - ? ScriptedBackend.toolCallTurn("call_1", "read_file", "{\"path\":\"hello.txt\"}") + ? ScriptedBackend.toolCallTurn("call_1", "read_file", "{\"file_path\":\"hello.txt\"}") : ScriptedBackend.textTurn("The file says VALUE=42.")); ByteArrayOutputStream out = new ByteArrayOutputStream(); ByteArrayOutputStream err = new ByteArrayOutputStream(); @@ -76,8 +80,10 @@ void oneShotTurnReadsAWorkspaceFileThroughTheBuiltInFileTools() throws Exception assertThat(exit, is(0)); } String console = out.toString(StandardCharsets.UTF_8); - assertThat(console, containsString("⚙ read_file {path=hello.txt}")); - assertThat(console, containsString("↳ VALUE=42")); + // the tool line and its result, as ConsoleSession renders them (unstyled here: not a terminal) + assertThat(console, containsString("● read_file {file_path=hello.txt}")); + // our read_file numbers the lines, so the result is " 1: VALUE=42" + assertThat(console, containsString("↳ 1: VALUE=42")); assertThat(console, containsString("The file says VALUE=42.")); List requests = backend.requests(); assertThat(requests, hasSize(2)); @@ -124,6 +130,251 @@ void interactiveModeRunsTurnsUntilExitAndKeepsHistory() throws Exception { assertThat(out.toString(StandardCharsets.UTF_8), containsString("answer 2")); } + @Test + void theSessionIsRecordedAndSavedWhereTheToolsWork() throws Exception { + // End to end through the real server: what was typed, what came back, written by /save with a + // timestamp on every line. /compact keeps the record -- it rewrites what the model is sent, + // not what happened -- and only /clear empties it. + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("answer " + call)); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + int exit = LocalAgent.run( + options, + new StringReader("what is two plus two\n/compact\n/save session.txt\n/exit\n"), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + + assertThat(exit, is(0)); + java.nio.file.Path written = workspace.resolve("session.txt"); + assertThat("the file lands in the workspace", java.nio.file.Files.exists(written), is(true)); + String text = java.nio.file.Files.readString(written, StandardCharsets.UTF_8); + assertThat(text, containsString("you: what is two plus two")); + assertThat("the answer survived the compaction", text, containsString("agent: answer 1")); + assertThat("and the compaction is noted rather than hidden", text, containsString("compacted")); + assertThat("every line is stamped", text.startsWith("["), is(true)); + } + } + + @Test + void clearingEmptiesTheRecordAsWell() throws Exception { + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("answer " + call)); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + LocalAgent.run( + options, + new StringReader("remember this\n/clear\n/save after-clear.txt\n/exit\n"), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + + String text = java.nio.file.Files.readString(workspace.resolve("after-clear.txt"), StandardCharsets.UTF_8); + assertThat("forget the session means the record too", text, not(containsString("remember this"))); + } + } + + @Test + void retryAsksAgainWithoutTheAnswerThatCameBack() throws Exception { + // The point of a retry: the model must not see what it said last time, or it says it again. + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("answer " + call)); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + LocalAgent.run( + options, + new StringReader("the question\n/retry\n/exit\n"), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + } + List requests = backend.requests(); + assertThat("it was asked twice", requests, hasSize(2)); + JsonNode second = requests.get(1).path("messages"); + assertThat("system and the question, and nothing else", second.size(), is(2)); + assertThat(second.get(1).path("content").asText(), is("the question")); + assertThat( + "the first answer is gone from the conversation", second.toString(), not(containsString("answer 1"))); + } + + @Test + void retryBeforeAnythingWasAskedSaysSo() throws Exception { + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("never")); + ByteArrayOutputStream out = new ByteArrayOutputStream(); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + LocalAgent.run( + options, + new StringReader("/retry\n/exit\n"), + new PrintStream(out, true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + } + assertThat(out.toString(StandardCharsets.UTF_8), containsString("nothing to retry")); + assertThat("and nothing was sent", backend.requests(), hasSize(0)); + } + + @Test + void aSystemPromptCanComeFromAFile() throws Exception { + java.nio.file.Path promptFile = workspace.resolve("persona.txt"); + java.nio.file.Files.writeString(promptFile, "You answer only in haiku."); + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("ok")); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", + "http://127.0.0.1:" + server.getPort() + "/v1", + "--workspace", + workspace.toString(), + "--system-file", + promptFile.toString() + }); + + LocalAgent.run( + options, + new StringReader("hello\n/exit\n"), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + } + JsonNode system = backend.requests().get(0).path("messages").get(0); + assertThat(system.path("role").asText(), is("system")); + assertThat(system.path("content").asText(), containsString("only in haiku")); + } + + @Test + void aSystemFileThatIsNotThereIsAUsageError() { + // Read at startup rather than at first use: a typo in a path is easy to miss, and a prompt long + // enough to be worth a file is long enough that its absence should not be a surprise mid-turn. + IllegalArgumentException thrown = org.junit.jupiter.api.Assertions.assertThrows( + IllegalArgumentException.class, + () -> AgentOptions.parse(new String[] { + "--base-url", + "http://localhost:1/v1", + "--system-file", + workspace.resolve("gone.txt").toString() + })); + + assertThat(thrown.getMessage(), containsString("--system-file cannot be read")); + } + + @Test + void aFencedBlockIsOneMessageWithItsNewlinesKept() throws Exception { + // A console sends on Enter, so a stack trace pasted into the prompt would become several + // questions. Between two fences it is one. + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("ok")); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + LocalAgent.run( + options, + new StringReader( + LocalAgent.BLOCK_FENCE + "\nline one\nline two\n" + LocalAgent.BLOCK_FENCE + "\n/exit\n"), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + } + assertThat("one question, not two", backend.requests(), hasSize(1)); + String asked = backend.requests() + .get(0) + .path("messages") + .get(1) + .path("content") + .asText(); + assertThat(asked, containsString("line one")); + assertThat(asked, containsString("line two")); + assertThat( + "the newline between them survived", + asked.contains("line one") && asked.indexOf("line two") > asked.indexOf("line one"), + is(true)); + assertThat("and the fences are not part of it", asked.contains("\"\"\""), is(false)); + } + + @Test + void endOfInputInsideABlockEndsTheSession() throws Exception { + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("never")); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + int exit = LocalAgent.run( + options, + new StringReader(LocalAgent.BLOCK_FENCE + "\nhalf a thought\n"), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + + assertThat("an unfinished block is not sent", exit, is(0)); + } + assertThat(backend.requests(), hasSize(0)); + } + + @Test + void loadingASavedTranscriptMakesItTheConversationAgain() throws Exception { + // Save in one session, load in the next: the model is sent what was said before, so it can be + // asked to carry on rather than to start over. + ScriptedBackend first = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("the earlier answer")); + try (OpenAiCompatServer server = server(first)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + LocalAgent.run( + options, + new StringReader("the earlier question" + NL + "/save earlier.txt" + NL + "/exit" + NL), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + } + + ScriptedBackend second = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("carrying on")); + ByteArrayOutputStream out = new ByteArrayOutputStream(); + try (OpenAiCompatServer server = server(second)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + LocalAgent.run( + options, + new StringReader("/load earlier.txt" + NL + "and then?" + NL + "/exit" + NL), + new PrintStream(out, true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + } + + assertThat(out.toString(StandardCharsets.UTF_8), containsString("loaded")); + JsonNode messages = second.requests().get(0).path("messages"); + assertThat("system, the earlier pair, and the new question", messages.size(), is(4)); + assertThat(messages.get(1).path("content").asText(), is("the earlier question")); + assertThat(messages.get(2).path("role").asText(), is("assistant")); + assertThat(messages.get(2).path("content").asText(), is("the earlier answer")); + assertThat(messages.get(3).path("content").asText(), is("and then?")); + } + + @Test + void loadingSomethingThatIsNotThereSaysSoAndChangesNothing() throws Exception { + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("ok")); + ByteArrayOutputStream out = new ByteArrayOutputStream(); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + LocalAgent.run( + options, + new StringReader("/load nowhere.txt" + NL + "still working?" + NL + "/exit" + NL), + new PrintStream(out, true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + } + assertThat(out.toString(StandardCharsets.UTF_8), containsString("cannot read")); + assertThat("the session carries on", backend.requests(), hasSize(1)); + assertThat( + "with nothing loaded into it", + backend.requests().get(0).path("messages").size(), + is(2)); + } + @Test void failedTurnExitsNonZero() throws Exception { ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("never")); @@ -185,4 +436,169 @@ void mavenJvmConfigPinsAUtf8ConsoleForExecJava() throws Exception { assertThat(content, containsString("-Dstdout.encoding=UTF-8")); assertThat(content, containsString("-Dstderr.encoding=UTF-8")); } + + @Test + void aToolCallStaysInTheHistorySoTheNextTurnSeesItHappened() throws Exception { + // The failure this pins: with only user text and the model's prose in the history, a small + // model stops calling tools after a few turns and starts DESCRIBING the work instead -- + // reporting exit codes and files that never existed. The evidence has to stay in the context. + Files.writeString(workspace.resolve("hello.txt"), "VALUE=42\n"); + ByteArrayOutputStream out = new ByteArrayOutputStream(); + ByteArrayOutputStream err = new ByteArrayOutputStream(); + ScriptedBackend backend = new ScriptedBackend((call, request) -> switch (call) { + case 1 -> ScriptedBackend.toolCallTurn("call_1", "read_file", "{\"file_path\":\"hello.txt\"}"); + case 2 -> ScriptedBackend.textTurn("The file says VALUE=42."); + default -> ScriptedBackend.textTurn("Understood."); + }); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + LocalAgent.run( + options, + new StringReader( + "what does hello.txt say?" + System.lineSeparator() + "and now?" + System.lineSeparator()), + new PrintStream(out, true, StandardCharsets.UTF_8), + new PrintStream(err, true, StandardCharsets.UTF_8)); + } + + // third request = second turn: the calls of turn one ride along with the new user message, so + // the model sees that they happened. Not as tool_calls messages (Atmosphere's assembleMessages + // rebuilds history as new ChatMessage(role, content) and drops the rest), and not as assistant + // text either (the model copied that into its own answers). + List requests = backend.requests(); + assertThat(requests.size(), is(3)); + List roles = new ArrayList<>(); + for (JsonNode message : requests.get(2).path("messages")) { + roles.add(message.path("role").asText()); + } + assertThat( + "strict alternation keeps every chat template happy", + roles, + contains("system", "user", "assistant", "user")); + String secondUserMessage = + requests.get(2).path("messages").get(3).path("content").asText(); + assertThat(secondUserMessage, containsString("read_file")); + assertThat(secondUserMessage, containsString("VALUE=42")); + assertThat(secondUserMessage, containsString("do not repeat it")); + assertThat("and the user's own words are still there", secondUserMessage, containsString("and now?")); + assertThat( + "the answer itself stays the model's own", + requests.get(2).path("messages").get(2).path("content").asText(), + is("The file says VALUE=42.")); + } + + @Test + void theActivityLineNamesTheRunningToolSoALongBuildLooksAlive() { + // "working… (90s)" during a two-minute mvn test is indistinguishable from a hang, so the line + // names the tool and how long IT has been running, next to the turn's total. + assertThat( + LocalAgent.activityLine('x', "Fettling", 12, "run_command", 9, 2), + is("x Fettling… (run_command 9s of 12s · 2 tool calls)")); + assertThat(LocalAgent.activityLine('x', "Fettling", 5, null, 0, 0), is("x Fettling… (5s)")); + assertThat(LocalAgent.activityLine('x', "Fettling", 30, null, 0, 3), is("x Fettling… (30s · 3 tool calls)")); + } + + @Test + void theSpinnerWordsAreOursAndHarmless() { + List words = LocalAgent.prompt(LocalAgent.SPINNER_WORDS) + .lines() + .map(String::strip) + .filter(word -> !word.isEmpty()) + .toList(); + + assertThat("enough variety to not repeat every other turn", words.size() > 15, is(true)); + assertThat("no duplicates", words.size(), is((int) + words.stream().distinct().count())); + for (String word : words) { + assertThat(word, word.matches("[A-Z][a-z-]+"), is(true)); + } + // Claude Code's own list is extracted from a proprietary binary and the public copies of it are + // unlicensed or CC BY-NC-SA; none of its words may appear here. + for (String theirs : List.of("Razzmatazzing", "Clauding", "Flibbertigibbeting", "Simmering", "Vibing")) { + assertThat(words.contains(theirs), is(false)); + } + assertThat(words.contains(LocalAgent.spinnerWord()), is(true)); + } + + @Test + void theToolNoteIsAddressedToTheModelAndNotWrittenAsItsOwnWords() { + // It first rode in front of the assistant's answer -- and the model copied it into its next + // reply, so the user read "(tools I actually ran this turn: …)" as the first line of an answer. + String note = LocalAgent.toolNote( + List.of(new ConsoleSession.ToolRound("grep", "{pattern=Test}", "3 matches in 2 files"))); + + assertThat(note, containsString("do not repeat it")); + assertThat(note, containsString("grep")); + assertThat(note, containsString("3 matches in 2 files")); + assertThat(LocalAgent.toolNote(List.of()), is("")); + } + + @Test + void compactionIsDecidedBeforeTheRequestAndOnlyWhenTheWindowIsKnown() { + AgentOptions on = AgentOptions.parse(new String[] {"--base-url", "u"}); + assertThat("on by default", on.isAutoCompact(), is(true)); + assertThat(on.getCompactAt(), is(AgentOptions.DEFAULT_COMPACT_AT)); + + // 70 % of 1000 tokens: 699 still fits, 700 does not + assertThat(LocalAgent.needsCompaction(on, 1000, 699), is(false)); + assertThat(LocalAgent.needsCompaction(on, 1000, 700), is(true)); + + // an unknown window is never guessed at + assertThat(LocalAgent.needsCompaction(on, StatusLine.UNKNOWN_CONTEXT, 1_000_000), is(false)); + + AgentOptions off = AgentOptions.parse(new String[] {"--base-url", "u", "--auto-compact", "false"}); + assertThat(off.isAutoCompact(), is(false)); + assertThat(LocalAgent.needsCompaction(off, 1000, 999), is(false)); + + AgentOptions early = AgentOptions.parse(new String[] {"--base-url", "u", "--compact-at", "50"}); + assertThat(LocalAgent.needsCompaction(early, 1000, 500), is(true)); + assertThat(LocalAgent.needsCompaction(early, 1000, 499), is(false)); + } + + @Test + void aMalformedCompactionFlagIsRejectedWithItsReason() { + assertThat( + org.junit.jupiter.api.Assertions.assertThrows( + IllegalArgumentException.class, + () -> AgentOptions.parse(new String[] {"--base-url", "u", "--auto-compact", "maybe"})) + .getMessage(), + containsString("Expected true or false")); + assertThat( + org.junit.jupiter.api.Assertions.assertThrows( + IllegalArgumentException.class, + () -> AgentOptions.parse(new String[] {"--base-url", "u", "--compact-at", "99"})) + .getMessage(), + containsString("between 10 and 95")); + } + + @Test + void compactingAnAlreadyCompactedHistoryIsRefusedInsteadOfRepeated() throws Exception { + // A compacted history is the summary plus its acknowledgement. Summarizing that again returns + // the same text for another model call -- and re-sends a byte-identical prompt, which llama.cpp + // answers with "need to evaluate at least 1 token for each active slot". + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("a summary")); + ByteArrayOutputStream out = new ByteArrayOutputStream(); + try (OpenAiCompatServer server = server(backend)) { + AgentOptions options = AgentOptions.parse(new String[] { + "--base-url", "http://127.0.0.1:" + server.getPort() + "/v1", "--workspace", workspace.toString() + }); + + LocalAgent.run( + options, + new StringReader("hello" + System.lineSeparator() + "/compact" + System.lineSeparator() + "/compact" + + System.lineSeparator()), + new PrintStream(out, true, StandardCharsets.UTF_8), + new PrintStream(new ByteArrayOutputStream(), true, StandardCharsets.UTF_8)); + } + + // one turn + one compaction = two requests; the second /compact must not add a third + assertThat(backend.requests(), hasSize(2)); + assertThat( + "the first compaction really happened", + out.toString(StandardCharsets.UTF_8), + containsString("compacted")); + assertThat(out.toString(StandardCharsets.UTF_8), containsString("already a summary")); + } } diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/PlainTerminalTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/PlainTerminalTest.java new file mode 100644 index 000000000..8f64e9987 --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/PlainTerminalTest.java @@ -0,0 +1,89 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.hasItem; +import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.not; +import static org.hamcrest.Matchers.nullValue; + +import java.io.BufferedReader; +import java.io.ByteArrayOutputStream; +import java.io.PrintStream; +import java.io.StringReader; +import java.nio.charset.StandardCharsets; +import java.util.List; +import org.junit.jupiter.api.Test; + +/** The console used whenever there is no terminal: one-shot runs, piped input, and these tests. */ +class PlainTerminalTest { + + private final ByteArrayOutputStream out = new ByteArrayOutputStream(); + + private PlainTerminal terminal(String typed) { + return new PlainTerminal( + new PrintStream(out, true, StandardCharsets.UTF_8), + typed == null ? null : new BufferedReader(new StringReader(typed)), + Ansi.PLAIN); + } + + private String written() { + return out.toString(StandardCharsets.UTF_8); + } + + @Test + void linesAreWrittenAndInputIsReadBackLineByLine() { + PlainTerminal terminal = terminal("first" + System.lineSeparator() + "second" + System.lineSeparator()); + + terminal.line("hello"); + + assertThat(written(), containsString("hello")); + assertThat(terminal.readLine("you> "), is("first")); + assertThat(written(), containsString("you> ")); + assertThat(terminal.readKey("allow? "), is("second")); + assertThat(terminal.readLine("you> "), is(nullValue())); + } + + @Test + void aPinnedStatusIsDroppedRatherThanPrintedRepeatedly() { + // Nothing can be pinned on a plain stream: a spinner refreshed four times a second would + // otherwise produce four lines a second in a piped log. The caller prints the status itself, + // once, above the prompt. + PlainTerminal terminal = terminal("x" + System.lineSeparator()); + terminal.status(java.util.List.of("⠙ Fettling… (5s)", "[state]")); + + terminal.readLine("you> "); + + assertThat(written(), not(containsString("Fettling"))); + assertThat(written(), containsString("you> ")); + } + + @Test + void anAnswerIsNormalisedSoTheCallerCanCompareItToOneCase() { + assertThat(terminal("YES" + System.lineSeparator()).readKey("? "), is("yes")); + assertThat(terminal(" A " + System.lineSeparator()).readKey("? "), is("a")); + } + + @Test + void withoutInputEverythingReadsAsEndOfInput() { + PlainTerminal terminal = terminal(null); + + assertThat(terminal.readLine("you> "), is(nullValue())); + assertThat(terminal.readKey("? "), is(nullValue())); + } + + @Test + void everyCommandNameIsOfferedForCompletion() { + List names = LocalAgent.commandNames(); + + for (SlashCommands.Command command : SlashCommands.Command.values()) { + for (String name : command.names()) { + assertThat(names, hasItem(name)); + } + } + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ServerPropsTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ServerPropsTest.java new file mode 100644 index 000000000..169df860c --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ServerPropsTest.java @@ -0,0 +1,45 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.is; + +import java.util.List; +import org.junit.jupiter.api.Test; + +class ServerPropsTest { + + @Test + void propsIsTriedNextToTheV1RoutesFirst() { + // llama-server serves /props at the root; this project's server serves it under both. + assertThat( + ServerProps.candidates("http://127.0.0.1:8080/v1"), + is(List.of("http://127.0.0.1:8080/props", "http://127.0.0.1:8080/v1/props"))); + assertThat( + ServerProps.candidates("http://127.0.0.1:8080/v1/"), + is(List.of("http://127.0.0.1:8080/props", "http://127.0.0.1:8080/v1/props"))); + assertThat(ServerProps.candidates("http://host/api"), is(List.of("http://host/api/props"))); + } + + @Test + void theContextSizeIsReadFromTheGenerationDefaults() { + assertThat( + ServerProps.parseContextSize("{\"default_generation_settings\":{\"n_ctx\":16384,\"model\":\"m\"}}"), + is(16384)); + } + + @Test + void anythingElseIsUnknownRatherThanAGuess() { + assertThat(ServerProps.parseContextSize("{}"), is(StatusLine.UNKNOWN_CONTEXT)); + assertThat(ServerProps.parseContextSize("not json"), is(StatusLine.UNKNOWN_CONTEXT)); + assertThat(ServerProps.parseContextSize("{\"n_ctx\":\"many\"}"), is(StatusLine.UNKNOWN_CONTEXT)); + } + + @Test + void anUnreachableServerIsUnknownAndDoesNotThrow() { + assertThat(ServerProps.contextSize("http://127.0.0.1:1/v1", "k"), is(StatusLine.UNKNOWN_CONTEXT)); + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ShellToolTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ShellToolTest.java index 0d92693e1..2b1e8fe1d 100644 --- a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ShellToolTest.java +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/ShellToolTest.java @@ -12,6 +12,7 @@ import java.nio.file.Files; import java.nio.file.Path; import java.time.Duration; +import java.util.List; import java.util.Map; import org.atmosphere.ai.tool.ToolDefinition; import org.junit.jupiter.api.Test; @@ -90,4 +91,22 @@ void timeoutKillsTheProcess() throws Exception { assertThat(result, startsWith("exit code: (killed after 0 s)")); } + + @Test + void outputIsReportedLineByLineWhileTheCommandStillRuns() throws Exception { + List live = new java.util.concurrent.CopyOnWriteArrayList<>(); + + String result = ShellTool.run( + workspace, + shell("echo one; echo two", "echo one& echo two"), + Duration.ofSeconds(30), + 10_000, + live::add); + + // the same lines reach the console while the process runs and the model afterwards + assertThat(live, is(java.util.List.of("one", "two"))); + assertThat(result, containsString("one")); + assertThat(result, containsString("two")); + assertThat(result, startsWith("exit code: 0")); + } } diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/SlashCommandsTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/SlashCommandsTest.java new file mode 100644 index 000000000..9426a00b0 --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/SlashCommandsTest.java @@ -0,0 +1,57 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; + +import java.util.Optional; +import org.junit.jupiter.api.Test; + +class SlashCommandsTest { + + @Test + void aCommandIsRecognisedWithItsAliasesAndCase() { + assertThat(SlashCommands.parse("/help").orElseThrow().command(), is(SlashCommands.Command.HELP)); + assertThat(SlashCommands.parse("/?").orElseThrow().command(), is(SlashCommands.Command.HELP)); + assertThat(SlashCommands.parse(" /HELP ").orElseThrow().command(), is(SlashCommands.Command.HELP)); + assertThat(SlashCommands.parse("/quit").orElseThrow().command(), is(SlashCommands.Command.EXIT)); + assertThat(SlashCommands.parse("/reset").orElseThrow().command(), is(SlashCommands.Command.CLEAR)); + assertThat(SlashCommands.parse("/approve").orElseThrow().command(), is(SlashCommands.Command.MODE)); + } + + @Test + void theRestOfTheLineIsTheArgument() { + SlashCommands compact = + SlashCommands.parse("/compact focus on the build errors").orElseThrow(); + + assertThat(compact.command(), is(SlashCommands.Command.COMPACT)); + assertThat(compact.arguments(), is("focus on the build errors")); + assertThat(compact.hasArguments(), is(true)); + assertThat(SlashCommands.parse("/compact").orElseThrow().hasArguments(), is(false)); + assertThat(SlashCommands.parse("/mode auto").orElseThrow().arguments(), is("auto")); + } + + @Test + void anythingElseIsAMessageForTheModel() { + // An unknown command is NOT rejected: a line may legitimately start with a slash, and a user + // who types /halp would rather get an answer than an error. The trade-off is deliberate. + assertThat(SlashCommands.parse("/halp"), is(Optional.empty())); + assertThat(SlashCommands.parse("/usr/bin/env is where?"), is(Optional.empty())); + assertThat(SlashCommands.parse("what does /help do?"), is(Optional.empty())); + assertThat(SlashCommands.parse(""), is(Optional.empty())); + assertThat(SlashCommands.parse(null), is(Optional.empty())); + } + + @Test + void everyCommandIsDocumentedInTheHelpText() { + String help = LocalAgent.prompt(LocalAgent.HELP_TEXT); + + for (SlashCommands.Command command : SlashCommands.Command.values()) { + assertThat(help, containsString(command.canonicalName())); + } + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TaskLoopTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TaskLoopTest.java new file mode 100644 index 000000000..49facd04e --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TaskLoopTest.java @@ -0,0 +1,245 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.nullValue; + +import java.io.ByteArrayOutputStream; +import java.io.PrintStream; +import java.nio.charset.StandardCharsets; +import java.nio.file.Files; +import java.nio.file.Path; +import java.time.Duration; +import java.util.List; +import net.ladenthin.llama.server.OpenAiCompatServer; +import net.ladenthin.llama.server.OpenAiServerConfig; +import org.atmosphere.ai.RetryPolicy; +import org.atmosphere.ai.fs.AgentFileSystem; +import org.atmosphere.ai.fs.WorkspaceAgentFileSystem; +import org.junit.jupiter.api.Test; +import org.junit.jupiter.api.io.TempDir; + +/** The loop: when it stops, when it does not, and what it sends. */ +class TaskLoopTest { + + private static final String MODEL_ID = "local-model"; + private static final Duration BUDGET = Duration.ofMinutes(5); + + @TempDir + Path workspace; + + private final ByteArrayOutputStream console = new ByteArrayOutputStream(); + + private AgentTerminal terminal() { + return new PlainTerminal(new PrintStream(console, true, StandardCharsets.UTF_8), null, Ansi.PLAIN); + } + + private TaskLoop.Outcome runLoop(ScriptedBackend backend, LoopOptions options) throws Exception { + OpenAiServerConfig config = OpenAiServerConfig.builder() + .host("127.0.0.1") + .port(0) + .apiKey("k") + .modelId(MODEL_ID) + .build(); + try (OpenAiCompatServer server = new OpenAiCompatServer(backend, config).start()) { + AgentRunner runner = new AgentRunner( + "http://127.0.0.1:" + server.getPort() + "/v1", + "k", + MODEL_ID, + List.of(), + "You are a test agent.", + 0.0, + 64, + 5) + .retryPolicy(RetryPolicy.NONE); + AgentFileSystem fs = new WorkspaceAgentFileSystem(workspace, AgentFileSystem.Limits.defaults()); + return TaskLoop.run(runner, fs, terminal(), workspace, options, () -> false, BUDGET, new ToolCallLog()); + } + } + + // ----- the stop marker ----- + + @Test + void onlyAWholeLineEndsTheLoop() { + assertThat(TaskLoop.isComplete("all done\n" + TaskLoop.SENTINEL), is(true)); + assertThat(TaskLoop.isComplete(" " + TaskLoop.SENTINEL + " "), is(true)); + // the failure this guards against: the model talking ABOUT the marker + assertThat(TaskLoop.isComplete("I will answer " + TaskLoop.SENTINEL + " when I am done."), is(false)); + assertThat(TaskLoop.isComplete("TASK_COMPLETE"), is(false)); + assertThat(TaskLoop.isComplete("still working"), is(false)); + } + + @Test + void everyStepSendsTheTaskVerbatimAndNamesTheFile() { + String prompt = TaskLoop.stepPrompt(new LoopOptions("rename the class", null, 5, null)); + + assertThat(prompt, containsString("rename the class")); + assertThat(prompt, containsString(TaskLoop.LOOP_FILE)); + assertThat(prompt, containsString(TaskLoop.SENTINEL)); + assertThat(prompt.contains("{"), is(false)); + } + + @Test + void theCheckCommandIsPartOfTheInstructionsWhenThereIsOne() { + assertThat( + TaskLoop.stepPrompt(new LoopOptions("fix it", null, 5, "mvn -q test")), containsString("mvn -q test")); + } + + // ----- the file ----- + + @Test + void theLoopFileIsCreatedOnceAndKeepsWhatIsInIt() throws Exception { + Path file = TaskLoop.ensureLoopFile(workspace, "write a parser"); + + assertThat(Files.readString(file), containsString("write a parser")); + Files.writeString(file, "edited by the agent"); + assertThat(TaskLoop.ensureLoopFile(workspace, "write a parser"), is(file)); + assertThat(Files.readString(file), is("edited by the agent")); + } + + @Test + void theFingerprintChangesWithTheContent() throws Exception { + Path file = TaskLoop.ensureLoopFile(workspace, "task"); + long before = TaskLoop.fingerprint(file); + + Files.writeString(file, "something else"); + + assertThat(TaskLoop.fingerprint(file) != before, is(true)); + assertThat(TaskLoop.fingerprint(workspace.resolve("absent.md")), is(-1L)); + } + + // ----- the loop itself, over the real server ----- + + @Test + void theLoopEndsWhenTheModelAnswersWithTheMarker() throws Exception { + ScriptedBackend backend = new ScriptedBackend((call, request) -> call < 3 + ? ScriptedBackend.textTurn("working on step ", String.valueOf(call)) + : ScriptedBackend.textTurn(TaskLoop.SENTINEL)); + + TaskLoop.Outcome outcome = runLoop(backend, new LoopOptions("do the thing", null, 10, null)); + + assertThat(outcome.completed(), is(true)); + assertThat(outcome.reason(), containsString("3 steps")); + assertThat(backend.requests().size(), is(3)); + } + + @Test + void everyStepStartsFromAnEmptyHistorySoTheContextCannotGrow() throws Exception { + ScriptedBackend backend = new ScriptedBackend((call, request) -> + call < 2 ? ScriptedBackend.textTurn("step") : ScriptedBackend.textTurn(TaskLoop.SENTINEL)); + + runLoop(backend, new LoopOptions("do the thing", null, 10, null)); + + // system + user, every single time: the file is the memory, the conversation is not + for (var request : backend.requests()) { + assertThat(request.path("messages").size(), is(2)); + } + } + + @Test + void aModelThatRepeatsItselfWithoutChangingAnythingIsStopped() throws Exception { + // no tool calls, no file change: exactly the shape of a model looping on its own answer + ScriptedBackend backend = new ScriptedBackend((call, request) -> ScriptedBackend.textTurn("thinking…")); + + TaskLoop.Outcome outcome = runLoop(backend, new LoopOptions("do the thing", null, 50, null)); + + assertThat(outcome.completed(), is(false)); + assertThat(outcome.reason(), containsString("no progress")); + assertThat(backend.requests().size(), is(TaskLoop.STALL_LIMIT)); + } + + @Test + void theStepLimitStopsALoopThatWouldOtherwiseRunOn() throws Exception { + ScriptedBackend backend = new ScriptedBackend((call, request) -> { + // something changes every step, so the stall detector never fires + Files.writeString(workspace.resolve(TaskLoop.LOOP_FILE), "step " + call); + return ScriptedBackend.textTurn("still working"); + }); + + TaskLoop.Outcome outcome = runLoop(backend, new LoopOptions("endless", null, 2, null)); + + assertThat(outcome.completed(), is(false)); + assertThat(outcome.reason(), containsString("step limit")); + assertThat(backend.requests().size(), is(2)); + } + + @Test + void aFailingCheckRejectsTheMarkerAndFeedsTheOutputBack() throws Exception { + String failing = ShellTool.isWindows() ? "exit 7" : "exit 7"; + ScriptedBackend backend = new ScriptedBackend((call, request) -> { + Files.writeString(workspace.resolve(TaskLoop.LOOP_FILE), "step " + call); + return ScriptedBackend.textTurn(TaskLoop.SENTINEL); + }); + + TaskLoop.Outcome outcome = runLoop(backend, new LoopOptions("claim it works", null, 2, failing)); + + assertThat("the model said done, the check said otherwise", outcome.completed(), is(false)); + assertThat( + backend.requests() + .get(1) + .path("messages") + .get(1) + .path("content") + .asText(), + containsString("the check")); + } + + // ----- the argument parser ----- + + @Test + void theTaskIsEverythingAfterTheFlags() { + LoopOptions plain = LoopOptions.parse("keep the README in sync with the code"); + + assertThat(plain.task(), is("keep the README in sync with the code")); + assertThat(plain.interval(), is(nullValue())); + assertThat(plain.maxSteps(), is(LoopOptions.DEFAULT_MAX_STEPS)); + assertThat(plain.check(), is(nullValue())); + } + + @Test + void flagsAreReadFromTheFrontAndAQuotedCheckKeepsItsSpaces() { + LoopOptions options = LoopOptions.parse("--every 90s --max 3 --check 'mvn -q test' fix the build"); + + assertThat(options.interval(), is(Duration.ofSeconds(90))); + assertThat(options.maxSteps(), is(3)); + assertThat(options.check(), is("mvn -q test")); + assertThat(options.task(), is("fix the build")); + } + + @Test + void durationsAreWrittenAsPeopleWriteThem() { + assertThat(LoopOptions.parseDuration("30s"), is(Duration.ofSeconds(30))); + assertThat(LoopOptions.parseDuration("5m"), is(Duration.ofMinutes(5))); + assertThat(LoopOptions.parseDuration("2h"), is(Duration.ofHours(2))); + assertThat(LoopOptions.parseDuration("10"), is(Duration.ofMinutes(10))); + } + + @Test + void aMalformedLoopCommandSaysWhatIsWrong() { + assertThat( + org.junit.jupiter.api.Assertions.assertThrows( + IllegalArgumentException.class, () -> LoopOptions.parse("")) + .getMessage(), + containsString("Usage: /loop")); + assertThat( + org.junit.jupiter.api.Assertions.assertThrows( + IllegalArgumentException.class, () -> LoopOptions.parse("--every soon do it")) + .getMessage(), + containsString("Not a duration")); + assertThat( + org.junit.jupiter.api.Assertions.assertThrows( + IllegalArgumentException.class, () -> LoopOptions.parse("--max zero do it")) + .getMessage(), + containsString("--max expects a number")); + assertThat( + org.junit.jupiter.api.Assertions.assertThrows( + IllegalArgumentException.class, () -> LoopOptions.parse("--bogus x do it")) + .getMessage(), + containsString("Unknown /loop flag")); + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TextEditsTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TextEditsTest.java new file mode 100644 index 000000000..2ab35252f --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TextEditsTest.java @@ -0,0 +1,138 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.not; +import static org.junit.jupiter.api.Assertions.assertThrows; + +import java.util.List; +import org.junit.jupiter.api.Test; + +/** The matching and rewriting rules that decide whether an edit lands. */ +class TextEditsTest { + + private static TextEdits.Edit edit(String oldString, String newString) { + return new TextEdits.Edit(oldString, newString, false); + } + + @Test + void aCrlfFileIsEditedByAModelThatOnlyWritesLineFeeds() { + // THE bug this class exists for: the framework compares raw content, so on Windows the + // model's "a\nb" never matches the file's "a\r\nb" and every edit silently fails. + String file = "int a = 1;\r\nint b = 2;\r\n"; + + String edited = TextEdits.apply(file, List.of(edit("int a = 1;\nint b = 2;", "int a = 3;\nint b = 4;"))); + + assertThat(edited, is("int a = 3;\r\nint b = 4;\r\n")); + } + + @Test + void theFilesOwnLineEndingAndByteOrderMarkSurvive() { + assertThat(TextEdits.apply("alpha\r\nbeta\r\n", List.of(edit("beta", "gamma"))), is("alpha\r\ngamma\r\n")); + assertThat(TextEdits.apply("alpha\nbeta\n", List.of(edit("beta", "gamma"))), is("alpha\ngamma\n")); + assertThat(TextEdits.lineEnding("a\r\nb"), is("\r\n")); + assertThat(TextEdits.lineEnding("a\nb"), is("\n")); + assertThat(TextEdits.lineEnding("no newline at all"), is("\n")); + } + + @Test + void anAmbiguousMatchNamesTheLineNumbers() { + String file = "x();\ny();\nx();\n"; + + TextEdits.EditException error = + assertThrows(TextEdits.EditException.class, () -> TextEdits.apply(file, List.of(edit("x();", "z();")))); + + assertThat(error.getMessage(), containsString("occurs 2 times")); + assertThat( + "naming the lines is what makes the next attempt possible", + error.getMessage(), + containsString("lines 1, 3")); + assertThat(error.getMessage(), containsString("replace_all")); + } + + @Test + void replaceAllTakesEveryOccurrence() { + assertThat( + TextEdits.apply("x();\ny();\nx();\n", List.of(new TextEdits.Edit("x();", "z();", true))), + is("z();\ny();\nz();\n")); + } + + @Test + void aMissShowsTheNearestLinesInsteadOfJustSayingNo() { + String file = "public void handle(String name) {\n log(name);\n}\n"; + + TextEdits.EditException error = assertThrows( + TextEdits.EditException.class, + () -> TextEdits.apply(file, List.of(edit("public void handle(String value) {", "x")))); + + assertThat(error.getMessage(), containsString("not found")); + assertThat(error.getMessage(), containsString("closest lines")); + assertThat(error.getMessage(), containsString("1: public void handle(String name) {")); + } + + @Test + void aMissWithNothingSimilarSaysToReadTheFileAgain() { + TextEdits.EditException error = assertThrows( + TextEdits.EditException.class, + () -> TextEdits.apply("alpha\nbeta\n", List.of(edit("zzzzzzzzzzzzzzzz", "x")))); + + assertThat(error.getMessage(), containsString("read the file again")); + // and it warns about the trap that causes many of these misses + assertThat(error.getMessage(), containsString("line numbers")); + } + + @Test + void severalEditsAreAllAppliedOrNoneAreAtAll() { + String file = "one\ntwo\nthree\n"; + + String edited = TextEdits.apply(file, List.of(edit("one", "1"), edit("three", "3"))); + assertThat(edited, is("1\ntwo\n3\n")); + + TextEdits.EditException error = assertThrows( + TextEdits.EditException.class, + () -> TextEdits.apply(file, List.of(edit("one", "1"), edit("absent", "x")))); + assertThat(error.getMessage(), containsString("edit 2 of 2")); + } + + @Test + void anEmptyOrUnchangedEditIsRejectedRatherThanSilentlyDoingNothing() { + assertThat( + assertThrows(TextEdits.EditException.class, () -> TextEdits.apply("a\n", List.of(edit("", "x")))) + .getMessage(), + containsString("must not be empty")); + assertThat( + assertThrows(TextEdits.EditException.class, () -> TextEdits.apply("a\n", List.of(edit("a", "a")))) + .getMessage(), + containsString("identical")); + assertThat( + assertThrows(TextEdits.EditException.class, () -> TextEdits.apply("a\n", List.of())) + .getMessage(), + containsString("No edits")); + } + + @Test + void theAnswerShowsTheEditedRegionWithLineNumbers() { + String edited = TextEdits.apply("a\nb\nc\nd\ne\n", List.of(edit("c", "CHANGED"))); + + String snippet = TextEdits.snippet(TextEdits.normalize(edited), "c", "CHANGED", 1); + + assertThat(snippet, containsString("3: CHANGED")); + assertThat(snippet, containsString("2: b")); + assertThat("the whole file is never echoed back", snippet, not(containsString("5: e"))); + } + + @Test + void similarityIsBetweenZeroAndOneAndOrdersTheCandidates() { + assertThat(TextEdits.similarity("abc", "abc"), is(1.0)); + assertThat(TextEdits.similarity("", "abc"), is(0.0)); + assertThat(TextEdits.similarity("int a = 1;", "int a = 2;") > TextEdits.CANDIDATE_MIN_SIMILARITY, is(true)); + assertThat( + TextEdits.similarity("int a = 1;", "completely different") < TextEdits.CANDIDATE_MIN_SIMILARITY, + is(true)); + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TranscriptTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TranscriptTest.java new file mode 100644 index 000000000..aba997165 --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/TranscriptTest.java @@ -0,0 +1,188 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.contains; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.not; + +import java.nio.charset.StandardCharsets; +import java.nio.file.Files; +import java.nio.file.Path; +import java.util.List; +import org.junit.jupiter.api.Test; +import org.junit.jupiter.api.io.TempDir; + +/** + * The record of what was said, which is a different thing from the conversation the model is sent. + */ +class TranscriptTest { + + private static final String LINE_SEPARATOR = System.lineSeparator(); + + @TempDir + Path directory; + + @Test + void entriesComeBackInTheOrderTheyWereSaid() { + Transcript transcript = new Transcript(); + + transcript.add(Transcript.Kind.USER, "first"); + transcript.add(Transcript.Kind.TOOL, "ls {} -> a b"); + transcript.add(Transcript.Kind.AGENT, "second"); + + assertThat( + transcript.entries().stream().map(Transcript.Entry::text).toList(), + contains("first", "ls {} -> a b", "second")); + } + + @Test + void twoEntriesInTheSameMomentAreBothKept() { + // The reason this is a list and not a map keyed by the time: a tool result and the answer that + // follows it regularly land in the same millisecond, and a map would keep one and lose the + // other without saying so. + Transcript transcript = new Transcript(); + + for (int i = 0; i < 50; i++) { + transcript.add(Transcript.Kind.AGENT, "entry " + i); + } + + assertThat(transcript.size(), is(50)); + } + + @Test + void blankTextIsNotWorthALine() { + Transcript transcript = new Transcript(); + + transcript.add(Transcript.Kind.AGENT, ""); + transcript.add(Transcript.Kind.AGENT, " "); + transcript.add(Transcript.Kind.AGENT, null); + + assertThat("a session of interrupted turns would otherwise fill the file", transcript.size(), is(0)); + } + + @Test + void everyLineCarriesTheTimeAndWhoSaidIt() { + Transcript transcript = new Transcript(); + transcript.add(Transcript.Kind.USER, "what time is it"); + + String rendered = transcript.render(); + + assertThat(rendered, containsString("you: what time is it")); + assertThat( + "a date and a clock time", + rendered.matches("(?s)\\[\\d{4}-\\d{2}-\\d{2} \\d{2}:\\d{2}:\\d{2}\\].*"), + is(true)); + } + + @Test + void savingWritesEverythingToTheNamedFile() throws Exception { + Transcript transcript = new Transcript(); + transcript.add(Transcript.Kind.USER, "hello"); + transcript.add(Transcript.Kind.AGENT, "hi"); + + Path written = transcript.save(directory, "session.txt"); + + assertThat(written.getFileName().toString(), is("session.txt")); + String text = Files.readString(written, StandardCharsets.UTF_8); + assertThat(text, containsString("you: hello")); + assertThat(text, containsString("agent: hi")); + } + + @Test + void savingWithoutANameUsesTheTime() throws Exception { + Transcript transcript = new Transcript(); + transcript.add(Transcript.Kind.USER, "hello"); + + Path written = transcript.save(directory, null); + + assertThat(written.getFileName().toString(), containsString("transcript-")); + assertThat(written.getFileName().toString().endsWith(".txt"), is(true)); + } + + @Test + void aConfiguredFileIsAppendedToAsThingsAreSaid() throws Exception { + // The point of it: a session that is killed still leaves what it had. Nothing is written at + // the end, so there is no end to miss. + Path live = directory.resolve("logs").resolve("live.txt"); + Transcript transcript = new Transcript(live); + + transcript.add(Transcript.Kind.USER, "one"); + assertThat("written already, not at the end", Files.readString(live), containsString("one")); + + transcript.add(Transcript.Kind.AGENT, "two"); + + List lines = Files.readAllLines(live, StandardCharsets.UTF_8); + assertThat(lines.size(), is(2)); + assertThat(lines.get(0), containsString("you: one")); + assertThat(lines.get(1), containsString("agent: two")); + } + + @Test + void aFileThatCannotBeWrittenDoesNotEndTheSession() { + // It exists to survive a bad ending, so it may not cause one. + Transcript transcript = new Transcript(directory); + + transcript.add(Transcript.Kind.USER, "the path is a directory, not a file"); + + assertThat("kept in memory regardless", transcript.size(), is(1)); + } + + @Test + void whatWasWrittenCanBeReadBack() { + Transcript written = new Transcript(); + written.add(Transcript.Kind.USER, "the question"); + written.add(Transcript.Kind.AGENT, "the answer"); + + List read = Transcript.parse(written.render()); + + assertThat(read.size(), is(2)); + assertThat(read.get(0).kind(), is(Transcript.Kind.USER)); + assertThat(read.get(0).text(), is("the question")); + assertThat(read.get(1).kind(), is(Transcript.Kind.AGENT)); + assertThat(read.get(1).text(), is("the answer")); + assertThat( + "the time survives, to the second the file records it in", + read.get(0).at(), + is(written.entries().get(0).at().truncatedTo(java.time.temporal.ChronoUnit.SECONDS))); + } + + @Test + void anAnswerWithNewlinesComesBackAsOneEntry() { + // An entry is not a line: an answer keeps its newlines when it is written, so reading line by + // line would turn one answer into several, each of them nonsense on its own. + Transcript written = new Transcript(); + written.add(Transcript.Kind.AGENT, "first line" + LINE_SEPARATOR + "second line"); + + List read = Transcript.parse(written.render()); + + assertThat(read.size(), is(1)); + assertThat(read.get(0).text(), containsString("first line")); + assertThat(read.get(0).text(), containsString("second line")); + } + + @Test + void aFileThatIsNotATranscriptYieldsNothing() { + // Rather than one wrong entry: a guess here would be replayed to the model as if it were said. + assertThat( + Transcript.parse("just some notes" + LINE_SEPARATOR + "and more") + .size(), + is(0)); + assertThat(Transcript.parse("").size(), is(0)); + } + + @Test + void clearingForgetsTheSession() { + Transcript transcript = new Transcript(); + transcript.add(Transcript.Kind.USER, "hello"); + + transcript.clear(); + + assertThat(transcript.size(), is(0)); + assertThat(transcript.render(), not(containsString("hello"))); + } +} diff --git a/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/WorkspaceToolsTest.java b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/WorkspaceToolsTest.java new file mode 100644 index 000000000..2bf03d4dd --- /dev/null +++ b/llama-atmosphere-agent/src/test/java/net/ladenthin/llama/atmosphere/WorkspaceToolsTest.java @@ -0,0 +1,205 @@ +// SPDX-FileCopyrightText: 2026 Bernard Ladenthin +// +// SPDX-License-Identifier: MIT + +package net.ladenthin.llama.atmosphere; + +import static org.hamcrest.MatcherAssert.assertThat; +import static org.hamcrest.Matchers.contains; +import static org.hamcrest.Matchers.containsString; +import static org.hamcrest.Matchers.is; +import static org.hamcrest.Matchers.not; + +import java.nio.charset.StandardCharsets; +import java.nio.file.Files; +import java.nio.file.Path; +import java.util.List; +import java.util.Map; +import org.atmosphere.ai.fs.AgentFileSystem; +import org.atmosphere.ai.fs.WorkspaceAgentFileSystem; +import org.atmosphere.ai.tool.ToolDefinition; +import org.junit.jupiter.api.BeforeEach; +import org.junit.jupiter.api.Test; +import org.junit.jupiter.api.io.TempDir; + +/** The three replaced tools, driven the way Atmosphere drives them. */ +class WorkspaceToolsTest { + + @TempDir + Path workspace; + + private Map, Object> scope; + private WorkspaceTools.ReadTracker tracker; + + @BeforeEach + void bindFileSystem() { + AgentFileSystem fs = new WorkspaceAgentFileSystem(workspace, AgentFileSystem.Limits.defaults()); + scope = Map.of(AgentFileSystem.class, fs); + tracker = new WorkspaceTools.ReadTracker(); + } + + private String call(ToolDefinition tool, Map arguments) throws Exception { + return String.valueOf(tool.executor().execute(arguments, scope)); + } + + private void write(String name, String content) throws Exception { + Files.writeString(workspace.resolve(name), content, StandardCharsets.UTF_8); + } + + // ----- read_file ----- + + @Test + void readingReturnsNumberedLinesAndSaysWhatItLeftOut() throws Exception { + write("big.txt", "l1\nl2\nl3\nl4\nl5\n"); + + String all = call(WorkspaceTools.readFile(tracker), Map.of("file_path", "big.txt")); + assertThat(all, containsString(" 1: l1")); + assertThat("a file that fits is shown without a footer", all, not(containsString("showing lines"))); + + String window = call(WorkspaceTools.readFile(tracker), Map.of("file_path", "big.txt", "offset", 2, "limit", 2)); + assertThat(window, containsString(" 2: l2")); + assertThat(window, containsString(" 3: l3")); + assertThat(window, not(containsString("l4"))); + assertThat(window, containsString("showing lines 2–3 of 5")); + } + + @Test + void readingPastTheEndAndReadingNothingAreExplained() throws Exception { + write("small.txt", "only\n"); + write("empty.txt", ""); + + assertThat( + call(WorkspaceTools.readFile(tracker), Map.of("file_path", "small.txt", "offset", 99)), + containsString("past the end")); + assertThat(call(WorkspaceTools.readFile(tracker), Map.of("file_path", "empty.txt")), containsString("empty")); + assertThat(call(WorkspaceTools.readFile(tracker), Map.of()), containsString("file_path is required")); + } + + // ----- edit_file ----- + + @Test + void anEditIsRefusedUntilTheFileWasRead() throws Exception { + write("code.txt", "alpha\n"); + + String refused = call( + WorkspaceTools.editFile(tracker), + Map.of("file_path", "code.txt", "old_string", "alpha", "new_string", "beta")); + assertThat(refused, containsString("read code.txt before editing")); + assertThat("nothing was written", Files.readString(workspace.resolve("code.txt")), is("alpha\n")); + + call(WorkspaceTools.readFile(tracker), Map.of("file_path", "code.txt")); + String edited = call( + WorkspaceTools.editFile(tracker), + Map.of("file_path", "code.txt", "old_string", "alpha", "new_string", "beta")); + assertThat(edited, containsString("Edited code.txt")); + assertThat(Files.readString(workspace.resolve("code.txt")), is("beta\n")); + } + + @Test + void theAnswerOfASuccessfulEditShowsTheChangedRegion() throws Exception { + write("code.txt", "a\nb\nc\nd\n"); + call(WorkspaceTools.readFile(tracker), Map.of("file_path", "code.txt")); + + String answer = call( + WorkspaceTools.editFile(tracker), + Map.of("file_path", "code.txt", "old_string", "c", "new_string", "CHANGED")); + + assertThat(answer, containsString("3: CHANGED")); + } + + @Test + void severalEditsArriveAsAListAndAreAllOrNothing() throws Exception { + write("code.txt", "one\ntwo\nthree\n"); + call(WorkspaceTools.readFile(tracker), Map.of("file_path", "code.txt")); + + String failed = call( + WorkspaceTools.editFile(tracker), + Map.of( + "file_path", + "code.txt", + "edits", + List.of( + Map.of("old_string", "one", "new_string", "1"), + Map.of("old_string", "absent", "new_string", "x")))); + assertThat(failed, containsString("edit 2 of 2")); + assertThat( + "a failed batch leaves the file untouched", + Files.readString(workspace.resolve("code.txt")), + is("one\ntwo\nthree\n")); + + String applied = call( + WorkspaceTools.editFile(tracker), + Map.of( + "file_path", + "code.txt", + "edits", + List.of( + Map.of("old_string", "one", "new_string", "1"), + Map.of("old_string", "three", "new_string", "3")))); + assertThat(applied, containsString("2 edits")); + assertThat(Files.readString(workspace.resolve("code.txt")), is("1\ntwo\n3\n")); + } + + @Test + void aFailedEditExplainsItselfInsteadOfSayingNo() throws Exception { + write("code.txt", "public void handle(String name) {\n}\n"); + call(WorkspaceTools.readFile(tracker), Map.of("file_path", "code.txt")); + + String answer = call( + WorkspaceTools.editFile(tracker), + Map.of("file_path", "code.txt", "old_string", "public void handle(String value) {", "new_string", "x")); + + assertThat(answer, containsString("closest lines")); + assertThat(answer, containsString("1: public void handle(String name) {")); + } + + // ----- grep ----- + + @Test + void theSearchSkipsBuildOutputAndRepositoryInternals() throws Exception { + Files.createDirectories(workspace.resolve("src")); + Files.createDirectories(workspace.resolve("target/classes")); + Files.createDirectories(workspace.resolve(".git")); + Files.writeString(workspace.resolve("src/Main.java"), "class Main { void needle() {} }\n"); + Files.writeString(workspace.resolve("target/classes/Main.txt"), "needle in build output\n"); + Files.writeString(workspace.resolve(".git/config"), "needle in the repository\n"); + + String answer = call(WorkspaceTools.grep(), Map.of("pattern", "needle")); + + assertThat(answer, containsString("src/Main.java")); + assertThat("build output is not source", answer, not(containsString("target/"))); + assertThat("the repository's own files are not source", answer, not(containsString(".git"))); + assertThat(answer, containsString("1 matches in 1 files")); + } + + @Test + void theSearchCanBeNarrowedAndCanListFilesOnly() throws Exception { + Files.writeString(workspace.resolve("A.java"), "needle\n"); + Files.writeString(workspace.resolve("B.txt"), "needle\n"); + + assertThat( + call(WorkspaceTools.grep(), Map.of("pattern", "needle", "glob", "*.java")), + not(containsString("B.txt"))); + String filesOnly = call(WorkspaceTools.grep(), Map.of("pattern", "needle", "files_only", true)); + assertThat(filesOnly, containsString("A.java")); + assertThat("files_only means no line content", filesOnly, not(containsString(" 1: needle"))); + } + + @Test + void anInvalidPatternIsAnErrorMessageNotAnException() throws Exception { + assertThat(call(WorkspaceTools.grep(), Map.of("pattern", "[unclosed")), containsString("Invalid regular")); + assertThat(call(WorkspaceTools.grep(), Map.of()), containsString("pattern is required")); + } + + @Test + void theToolSetReplacesTheFrameworksReadEditAndGrepWithoutDuplicates() { + List names = + WorkspaceTools.all(tracker).stream().map(ToolDefinition::name).toList(); + + assertThat(names, contains("ls", "read_file", "write_file", "edit_file", "glob", "grep", "delete", "rename")); + assertThat( + "no tool name may appear twice", + names.size(), + is(names.stream().distinct().count() == 8L ? 8 : -1)); + } +}