Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
115 changes: 115 additions & 0 deletions .github/smoke-agent-jar.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,115 @@
#!/usr/bin/env bash

# SPDX-FileCopyrightText: 2026 Bernard Ladenthin <bernard.ladenthin@gmail.com>
#
# SPDX-License-Identifier: MIT

# Smoke test for the llama-atmosphere-agent release asset, run exactly the way the README
# tells a user to run it: the agent jar lies next to a core fat jar and is started with
# `java -jar`, so the core is found only through the agent manifest's Class-Path. That
# makes this the check that the two release assets actually fit together (same version in
# the file names, nothing missing on either side), which no test run from the source tree
# can see.
#
# 1. the agent jar carries no core: started alone, loading a model fails with
# NoClassDefFoundError for net.ladenthin.llama.LlamaModel;
# 2. --help exits 0 and prints the usage;
# 3. a one-shot prompt with the model loaded in-process answers "2+2" with a 4;
# 4. a one-shot prompt that needs a reading tool (read_file) reads a marker file from
# --workspace, so the whole tool loop runs through the shipped jars.
#
# Usage: smoke-agent-jar.sh <jar-dir> <model-path>
# <jar-dir> must hold exactly one llama-atmosphere-agent-*-jar-with-dependencies.jar and at
# least one core fat jar its Class-Path names. Output of each run is kept in agent-*.log in
# the working directory (uploaded by the CI job on failure).
set -euo pipefail

JAR_DIR="${1:?usage: smoke-agent-jar.sh <jar-dir> <model-path>}"
MODEL="${2:?usage: smoke-agent-jar.sh <jar-dir> <model-path>}"
TIMEOUT="${AGENT_SMOKE_TIMEOUT:-600}"
# The checks grep plain text; never let a CI runner that forces colour put escapes in it.
export NO_COLOR=1
unset CLICOLOR_FORCE

fail() {
echo "::error::$*" >&2
exit 1
}

[ -f "$MODEL" ] || fail "model not found: $MODEL"
MODEL="$(cd "$(dirname "$MODEL")" && pwd)/$(basename "$MODEL")"
JAR_DIR="$(cd "$JAR_DIR" && pwd)"

mapfile -t AGENTS < <(find "$JAR_DIR" -maxdepth 1 -name 'llama-atmosphere-agent-*-jar-with-dependencies.jar' | sort)
[ "${#AGENTS[@]}" -eq 1 ] || fail "expected exactly 1 agent jar in $JAR_DIR, got ${#AGENTS[@]}: ${AGENTS[*]:-none}"
AGENT="${AGENTS[0]}"
echo "Agent jar: $(basename "$AGENT") ($(du -h "$AGENT" | cut -f1))"

# The manifest names the core jars by file name; at least one of them must be here, or
# `java -jar` would start without a core. unzip wraps manifest lines at 72 bytes with a
# leading space, so the continuation lines are joined first.
CLASS_PATH="$(unzip -p "$AGENT" META-INF/MANIFEST.MF | tr -d '\r' | sed -e ':a' -e 'N' -e '$!ba' -e 's/\n //g' \
| sed -n 's/^Class-Path: //p')"
[ -n "$CLASS_PATH" ] || fail "agent manifest has no Class-Path"
found=""
for entry in $CLASS_PATH; do
if [ -f "$JAR_DIR/$entry" ]; then
found="$entry"
break
fi
done
[ -n "$found" ] || fail "none of the core jars the agent manifest names is in $JAR_DIR: $CLASS_PATH (present: $(ls "$JAR_DIR"))"
echo "Core jar picked up via Class-Path: $found"

# 1. Without a core next to it the agent must not work: that is what makes it small.
ALONE="$(mktemp -d)"
cp "$AGENT" "$ALONE/"
set +e
timeout "$TIMEOUT" java -jar "$ALONE/$(basename "$AGENT")" --model "$MODEL" --plain --prompt hi \
> agent-alone.log 2>&1 < /dev/null
rc=$?
set -e
rm -rf "$ALONE"
[ "$rc" -ne 0 ] || fail "the agent jar ran without a core jar next to it - does it bundle the core?"
grep -q 'NoClassDefFoundError: net/ladenthin/llama/LlamaModel' agent-alone.log \
|| { cat agent-alone.log; fail "agent without core failed, but not for the missing core (see above)"; }
echo "OK: the agent jar carries no core"

# 2. --help
java -jar "$AGENT" --help > agent-help.log 2>&1 < /dev/null || { cat agent-help.log; fail "--help exited non-zero"; }
grep -q 'Usage: LocalAgent' agent-help.log || { cat agent-help.log; fail "--help printed no usage"; }
echo "OK: --help"

run_agent() {
local log="$1"
shift
set +e
timeout "$TIMEOUT" java -jar "$AGENT" --model "$MODEL" --ngl 0 --plain --temperature 0 "$@" \
> "$log" 2>"${log%.log}.err.log" < /dev/null
local status=$?
set -e
if [ "$status" -ne 0 ]; then
echo "===== $log =====" && cat "$log"
echo "===== ${log%.log}.err.log (last 80 lines) =====" && tail -n 80 "${log%.log}.err.log"
fail "agent exited with $status ($log)"
fi
}

# 3. A plain answer through the in-process server.
run_agent agent-answer.log --prompt 'What is 2 + 2? Answer with one short sentence.'
grep -q '4' agent-answer.log || { cat agent-answer.log; fail "the answer does not contain 4"; }
echo "OK: plain answer"

# 4. A tool round: the marker exists only in the file, so it reaches the output only
# through read_file (whose result the console prints) or an answer built from it.
WORKSPACE="$(mktemp -d)"
MARKER="AGENT_SMOKE_$(date +%s)_$RANDOM"
printf '%s\n' "$MARKER" > "$WORKSPACE/marker.txt"
run_agent agent-tool.log --workspace "$WORKSPACE" \
--prompt 'Read the file marker.txt with the read_file tool and tell me its exact content.'
rm -rf "$WORKSPACE"
grep -q 'read_file' agent-tool.log || { cat agent-tool.log; fail "the model did not call read_file"; }
grep -q "$MARKER" agent-tool.log || { cat agent-tool.log; fail "the marker never reached the output"; }
echo "OK: tool round (read_file)"

echo "Agent release asset smoke test passed."
104 changes: 92 additions & 12 deletions .github/workflows/publish.yml
Original file line number Diff line number Diff line change
Expand Up @@ -533,17 +533,19 @@ jobs:
# job — never in the dockcross cross-compilers (which have no node) or per-platform.
# ---------------------------------------------------------------------------
# ---------------------------------------------------------------------------
# llama-atmosphere-agent: the standalone (non-reactor, unpublished) local coding-agent
# project that wires Atmosphere's built-in OpenAI-compatible agent runtime to this
# project's OpenAiCompatServer. Two jobs, mirroring the langchain4j pair:
# llama-atmosphere-agent: the standalone (non-reactor, never on Maven Central) local
# coding-agent project that wires Atmosphere's built-in OpenAI-compatible agent runtime to
# this project's OpenAiCompatServer. Three jobs, all publish gates:
# - model-free: unit tests + the wire-contract tests, which drive the REAL
# OpenAiCompatServer over a loopback socket with a scripted backend (no native lib,
# no GGUF) and pin the streamed tool_calls / role=tool / multi-round shape — seconds,
# on every PR.
# on every PR. It also builds the GitHub Release asset (the agent jar WITHOUT the core).
# - model-backed: the same loop against the cached Qwen2.5-1.5B tool model through the
# downloaded Linux native library (chat, streaming, tool call + result, read/write/read
# loop). Validation-only, not a publish gate: a small model's wording is not a release
# signal, the deterministic contract is the model-free job.
# loop). A gate since its assertions stopped pinning wording (content checks are limited
# to facts no instruct model gets wrong and to tool results) and it ran green throughout.
# - smoke-agent-linux (further down, after package-fatjars): the release asset itself,
# started next to the real core fat jar.
# The project is built with -Dllama.version=<reactor version> against the core that was just
# installed to the local repo, so it always tests the code of this checkout.
# ---------------------------------------------------------------------------
Expand All @@ -569,6 +571,28 @@ jobs:
run: mvn -B --no-transfer-progress -f llama-atmosphere-agent/pom.xml "-Dllama.version=${VERSION}" spotless:check
- name: Build and test (unit + model-free wire contract against the real OpenAiCompatServer)
run: mvn -B --no-transfer-progress -f llama-atmosphere-agent/pom.xml "-Dllama.version=${VERSION}" verify
# The GitHub Release asset llama-atmosphere-agent-<core version>-jar-with-dependencies.jar:
# the agent plus Atmosphere/JLine, WITHOUT the core (src/assembly/agent-jar.xml), so it is a
# few MB and the natives are not in the release twice. Never deployed to Maven Central.
# smoke-agent-linux launches it next to the real core fat jar; the attach jobs sign it.
- name: Build the agent release jar (without the core)
run: >
mvn -B --no-transfer-progress -f llama-atmosphere-agent/pom.xml "-Dllama.version=${VERSION}"
-P assembly -DskipTests package
- name: Collect the agent release jar + sha256
run: |
mkdir -p agent-jar
cp "llama-atmosphere-agent/target/llama-atmosphere-agent-${VERSION}-jar-with-dependencies.jar" agent-jar/
(cd agent-jar && for f in *.jar; do sha256sum "$f" > "$f.sha256"; done)
ls -la agent-jar
- name: Upload the agent release jar
uses: actions/upload-artifact@v7
with:
name: llama-atmosphere-agent-jar
path: agent-jar/
compression-level: 0 # jars are already deflated
retention-days: 7
if-no-files-found: error

test-java-llama-atmosphere-agent-integration:
name: Integration Test llama-atmosphere-agent (model-backed)
Expand Down Expand Up @@ -3665,6 +3689,50 @@ jobs:
server-err.log
if-no-files-found: warn

# The agent release asset, launched the way the README tells a user to: `java -jar` on the agent
# jar lying next to the real all-backends Linux fat jar, so the core is found only through the
# agent manifest's Class-Path. Proves the two assets fit together (matching version in the file
# names, nothing missing on either side), that the agent jar carries no core, and that a
# one-shot answer and a read_file tool round work through the shipped jars.
smoke-agent-linux:
name: Smoke test the agent release jar (Linux)
needs: [test-java-llama-atmosphere-agent, package-fatjars, verify-model-cache]
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: actions/download-artifact@v8
with:
name: llama-atmosphere-agent-jar
path: agent-assets/
- uses: actions/download-artifact@v8
with:
name: llama-fatjar-smoke-linux
path: agent-assets/
- name: Restore shared GGUF model cache (populated by download-models; no re-download)
uses: actions/cache/restore@v6
with:
path: models/
key: gguf-models-${{ hashFiles('.github/models.csv') }}
enableCrossOsArchive: true
- name: Validate model files
run: .github/validate-models.sh
- uses: actions/setup-java@v6
with:
distribution: 'temurin'
java-version: ${{ env.JAVA_VERSION }}
# The agent is Java 21 (Atmosphere's floor), unlike the Java 8 core: its ceiling is 65.
- name: Verify Java 21 bytecode (no class newer than major 65)
run: .github/verify-bytecode-version.sh --max-major 65 agent-assets/llama-atmosphere-agent-*-jar-with-dependencies.jar
- name: Run the agent release-jar smoke test
run: .github/smoke-agent-jar.sh agent-assets "models/${TOOL_MODEL_NAME}"
- name: Upload agent logs
if: failure()
uses: actions/upload-artifact@v7
with:
name: agent-smoke-linux-logs
path: agent-*.log
if-no-files-found: warn

smoke-fatjar-windows:
name: Smoke test all-backends fat jar (Windows)
needs: [package-fatjars, verify-model-cache]
Expand Down Expand Up @@ -3880,7 +3948,7 @@ jobs:

publish-snapshot:
name: Publish Snapshot to Central
needs: [check-snapshot, crosscompile-linux-x86_64-cuda, crosscompile-android-aarch64-opencl, package-android-aar, test-android-emulator, code-style, test-java-llama-langchain4j, test-java-llama-kotlin, package-fatjars, smoke-fatjar-linux, smoke-fatjar-windows, smoke-fatjar-linux-aarch64, smoke-fatjar-windows-arm64, smoke-fatjar-macos]
needs: [check-snapshot, crosscompile-linux-x86_64-cuda, crosscompile-android-aarch64-opencl, package-android-aar, test-android-emulator, code-style, test-java-llama-langchain4j, test-java-llama-kotlin, test-java-llama-atmosphere-agent, test-java-llama-atmosphere-agent-integration, package-fatjars, smoke-fatjar-linux, smoke-fatjar-windows, smoke-fatjar-linux-aarch64, smoke-fatjar-windows-arm64, smoke-fatjar-macos, smoke-agent-linux]
if: needs.check-snapshot.result == 'success' && inputs.publish_to_central
runs-on: ubuntu-latest
environment: maven-central
Expand Down Expand Up @@ -4059,11 +4127,11 @@ jobs:

github-snapshot:
name: Update Snapshot Pre-release on GitHub
needs: [publish-snapshot, package-fatjars]
needs: [publish-snapshot, package-fatjars, test-java-llama-atmosphere-agent]
# Also runs when publish-snapshot FAILED (not when skipped/cancelled): a Central
# publish-poll timeout reds that job after the artifacts were already uploaded —
# the GitHub pre-release assets must not be lost in that case.
if: ${{ !cancelled() && (needs.publish-snapshot.result == 'success' || needs.publish-snapshot.result == 'failure') && needs.package-fatjars.result == 'success' }}
if: ${{ !cancelled() && (needs.publish-snapshot.result == 'success' || needs.publish-snapshot.result == 'failure') && needs.package-fatjars.result == 'success' && needs.test-java-llama-atmosphere-agent.result == 'success' }}
runs-on: ubuntu-latest
# maven-central so the GPG_PRIVATE_KEY / GPG_PASSPHRASE secret is delivered (it is
# scoped to this environment) for signing the fat jars below. This environment has
Expand All @@ -4086,6 +4154,12 @@ jobs:
with:
name: llama-fatjars
path: snapshot-assets/
# The agent jar (+ sha256) — built without the core, run next to one of the fat jars above.
# Same directory, so sign-fatjars.sh signs it and the upload glob attaches it.
- uses: actions/download-artifact@v8
with:
name: llama-atmosphere-agent-jar
path: snapshot-assets/
# GPG-sign the fat jars so each carries a detached .asc signature alongside its
# .sha256 checksum — signature parity with the thin jars (which maven-gpg signs at
# deploy) and with the BAF / srcmorph sibling fat jars. The .sha256 files (integrity)
Expand Down Expand Up @@ -4147,7 +4221,7 @@ jobs:
publish-release:
name: Publish Release to Central
if: needs.check-tag.result == 'success' && inputs.publish_to_central
needs: [check-tag, crosscompile-linux-x86_64-cuda, crosscompile-android-aarch64-opencl, package-android-aar, test-android-emulator, code-style, test-java-llama-langchain4j, test-java-llama-kotlin, package-fatjars, smoke-fatjar-linux, smoke-fatjar-windows, smoke-fatjar-linux-aarch64, smoke-fatjar-windows-arm64, smoke-fatjar-macos]
needs: [check-tag, crosscompile-linux-x86_64-cuda, crosscompile-android-aarch64-opencl, package-android-aar, test-android-emulator, code-style, test-java-llama-langchain4j, test-java-llama-kotlin, test-java-llama-atmosphere-agent, test-java-llama-atmosphere-agent-integration, package-fatjars, smoke-fatjar-linux, smoke-fatjar-windows, smoke-fatjar-linux-aarch64, smoke-fatjar-windows-arm64, smoke-fatjar-macos, smoke-agent-linux]
runs-on: ubuntu-latest
environment: maven-central
permissions:
Expand Down Expand Up @@ -4326,11 +4400,11 @@ jobs:

github-release-signed:
name: Attach Signed Binaries to GitHub Release
needs: [publish-release, package-fatjars]
needs: [publish-release, package-fatjars, test-java-llama-atmosphere-agent]
# Also runs when publish-release FAILED (not when skipped/cancelled): a Central
# publish-poll timeout reds that job after the artifacts were already uploaded —
# the GitHub release assets must not be lost in that case.
if: ${{ !cancelled() && (needs.publish-release.result == 'success' || needs.publish-release.result == 'failure') && needs.package-fatjars.result == 'success' }}
if: ${{ !cancelled() && (needs.publish-release.result == 'success' || needs.publish-release.result == 'failure') && needs.package-fatjars.result == 'success' && needs.test-java-llama-atmosphere-agent.result == 'success' }}
runs-on: ubuntu-latest
# maven-central so the GPG_PRIVATE_KEY / GPG_PASSPHRASE secret is delivered (it is
# scoped to this environment) for signing the fat jars below. This environment has
Expand All @@ -4353,6 +4427,12 @@ jobs:
with:
name: llama-fatjars
path: release-assets/
# The agent jar (+ sha256) — built without the core, run next to one of the fat jars above.
# Same directory, so sign-fatjars.sh signs it and the upload glob attaches it.
- uses: actions/download-artifact@v8
with:
name: llama-atmosphere-agent-jar
path: release-assets/
# GPG-sign the fat jars so each carries a detached .asc signature alongside its
# .sha256 checksum — signature parity with the thin jars (which maven-gpg signs at
# deploy) and with the BAF / srcmorph sibling fat jars. The .sha256 files (integrity)
Expand Down
17 changes: 17 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,15 @@ from version 5.0.0 onward. Pre-fork releases (`1.x`–`4.2.0`) were authored by
## [Unreleased]

### Fixed
- **`ToolCallingIntegrationTest#requiredToolCallIsParsedFromStreamingResponse` failed on both Windows
x86-64 jobs after the b11211 bump** (the Ubuntu run and the blocking twin stayed green). The streamed
request generated its full 512 tokens without a tool call. The prompt ("Write an example") never asked
for the tool, so the grammar-constrained answer was a fragile ~90-token call even when it worked, and a
numerically different CPU path on those runners (most likely upstream's new tiled k-quant matmul,
which picks its microkernel by ISA) tipped greedy decoding over. The test pins how a tool call is
parsed and streamed, not whether a 1.5B model infers one, so the user message now asks for the call
outright; both assertions carry the streamed content, `finish_reason` and chunk count, so a future
failure says what the model did instead of `but: was ""`.
- **`LlamaModel.setLogger` was silently overridden by every model load, and never saw the server's own
log lines.** llama.cpp's `common_init()` — run on each load — re-points `llama_log_set()` at its own
default callback, so a logger set *before* `new LlamaModel(…)` (the natural order) stopped receiving
Expand All @@ -30,6 +39,14 @@ from version 5.0.0 onward. Pre-fork releases (`1.x`–`4.2.0`) were authored by
documented as the no-ops they are (`common_init()` forces both on), `setLogFile` as additive.

### Added
- **The agent is a release asset: `llama-atmosphere-agent-<version>-jar-with-dependencies.jar`**, with
`.sha256` and a GPG `.asc`, on every GitHub release and the rolling `snapshot` pre-release — never on
Maven Central. It carries **no core** (~7 MB instead of hundreds, natives not in the release twice):
put it next to a core fat jar of the same version and `java -jar` finds the core through its manifest
`Class-Path`, or name both with `java -cp`. A new CI job, `smoke-agent-linux`, launches exactly that
pair (bytecode ≤ Java 21, the jar alone must fail for the missing core, `--help`, a one-shot answer and
a `read_file` round on the cached tool model), and it, the model-free agent job and the model-backed
agent integration test now gate both publish jobs.
- **`llama-atmosphere-agent`: `--log-verbosity <n>` (default `2`) and `--verbose`** for the in-process
`--model` mode. llama.cpp's per-request INFO lines go to stderr, the console the streamed answer is
printed to, and interleaved with it; the agent now loads the model with warnings-and-errors only. A
Expand Down
Loading
Loading