Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 8 additions & 3 deletions .github/actions/openai-agent/INTENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,14 +17,16 @@ The workflow controls:

- Provider connection details and credentials.
- Model and instructions.
- An optional bounded prompt context appended to the configured prompt.
- Output schema.
- Read-only filesystem capabilities.
- Stage, stream-idle, request-retry, model-turn, tool-call, and output-repair limits.
- An optional synchronous task-specific output normalizer.
- An optional task-specific validator and its bounded metadata.
- An optional structured-output file target.

Configuration, prompts, schemas, normalizer modules, and validator modules are trusted workflow inputs.
Configuration, prompts, prompt contexts, schemas, normalizer modules, and validator modules are trusted workflow inputs.
A prompt context reaches the model as instructions rather than as a tool result, so it may carry only text whose shape the workflow has constrained, never raw evidence or model output.
Evidence and model-produced tool arguments never grant capabilities.
The model can call only the declared `read_file`, `list_files`, and `search_text` tools within the configured paths and resource limits.

Expand Down Expand Up @@ -90,7 +92,10 @@ A semantic repair may use only necessary read-only evidence lookup.
After evidence lookup, the corrected value is requested without tools so the configured provider output format applies where supported.
Every schema-invalid canonical candidate is supplied to the validator in canonical order wherever parsed candidates are retained, so repair cannot silently discard usable findings.
Normalizer failures create no repair-history entry.
Every rejected attempt records the validation layer and a bounded, sanitized reason.
A semantic rejection carries a short content-free reason and may add a longer content-free detail and repair-only guidance.
Repair feedback carries the detail in place of the reason, so a repair sees everything the short reason had to drop.
Guidance may quote trusted identifiers the model must copy; it is sent only as repair feedback and never appears in diagnostics, logs, or failure reasons.
Every rejected attempt records the validation layer, a bounded, sanitized reason, and any validator detail.
The final rejection reason is retained when the repair budget is exhausted.

## Structured output files
Expand All @@ -113,7 +118,7 @@ Expose these bounded diagnostics:
- Provider finish reason and bounded error code when available.
- Token usage and whether it is complete.
- Accumulated turn and tool-call counts, including on failure.
- Every rejected output attempt with its validation layer and sanitized reason.
- Every rejected output attempt with its validation layer, sanitized reason, and sanitized validator detail.
- The first stream structural violation as a closed static value with no provider data.
- Ignored empty post-finish choices, repeated terminal-choice subsets, and the first rejected post-finish shape as closed, content-free values.

Expand Down
4 changes: 4 additions & 0 deletions .github/actions/openai-agent/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,10 @@ inputs:
config-file:
description: Workflow-controlled action configuration relative to the workspace
required: true
prompt-context:
description: Trusted workflow-generated text appended to the configured prompt
required: false
default: ""
validator:
description: Trusted workflow validator module selector in path#export form
required: false
Expand Down
2 changes: 1 addition & 1 deletion .github/actions/openai-agent/dist/index.js

Large diffs are not rendered by default.

13 changes: 9 additions & 4 deletions .github/actions/openai-agent/src/agent.js
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@ const TOOLS = [
type: "function",
function: {
name: "list_files",
description: "List bounded entries in an allowed directory.",
description: "List bounded entries in an allowed directory; the path . lists the allowed directories and files.",
parameters: {
type: "object",
additionalProperties: false,
Expand Down Expand Up @@ -566,7 +566,8 @@ async function runAgent({
role: "user",
content: [
"Your previous final response was invalid.",
candidate.reason,
candidate.detail ?? candidate.reason,
...(candidate.guidance === undefined ? [] : [candidate.guidance]),
toolsPermitted
? "Use only necessary read-only tools to correct this validation error, not to begin a new investigation, then return the corrected JSON in your next message."
: "Do not call tools or investigate further.",
Expand Down Expand Up @@ -642,9 +643,13 @@ async function runAgent({
state.candidates.push(candidate.value);
if (validation.ok) return candidate;
metrics?.recordOutputRejection({
activity, layer: "semantic", reason: validation.reason,
activity, layer: "semantic", reason: validation.reason, detail: validation.detail,
});
return { ok: false, kind: "validator", layer: "semantic", reason: validation.reason };
// The exhaustion failure keeps only the short reason; detail and guidance are repair feedback.
return {
ok: false, kind: "validator", layer: "semantic", reason: validation.reason,
detail: validation.detail, guidance: validation.guidance,
};
}

async function completion(messages, allowTools, activity) {
Expand Down
14 changes: 12 additions & 2 deletions .github/actions/openai-agent/src/config.js
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@ const {
DEFAULT_OUTPUT_REPAIRS, DEFAULT_REQUEST_RETRIES, DEFAULT_STAGE_TIMEOUT_MS,
DEFAULT_STREAM_IDLE_TIMEOUT_MS,
MAX_CONFIG_BYTES, MAX_METHODOLOGY_BYTES, MAX_METHODOLOGY_TOTAL_BYTES, MAX_PROMPT_BYTES,
MAX_PROMPT_CONTEXT_BYTES,
MAX_OUTPUT_REPAIRS, MAX_REQUEST_RETRIES, MAX_STAGE_TIMEOUT_MS, MAX_STREAM_IDLE_TIMEOUT_MS,
MAX_SCHEMA_BYTES, MAX_TOOL_CALLS, MAX_TURNS,
} = require("./limits");
Expand Down Expand Up @@ -64,7 +65,13 @@ function parseJson(text, code) {
}
}

function loadConfiguration(workspace, configFile) {
// The prompt context is trusted workflow text computed for this invocation, such as the identifiers a
// stage must echo back, so the model has them before it spends any turn looking for them.
function loadConfiguration(workspace, configFile, promptContext = "") {
if (typeof promptContext !== "string" ||
Buffer.byteLength(promptContext, "utf8") > MAX_PROMPT_CONTEXT_BYTES) {
fail("prompt context exceeds byte limit");
}
const workspaceReader = new WorkspaceSandbox(workspace);
const rawConfig = workspaceReader.readWorkflowFile(configFile, MAX_CONFIG_BYTES);
const parsed = parseJson(rawConfig, "configuration is not valid JSON");
Expand All @@ -86,7 +93,10 @@ function loadConfiguration(workspace, configFile) {
allowedRoots: config.allowed_roots,
allowedFiles: config.allowed_files,
});
const prompt = workspaceReader.readWorkflowFile(config.prompt_file, MAX_PROMPT_BYTES);
const configuredPrompt = workspaceReader.readWorkflowFile(config.prompt_file, MAX_PROMPT_BYTES);
const prompt = promptContext === ""
? configuredPrompt
: `${configuredPrompt.trimEnd()}\n\n${promptContext}`;
const schemaText = workspaceReader.readWorkflowFile(config.schema_file, MAX_SCHEMA_BYTES);
const schema = parseJson(schemaText, "output schema is not valid JSON");
if (schema === null || typeof schema !== "object" || Array.isArray(schema)) {
Expand Down
3 changes: 3 additions & 0 deletions .github/actions/openai-agent/src/limits.js
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@ module.exports = Object.freeze({
MAX_CONFIG_BYTES: 256 * 1024,
MAX_SCHEMA_BYTES: 256 * 1024,
MAX_PROMPT_BYTES: 1024 * 1024,
MAX_PROMPT_CONTEXT_BYTES: 64 * 1024,
MAX_METHODOLOGY_BYTES: 512 * 1024,
MAX_METHODOLOGY_TOTAL_BYTES: 2 * 1024 * 1024,
MAX_SOURCE_FILE_BYTES: 4 * 1024 * 1024,
Expand Down Expand Up @@ -32,6 +33,8 @@ module.exports = Object.freeze({
MAX_OUTPUT_REPAIRS: 5,
MAX_VALIDATOR_METADATA_BYTES: 4 * 1024,
MAX_VALIDATION_REASON_BYTES: 512,
MAX_VALIDATION_DETAIL_BYTES: 2048,
MAX_VALIDATION_GUIDANCE_BYTES: 16 * 1024,
MAX_OUTPUT_REJECTIONS: 8,
MAX_OUTPUT_REJECTION_REASON_BYTES: 240,
DEFAULT_STREAM_IDLE_TIMEOUT_MS: 300_000,
Expand Down
3 changes: 2 additions & 1 deletion .github/actions/openai-agent/src/main.js
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,7 @@ async function main(core, environment = process.env, OpenAIClient = OpenAI) {

const baseUrlInput = requiredInput(core, "base-url", "base URL input is missing");
const configFile = requiredInput(core, "config-file", "config file input is missing");
const promptContext = core.getInput("prompt-context");
const validatorSelector = core.getInput("validator");
const normalizerSelector = core.getInput("normalizer");
const validatorMetadata = parseMetadata(core.getInput("validator-metadata"));
Expand All @@ -40,7 +41,7 @@ async function main(core, environment = process.env, OpenAIClient = OpenAI) {
if (typeof workspace !== "string" || workspace.length === 0) {
throw new ActionError("workspace is unavailable");
}
const loaded = loadConfiguration(workspace, configFile);
const loaded = loadConfiguration(workspace, configFile, promptContext);
const config = loaded.config;
const validator = loadValidator(workspace, validatorSelector, validatorMetadata);
const normalizer = loadNormalizer(workspace, normalizerSelector);
Expand Down
15 changes: 10 additions & 5 deletions .github/actions/openai-agent/src/provider.js
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
"use strict";

const {
MAX_OUTPUT_REJECTION_REASON_BYTES, MAX_OUTPUT_REJECTIONS,
MAX_OUTPUT_REJECTION_REASON_BYTES, MAX_OUTPUT_REJECTIONS, MAX_VALIDATION_DETAIL_BYTES,
} = require("./limits");

const SAFE_DIAGNOSTIC_VALUE = /^[A-Za-z0-9._:-]{1,128}$/;
Expand Down Expand Up @@ -169,14 +169,19 @@ class RuntimeMetrics {
}

// A rejected output attempt is the only evidence left of why a stage exhausted its repairs, so it
// is kept as bounded telemetry rather than being reduced to the exhaustion itself.
recordOutputRejection({ activity, layer, reason }) {
// is kept as bounded telemetry rather than being reduced to the exhaustion itself. The short reason
// alone can drop what the attempt got wrong, so a validator's full detail travels beside it.
recordOutputRejection({ activity, layer, reason, detail }) {
if (this.outputRejections.length >= MAX_OUTPUT_REJECTIONS) return;
const sanitizedDetail = detail === undefined
? ""
: sanitizeReason(detail, MAX_VALIDATION_DETAIL_BYTES);
this.outputRejections.push({
attempt: this.outputRejections.length + 1,
activity,
layer,
reason: sanitizeReason(reason),
...(sanitizedDetail === "" ? {} : { detail: sanitizedDetail }),
});
}

Expand Down Expand Up @@ -227,12 +232,12 @@ function saturatingIncrement(value) {

// The alphabet is entirely ASCII, so what survives it measures the same in characters as in bytes and
// the budget can be applied by slicing.
function sanitizeReason(reason) {
function sanitizeReason(reason, maximumBytes = MAX_OUTPUT_REJECTION_REASON_BYTES) {
return String(reason ?? "")
.replace(UNSAFE_REASON_CHARACTER, " ")
.replace(/ +/g, " ")
.trim()
.slice(0, MAX_OUTPUT_REJECTION_REASON_BYTES)
.slice(0, maximumBytes)
.trimEnd();
}

Expand Down
97 changes: 82 additions & 15 deletions .github/actions/openai-agent/src/sandbox.js
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,28 @@ function normalizeRepositoryPath(value) {
return segments.join("/");
}

// Models routinely spell a tool path as `./file`, `dir/`, or `.`. Those segments name nothing, so
// they are dropped before the strict check instead of costing the call. An absolute path passes
// through untouched, so it, `..`, and `.git` stay rejected. An empty result names the workspace root.
function normalizeToolPath(value) {
if (typeof value !== "string" || value.startsWith("/")) return value;
return value.split("/").filter((segment) => segment !== "" && segment !== ".").join("/");
}

// A line longer than the read limit is cut at a code point boundary rather than failing the whole
// read, so one long line, such as a pull request body, no longer hides the rest of its file.
function boundedLine(line, from = 0) {
let bytes = 0;
let end = from;
for (const character of line.slice(from)) {
const size = Buffer.byteLength(character, "utf8");
if (bytes + size > MAX_LINE_BYTES) break;
bytes += size;
end += character.length;
}
return line.slice(from, end);
}

function validateOutputFilePath(value) {
if (value === "") return "";
const relative = normalizeRepositoryPath(value);
Expand Down Expand Up @@ -203,14 +225,18 @@ class WorkspaceSandbox {
positiveInteger(args.end_line, "invalid end line");
if (requestedEnd < start || requestedEnd - start + 1 > MAX_READ_LINES) fail("invalid line range");

const target = this.resolve(args.path, "file");
const target = this.resolve(normalizeToolPath(args.path), "file");
const lines = this.readText(target).split(/\r?\n/);
if (start > lines.length) fail("start line exceeds file length");
const end = Math.min(requestedEnd, lines.length);
const selected = [];
const truncatedLines = [];
for (let index = start; index <= end; index++) {
const line = lines[index - 1];
if (Buffer.byteLength(line, "utf8") > MAX_LINE_BYTES) fail("source line exceeds byte limit");
let line = lines[index - 1];
if (Buffer.byteLength(line, "utf8") > MAX_LINE_BYTES) {
line = boundedLine(line);
truncatedLines.push(index);
}
selected.push(`${index}: ${line}`);
}
return boundJson({
Expand All @@ -219,6 +245,7 @@ class WorkspaceSandbox {
start_line: start,
end_line: end,
truncated: end < lines.length,
...(truncatedLines.length === 0 ? {} : { truncated_lines: truncatedLines }),
content: selected.join("\n"),
});
}
Expand All @@ -229,7 +256,10 @@ class WorkspaceSandbox {
(args.recursive !== undefined && typeof args.recursive !== "boolean")) {
fail("invalid list_files arguments");
}
const target = this.resolve(args.path, "directory");
const requested = normalizeToolPath(args.path);
// The workspace root is not itself a capability, so listing it names the ones that exist.
if (requested === "") return this.listCapabilities();
const target = this.resolve(requested, "directory");
const entries = [];
const traversal = this.walk(target, args.recursive === true, ({ relative, metadata }) => {
if (entries.length >= MAX_LIST_ENTRIES) return false;
Expand All @@ -245,6 +275,20 @@ class WorkspaceSandbox {
});
}

listCapabilities() {
const entries = [
...this.allowedRoots.map((root) => ({ path: root.relative, type: "directory" })),
...this.allowedFiles.map((file) => ({ path: file.relative, type: "file" })),
].sort((left, right) => left.path.localeCompare(right.path, "en"));
return boundJson({
ok: true,
path: ".",
recursive: false,
truncated: entries.length > MAX_LIST_ENTRIES,
entries: entries.slice(0, MAX_LIST_ENTRIES),
});
}

searchText(args) {
assertObject(args, ["path", "query"]);
if (typeof args.path !== "string" || typeof args.query !== "string" ||
Expand All @@ -253,22 +297,40 @@ class WorkspaceSandbox {
fail("invalid search_text arguments");
}

const requested = normalizeToolPath(args.path);
let target;
try {
target = this.resolve(args.path, "file");
target = this.resolve(requested, "file");
} catch (fileError) {
if (!(fileError instanceof ActionError) ||
!["path is not a regular file", "unsupported filesystem object"].includes(fileError.code)) {
throw fileError;
}
target = this.resolve(args.path, "directory");
target = this.resolve(requested, "directory");
}

const results = [];
let filesSearched = 0;
let traversal = { truncated: false };
// Matches are collected against the serialized budget, so enough long lines end the search with
// what fits instead of failing the whole call. The envelope is reserved at its largest shape.
let remainingBytes = MAX_TOOL_RESULT_BYTES - Buffer.byteLength(JSON.stringify({
ok: true, path: target.relative, files_searched: MAX_SEARCH_FILES, truncated: false, matches: [],
}), "utf8");
let budgetExhausted = false;
const collect = (match) => {
const size = Buffer.byteLength(JSON.stringify(match), "utf8") + (results.length === 0 ? 0 : 1);
if (size > remainingBytes) {
budgetExhausted = true;
return false;
}
remainingBytes -= size;
results.push(match);
return results.length < MAX_SEARCH_RESULTS;
};
const searchFile = (file) => {
if (filesSearched >= MAX_SEARCH_FILES || results.length >= MAX_SEARCH_RESULTS) return false;
if (filesSearched >= MAX_SEARCH_FILES || results.length >= MAX_SEARCH_RESULTS ||
budgetExhausted) return false;
filesSearched++;
let text;
try {
Expand All @@ -282,11 +344,16 @@ class WorkspaceSandbox {
throw error;
}
for (const [index, line] of text.split(/\r?\n/).entries()) {
if (line.includes(args.query)) {
if (Buffer.byteLength(line, "utf8") > MAX_LINE_BYTES) continue;
results.push({ path: file.relative, line: index + 1, text: line });
if (results.length >= MAX_SEARCH_RESULTS) return false;
}
const column = line.indexOf(args.query);
if (column === -1) continue;
// An overlong line is shown from its first match, so the returned text contains it.
const match = Buffer.byteLength(line, "utf8") <= MAX_LINE_BYTES
? { path: file.relative, line: index + 1, text: line }
: {
path: file.relative, line: index + 1, column: column + 1,
text: boundedLine(line, column), text_truncated: true,
};
if (!collect(match)) return false;
}
return true;
};
Expand All @@ -303,7 +370,7 @@ class WorkspaceSandbox {
ok: true,
path: target.relative,
files_searched: filesSearched,
truncated: traversal.truncated ||
truncated: traversal.truncated || budgetExhausted ||
filesSearched >= MAX_SEARCH_FILES || results.length >= MAX_SEARCH_RESULTS,
matches: results,
});
Expand Down Expand Up @@ -393,6 +460,6 @@ function positiveInteger(value, code) {
}

module.exports = {
OUTPUT_DIRECTORY, WorkspaceSandbox, boundJson, normalizeRepositoryPath, validateOutputFilePath,
writeOutputFile,
OUTPUT_DIRECTORY, WorkspaceSandbox, boundJson, normalizeRepositoryPath, normalizeToolPath,
validateOutputFilePath, writeOutputFile,
};
Loading
Loading