Repository navigation
[ENHANCEMENT] Restore bounded multi-file reading in a single native tool call #1948
Description
Activity
- added a commit that references this issue
on Oct 6, 2026 Follow-up: real-model A/B pilot results
We have now run real model-backed tasks, rather than only the local reader comparison described earlier. These results update the original investigation's statement that end-to-end performance had not been measured.
Summary: on the tested known-path workflow, one bounded batch reduced recorded API requests from 6 to 2 in every pair, while returning identical requested source content. Observed median task time was 39.0% lower on GPT-6.1 Sol and 49.9% lower on GPT-5.6 Sol. This is a small controlled pilot, not a universal or statistically established speedup claim.
1. What was tested
- GPT-6.1 Sol, through ChatGPT OAuth / Codex's Responses Lite path.
- GPT-5.6 Sol, through ChatGPT OAuth / Codex's ordinary Responses path.
- Three A/B pairs per model: 12 completed tasks in total.
- The development extension included the modern companion batch reader. The inspected source and separate read-only benchmark worktree used snapshot 78ef79de2, which remained unchanged.
- This tested the new bounded batch interface, not the legacy multi-file fallback. Batch processing remained sequential internally; general parallel dispatch was not introduced for this test.
The two arms were:
Arm Reading pattern A — single-file Exactly five calls to the modern single-file reader, strictly sequential, proceeding to the next file after each result. B — batch Exactly one call to the bounded companion reader, containing the same five files and slice parameters in the same order. Both prompts supplied already-known, independent paths. They prohibited searching, shell/MCP substitutes, modifications, subtasks, and repeated reads. After reading, the model had to produce the same five-row factual table in English, within 250 words.
The exact slices were:
Source Starting line Line count Shared batch parameter schema 1 56 Runtime reading defaults 1 8 Batch budget helper 1 27 Batch executor loop 65 21 Native batch tool definition 1 26 The answer checks covered batch/path limits, reader defaults, byte-budget reserves, sequential execution and stopping statuses, and image/optional-parameter support. This produced 138 requested numbered source lines in every task.
2. Run procedure and measurement
For each model, we started six fresh tasks with no prior conversation and ran A1 → B1 → B2 → A2 → A3 → B3. The prompts differed in their required invocation pattern, not the requested files, ranges, or final factual task. Read auto-approval within the workspace was part of the intended protocol, to avoid measuring human confirmation time.
Measurements came from the persisted task UI records, API conversation histories, and history summaries—not from grouped chat cards:
- API requests: counted recorded request-start events and cross-checked them against persisted assistant responses. This includes the final-answer request: five reading turns plus an answer for A, versus one batch turn plus an answer for B.
- Tool calls/content: inspected actual call arguments and results, verifying their order, offsets, limits, and exact source-line content.
- Task time: initial stored user-prompt timestamp → completion-ready marker, excluding the user's later completion acknowledgment. We did not use the streamed answer's original timestamp as its finish time: that timestamp is retained while the partial answer is updated. This is an extension-observed duration, not a browser-rendering measurement.
- Usage: summed per-request token/cache records and cross-checked them against persisted history totals.
- Errors/repeats: checked recorded failures/retries and actual reading calls. Unrecorded transport-level retries cannot be ruled out.
The recorded profile names were gpt-6.1-sol and gpt-5.6-sol, respectively, and all tasks recorded Code mode and the same benchmark workspace. The task records do not preserve a complete historical configuration snapshot: exact reasoning effort, service tier, approval settings, and compiled extension revision were not independently established after the runs. Fresh tasks also contain dynamic environment metadata, so complete API inputs are not byte-identical.
3. Individual timing results
All times are seconds. Each row compares the corresponding A/B pair within one model.
Model Pair A: five single-file calls B: one batch A minus B GPT-6.1 Sol 1 41.642 29.233 12.409 GPT-6.1 Sol 2 44.896 27.387 17.509 GPT-6.1 Sol 3 75.224 23.384 51.840 GPT-5.6 Sol 1 32.784 15.383 17.401 GPT-5.6 Sol 2 41.614 16.428 25.186 GPT-5.6 Sol 3 27.472 24.282 3.190 Model Median A Median B Difference between medians Observed reduction GPT-6.1 Sol 44.896 s 27.387 s 17.509 s 39.0% GPT-5.6 Sol 32.784 s 16.428 s 16.356 s 49.9% Batch was faster in all six pairs. All completed runs were retained. GPT-6.1 Sol's A3 included a 34.467-second interval between its third and fourth request-start records. Its cause is unknown; that interval includes model/network waiting, streaming, and local handling, and is not an isolated inference or disk-read measurement. It was not removed to improve the result.
An earlier incomplete task under the GPT-6.1 profile occurred before the GPT-5.6 series; it was not included in the six completed GPT-5.6 runs.
4. Correctness, output size, and usage
All 12 completed runs:
- Followed the specified five-single-call or one-batch-call protocol.
- Returned the same 138 requested source lines, checked directly against the unchanged snapshot.
- Answered 5/5 factual checks correctly, in the requested English table format.
- Had no recorded read errors, repeated reads, or unexpected additional tools.
Metric GPT-6.1 A GPT-6.1 B GPT-5.6 A GPT-5.6 B Recorded API requests per task 6 2 6 2 Reading calls per task 5 1 5 1 Cumulative input tokens per task 139,934 45,903 139,442 45,739 Median cache-read tokens 114,304 21,760 113,920 42,624 Median input tokens minus cache reads 25,630 24,143 25,522 3,115 Median output tokens 1,228 1,063 1,012 1,027 Reader-result text, UTF-8 bytes 6,632 6,963 6,632 6,963 Cumulative input is repeated context across requests, not unique context size. The roughly 67% reduction in cumulative input must not be presented as a 67% billing/quota saving. On GPT-6.1, input excluding cache reads fell by only about 5.8%. Caching varied in the GPT-5.6 runs: A2 had 134,912 cache-read tokens and 4,530 input tokens excluding cache reads, unlike A1/A3. These are therefore not cache-controlled latency comparisons. Cache-write tokens were zero throughout.
Batch reader output was about 5% larger, due to its result/status/continuation envelope. Byte counts include reading-result text but exclude surrounding environment messages and serialized message wrappers. Stored cost values were zero on these OAuth profiles; that does not establish free execution or monetary savings.
Batch entries carried range-limited/truncated labels because the requested slices did not cover entire files, sometimes excluding only a trailing empty line. Every requested line was present. There was no aggregate-budget loss of requested content or follow-up reading; the batch result was well below its 65,536-byte budget.
5. Interpretation and relationship to #375
These pilots support a narrow claim: a single bounded batch can avoid intervening model turns for already-known independent reads, even with sequential internal processing, and that correlated with lower task latency in these runs.
The distinction between the two models is important:
- GPT-6.1 Sol's current Responses Lite path enforces one tool call per response. A dispatcher cannot parallelize five separate calls that the provider does not return together. Batching addresses that round-trip constraint without changing the provider contract.
- GPT-5.6 Sol is not subject to that same Lite routing restriction. Here, A was deliberately forced to issue calls sequentially. Its result does not show that batch is better than asking GPT-5.6 to emit several independent single-file calls in one response.
This did not benchmark #375 or general parallel dispatch. Those changes remain complementary, but a third comparison arm is needed: multiple single-file calls emitted in one response, then serial versus concurrent execution where supported.
Finally, three pairs per model are only a smoke pilot. We did not test autonomous file selection, dependent exploration, large/near-budget batches, images/documents, or realistic long-running coding tasks. No statistical significance or fixed provider-independent speedup is claimed. The next useful step is more paired runs, explicit configuration capture and cache reporting, and the multi-call-response comparison—not extrapolating these percentages to every task.
- added 11 commits that reference this issue
on Oct 8, 2026 - added a commit that references this issue
on Oct 10, 2026
Problem (one or two sentences)
Zoo's current native file-reading tool exposes exactly one file per call, even when the agent already knows several independent files it needs for context. On provider paths that permit only one tool call per response, this requires additional model/tool round trips that a bounded multi-file read could avoid.
Context (who is affected and when)
This affects codebase exploration, comparing related implementations, and inspecting a source file together with its tests or configuration. The clearest case is an agent that already knows the relevant paths and reading ranges; it should not need another model response merely to retrieve the next independent file.
The current GPT-6.1 Sol / ChatGPT OAuth Codex path is particularly relevant: Zoo routes it through Responses Lite and explicitly disables multiple tool calls per response. Other providers can already return several calls in one response, so the round-trip benefit is provider-dependent; that capability should not be confused with concurrent local execution.
The goal is fewer model round trips, not faster disk I/O or unrestricted parallel tool execution. No end-to-end speedup has been benchmarked in this investigation.
Desired behavior (conceptual, not technical)
Allow the agent to request a bounded collection of independent files in one native tool call and receive a clearly separated result for each file.
Each file should retain the modern reading capabilities: a line slice or a structurally meaningful code block, with individual read parameters and clear clipping/continuation information. Keep the ordinary single-file operation for dependent exploration, where the next path is chosen only after examining an earlier result.
A batch may process its files sequentially internally. It should work without the general parallel-tool-execution feature being implemented or enabled.
Constraints / preferences
Research: what exists today
The following findings were checked against upstream revision d7963fc on 2026-10-06:
Historical context: restore the capability, not the old implementation
The proposed enhancement should retain the modern simplified per-file semantics instead of rolling back the refactor.
Related issues — not duplicates
presentAssistantMessage#375 / [Epic 4] Parallel Tool Execution #359 / [Tracking] Parallelizable Tasks — Implementation Roadmap #355 — General parallel dispatch, its epic, and the implementation roadmap. These address multiple independently executing tool calls and shared execution state. This enhancement can deliver round-trip savings without that dispatcher.Acceptance criteria
Proposed approach
Add a bounded multi-file interface built on the modern per-file reader, with one aggregate budget and per-entry validation/results. Start with sequential internal processing to avoid broad concurrency changes.
Whether this is an additive form of the existing tool or a companion batch tool is an implementation choice; preserve existing calls and avoid an ambiguous schema or misleading provider-independent parallel-call guidance.
Trade-offs / risks
Batching can increase unnecessary context if the agent speculatively requests files it does not need. Guidance should recommend batches only for already-known independent reads and retain single-file reads for dependent steps. Aggregate clipping must remain useful rather than starving later entries silently.
The main risks are context growth, approval/access-policy bypass, schema complexity, and compatibility with existing conversation history. None requires unrestricted concurrent execution.
Out of scope
Request checklist
This request was prepared with Zoo Code assistance from a read-only review of upstream source, Git history, and linked PR/issue discussions. No implementation changes or paid model benchmark runs were performed as part of the investigation.