Skip to content

Add bounded-memory GDS loading with aligned chunk reads - #126

Merged
takeshi-yoshimura merged 1 commit into
foundation-model-stack:mainfrom
takeshi-yoshimura:tyos/gds-membudget
Oct 4, 2026
Merged

takeshi-yoshimura merged 1 commit into
foundation-model-stack:mainfrom
takeshi-yoshimura:tyos/gds-membudget

Conversation

@takeshi-yoshimura

@takeshi-yoshimura takeshi-yoshimura commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Support budgeted GDS partial reads, tensor readiness, stable chunk allocation sizes, and fixed cuFile cache accounting. Align CUDA chunk file ranges and destination pointers to 4 KiB while charging padding to the memory planner. Coalesce overlapping boundary pages and preserve selected tensor coverage.

Validate planner bounds, sparse ranges, EOF reads, readiness, casting, and allocation modes. Selected suites pass locally (354 passed, 5 skipped) and on cccxc580 NVMe (356 passed, 3 skipped); pre-commit checks pass. Compare 8 GiB timings and verify full 78 GiB contents under near-full GPU memory for resident and preallocated-destination loading.

Refs #110

Support budgeted GDS partial reads, tensor readiness, stable chunk allocation
sizes, and fixed cuFile cache accounting. Align CUDA chunk file ranges and
destination pointers to 4 KiB while charging padding to the memory planner.
Coalesce overlapping boundary pages and preserve selected tensor coverage.
For AMD hipFile 0.2.x--0.4.x, validate the library version and reserve zero
fixed device cache bytes without calling its unimplemented property API.
For unsupported or unknown AMD hipFile versions, select nogds/unified before
planning and preserve the original budget, chunk limits, and tensor selection.
Use the fallback copier costs; infeasible fallback plans raise
BudgetInfeasibleError before tensor I/O. Preserve NVIDIA cache queries and
unbudgeted HIP direct loading. Restrict the nvidia-fs node check to NVIDIA.

Validate planner bounds, sparse ranges, EOF reads, readiness, casting, and
allocation modes. Selected suites pass locally (354 passed, 5 skipped) and
on cccxc580 NVMe (356 passed, 3 skipped); pre-commit checks pass. Compare
8 GiB timings and verify full 78 GiB contents under near-full GPU memory
for resident and preallocated-destination loading.

Native runtime discovery tests use isolated mock CUDA/HIP libraries to cover
version guards, missing APIs, and cache query errors. Related suites after the
HIP fallback fix pass (305 passed, 5 skipped); an additional UMA fallback
case passes with the final runtime tests (85 passed). Pre-commit checks pass.
AMD hardware validation remains outstanding.

Signed-off-by: Takeshi Yoshimura <tyos@jp.ibm.com>
@takeshi-yoshimura
takeshi-yoshimura merged commit fc863ca into foundation-model-stack:main Oct 4, 2026
13 checks passed
@takeshi-yoshimura
takeshi-yoshimura deleted the tyos/gds-membudget branch October 4, 2026 06:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant