Repository navigation
Add bounded-memory GDS loading with aligned chunk reads - #126
Merged
takeshi-yoshimura merged 1 commit intoOct 4, 2026
Merged
takeshi-yoshimura merged 1 commit into
takeshi-yoshimura merged 1 commit into
Conversation
Support budgeted GDS partial reads, tensor readiness, stable chunk allocation sizes, and fixed cuFile cache accounting. Align CUDA chunk file ranges and destination pointers to 4 KiB while charging padding to the memory planner. Coalesce overlapping boundary pages and preserve selected tensor coverage. For AMD hipFile 0.2.x--0.4.x, validate the library version and reserve zero fixed device cache bytes without calling its unimplemented property API. For unsupported or unknown AMD hipFile versions, select nogds/unified before planning and preserve the original budget, chunk limits, and tensor selection. Use the fallback copier costs; infeasible fallback plans raise BudgetInfeasibleError before tensor I/O. Preserve NVIDIA cache queries and unbudgeted HIP direct loading. Restrict the nvidia-fs node check to NVIDIA. Validate planner bounds, sparse ranges, EOF reads, readiness, casting, and allocation modes. Selected suites pass locally (354 passed, 5 skipped) and on cccxc580 NVMe (356 passed, 3 skipped); pre-commit checks pass. Compare 8 GiB timings and verify full 78 GiB contents under near-full GPU memory for resident and preallocated-destination loading. Native runtime discovery tests use isolated mock CUDA/HIP libraries to cover version guards, missing APIs, and cache query errors. Related suites after the HIP fallback fix pass (305 passed, 5 skipped); an additional UMA fallback case passes with the final runtime tests (85 passed). Pre-commit checks pass. AMD hardware validation remains outstanding. Signed-off-by: Takeshi Yoshimura <tyos@jp.ibm.com>
takeshi-yoshimura
force-pushed
the
tyos/gds-membudget
branch
from
October 4, 2026 00:33
4eb4172 to
5bfe7dc
Compare
takeshi-yoshimura
merged commit Oct 4, 2026
fc863ca
into
foundation-model-stack:main
13 checks passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Support budgeted GDS partial reads, tensor readiness, stable chunk allocation sizes, and fixed cuFile cache accounting. Align CUDA chunk file ranges and destination pointers to 4 KiB while charging padding to the memory planner. Coalesce overlapping boundary pages and preserve selected tensor coverage.
Validate planner bounds, sparse ranges, EOF reads, readiness, casting, and allocation modes. Selected suites pass locally (354 passed, 5 skipped) and on cccxc580 NVMe (356 passed, 3 skipped); pre-commit checks pass. Compare 8 GiB timings and verify full 78 GiB contents under near-full GPU memory for resident and preallocated-destination loading.
Refs #110