Skip to content

Default nogds to filesystem-aware O_DIRECT - #128

Merged
takeshi-yoshimura merged 1 commit into
foundation-model-stack:mainfrom
takeshi-yoshimura:tyos/nogds-odirect
Oct 8, 2026
Merged

takeshi-yoshimura merged 1 commit into
foundation-model-stack:mainfrom
takeshi-yoshimura:tyos/nogds-odirect

Conversation

@takeshi-yoshimura

Copy link
Copy Markdown
Collaborator

Move the unified copier's filesystem policy into common.is_odirect_enabled and reuse it for nogds. Local and unknown filesystems now use direct I/O by default; known network filesystems and platforms without O_DIRECT retain buffered reads.

Benchmark: H100 80GB (one GPU for Granite, four for DeepSeek). Use nogds, queue_size=2, device_memory_budget=-1, enforce-eager, max_model_len=32768, gpu_memory_utilization=0.97/0.8 respectively. NVMe is XFS at /opt/nvme/hf_home; /tmp/tyos/hf_cache is tmpfs.

Model (TP) Storage I/O Cold mean Warm mean
Granite (1) NVMe buffered 11.625691 4.198456
Granite (1) NVMe direct 4.890897 4.810586
Granite (1) tmpfs buffered 4.504860 4.365668
Granite (1) tmpfs direct 4.520546 4.242581
DeepSeek (4) NVMe buffered 50.294593 16.138033
DeepSeek (4) NVMe direct 15.921407 14.378360
DeepSeek (4) tmpfs buffered 11.303327 11.471072
DeepSeek (4) tmpfs direct 11.514589 13.038018

Cold requests POSIX_FADV_DONTNEED before every load to drop page cache, while warm performs no eviction.

I observed similar trends on different models. Some small models showed that bufferred I/O outperformed, but the gain was few seconds. So, in conclusion, I made O_DIRECT as default and let users disable it on some specific cases.

Refs #110

Reported-by: @gitbisector

Move the unified copier's filesystem policy into common.is_odirect_enabled
and reuse it for nogds. Local and unknown filesystems now use direct I/O
by default; known network filesystems and platforms without O_DIRECT
retain buffered reads.

Benchmark: H100 80GB (one GPU for Granite, four for DeepSeek).
Use nogds, queue_size=2, device_memory_budget=-1, enforce-eager,
max_model_len=32768, gpu_memory_utilization=0.97/0.8 respectively.
NVMe is XFS at /opt/nvme/hf_home; /tmp/tyos/hf_cache is tmpfs.

| Model (TP)   | Storage | I/O      | Cold mean | Warm mean |
|--------------|---------|----------|-----------|-----------|
| Granite (1)  | NVMe    | buffered | 11.625691 |  4.198456 |
| Granite (1)  | NVMe    | direct   |  4.890897 |  4.810586 |
| Granite (1)  | tmpfs   | buffered |  4.504860 |  4.365668 |
| Granite (1)  | tmpfs   | direct   |  4.520546 |  4.242581 |
| DeepSeek (4) | NVMe    | buffered | 50.294593 | 16.138033 |
| DeepSeek (4) | NVMe    | direct   | 15.921407 | 14.378360 |
| DeepSeek (4) | tmpfs   | buffered | 11.303327 | 11.471072 |
| DeepSeek (4) | tmpfs   | direct   | 11.514589 | 13.038018 |

Cold requests POSIX_FADV_DONTNEED before every load to drop page cache,
while warm performs no eviction.

I observed similar trends on different models. Some small models showed
that bufferred I/O outperformed, but the gain was few seconds. So, in
conclusion, I made O_DIRECT as default and let users disable it on some
specific cases.

Signed-off-by: Takeshi Yoshimura <tyos@jp.ibm.com>
@takeshi-yoshimura
takeshi-yoshimura merged commit 02560dc into foundation-model-stack:main Oct 8, 2026
13 checks passed
@takeshi-yoshimura
takeshi-yoshimura deleted the tyos/nogds-odirect branch October 8, 2026 03:02
@gitbisector

Copy link
Copy Markdown
Contributor

Thanks for landing this, and for the credit. I read the merged change and tried it on a GB10 (ext4/tmpfs/overlayfs) and a ZFS box; notes in #129, plus a small CI one in #130. None of it is urgent.

BTW, happy to review before merge whenever that's useful to you. Tag me on anything touching the loaders or the copiers and I'll turn it around quickly.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants