Skip to content

test: CUDA memory pool fallback in python backend - #8992

Open
mattwittwer wants to merge 2 commits into
mainfrom
mwittwer/python_backend_pinned_fallback_test
Open

mattwittwer wants to merge 2 commits into
mainfrom
mwittwer/python_backend_pinned_fallback_test

Conversation

@mattwittwer

@mattwittwer mattwittwer commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

What does the PR do?

Adds regression coverage for silent output corruption in the Python backend when the CUDA memory pool is exhausted and GPU output tensors fall back to pinned host memory (#7148). The runtime fix is python_backend#457.

The existing IOTest.test_ensemble_io already drives the ensemble_io pipeline of three chained dlpack_io_identity Python models, with per-request flags choosing which stage emits its output as a GPU tensor, and asserts exact equality against a 1000 x FP32 input. This change re-runs that test with --cuda-memory-pool-byte-size=0:1024, so every 4000-byte GPU output overflows the pool and the ensemble's response allocator falls back to pinned memory for each of them. The block also fails if the server log does not contain the core's falling back to pinned system memory warning, so it cannot pass without exercising the fallback path. No new models or Python code.

Checklist

  • PR title reflects the change and is of format <commit_type>: <Title>
  • Changes are described in the pull request.
  • Related issues are referenced.
  • Populated github labels field
  • Added test plan and verified test passes.
  • Verified that the PR passes existing CI.
  • Verified copyright is correct on all changed files.
  • Added succinct git squash message before merging.
  • All template sections are filled out.
  • Optional: Additional screenshots for behavior/output changes with before/after.

Commit Type:

  • test

Related PRs:

Where should the reviewer start?

  • qa/L0_backend_python/io/test.sh — the new block IOTest.test_ensemble_io with GPU outputs falling back to pinned memory: model setup mirrors the existing default trial, SERVER_ARGS adds the 1024-byte CUDA pool, and the post-run grep on the server log guards against a vacuous pass.

Test plan:

  • CI Pipeline ID: [71702518]

Caveats:

Background

Related Issues:

@mattwittwer mattwittwer self-assigned this Oct 1, 2026
@greptile-apps

greptile-apps Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 3/5

[Low risk] Adds a test case for GPU memory fallback behavior.

The PR is not ready to merge because the new test block can signal its process group on startup failure or stop the QA script during cleanup.

Findings

  1. P1 Startup failure signals the test group ▶
  2. P1 Cleanup failure skips remaining checks ▶

Summary

The PR reruns the Python-backend ensemble IO test with a 1024-byte CUDA memory pool and checks the server log for pinned-memory fallback.

  • The test also checks output equality through the existing ensemble IO test.
  • No new findings were identified since the previous review.

Reviews (2) · Last reviewed commit: "Merge branch 'main' into mwittwer/python..."

fi
set -e

kill $SERVER_PID

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Startup failure signals the test group
If the server fails to start, run_server leaves SERVER_PID=0. This block records the failure but continues to kill $SERVER_PID, so kill 0 signals the test's entire process group instead of a server. That can terminate the parent QA run. Skip server cleanup when startup fails.

Comment on lines +131 to +132
kill $SERVER_PID
wait $SERVER_PID

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Cleanup failure skips remaining checks
Cleanup runs under set -e. If the server has already exited, kill fails; if it exits with a nonzero status, wait fails. Either failure ends the script before it checks the fallback log or runs the remaining IO subtests. Handle cleanup failures without stopping those checks.

@mattwittwer mattwittwer changed the title draft: test: CUDA memory pool fallback in python backend test: CUDA memory pool fallback in python backend Oct 5, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

1 participant