test: CUDA memory pool fallback in python backend - #8992
mattwittwer wants to merge 2 commits into
Conversation
…s back to pinned memory
|
| fi | ||
| set -e | ||
|
|
||
| kill $SERVER_PID |
There was a problem hiding this comment.
Startup failure signals the test group
If the server fails to start, run_server leaves SERVER_PID=0. This block records the failure but continues to kill $SERVER_PID, so kill 0 signals the test's entire process group instead of a server. That can terminate the parent QA run. Skip server cleanup when startup fails.
| kill $SERVER_PID | ||
| wait $SERVER_PID |
There was a problem hiding this comment.
Cleanup failure skips remaining checks
Cleanup runs under set -e. If the server has already exited, kill fails; if it exits with a nonzero status, wait fails. Either failure ends the script before it checks the fallback log or runs the remaining IO subtests. Handle cleanup failures without stopping those checks.
What does the PR do?
Adds regression coverage for silent output corruption in the Python backend when the CUDA memory pool is exhausted and GPU output tensors fall back to pinned host memory (#7148). The runtime fix is
python_backend#457.The existing
IOTest.test_ensemble_ioalready drives theensemble_iopipeline of three chaineddlpack_io_identityPython models, with per-request flags choosing which stage emits its output as a GPU tensor, and asserts exact equality against a 1000 x FP32 input. This change re-runs that test with--cuda-memory-pool-byte-size=0:1024, so every 4000-byte GPU output overflows the pool and the ensemble's response allocator falls back to pinned memory for each of them. The block also fails if the server log does not contain the core'sfalling back to pinned system memorywarning, so it cannot pass without exercising the fallback path. No new models or Python code.Checklist
<commit_type>: <Title>Commit Type:
Related PRs:
test guards; the new block fails until it is in the test container)
Where should the reviewer start?
qa/L0_backend_python/io/test.sh— the new blockIOTest.test_ensemble_io with GPU outputs falling back to pinned memory: model setup mirrors the existingdefaulttrial,SERVER_ARGSadds the 1024-byte CUDA pool, and the post-rungrepon the server log guards against a vacuous pass.Test plan:
Caveats:
Background
Related Issues: