draft: fix: GPU tensor allocation fall back to host memory - #457
mattwittwer wants to merge 1 commit into
Conversation
|
| if (pb_memory->MemoryType() == TRITONSERVER_MEMORY_CPU || | ||
| pb_memory->MemoryType() == TRITONSERVER_MEMORY_CPU_PINNED) { |
There was a problem hiding this comment.
Pinned fallback lacks tests The new CPU_PINNED branch copies GPU output through a shared-memory buffer before filling Triton’s output buffer, but there is no regression test for that transfer in either ordinary or decoupled responses. A test that forces a pinned-host output and checks the returned bytes in both modes would help catch a future change that returns incorrect output.
Knowledge Base Used:
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
No description provided.