-
Notifications
You must be signed in to change notification settings - Fork 2.7k
All issues
Issue creation is restricted in this repository
- #15044 · laikhtewari opened
on Jun 6, 2026 1 - #3148 · juney-nvidia opened
on Mar 29, 2025 5 - #3124 · juney-nvidia opened
on Mar 27, 2025 11
Issues
is:issue state:open
is:issue state:open
Search results
[Bug] configure_cpu_affinity() is not idempotent: a second call un-pins every thread
Inference runtime<NV>General operational aspects of TRTLLM execution not in other categories.<NV>General operational aspects of TRTLLM execution not in other categories.Status: Open.#18848 In NVIDIA/TensorRT-LLM;[Bug] Worker CPU affinity is applied process-wide in shared-process deployments and never restored
Inference runtime<NV>General operational aspects of TRTLLM execution not in other categories.<NV>General operational aspects of TRTLLM execution not in other categories.Status: Open.#18847 In NVIDIA/TensorRT-LLM;fix: Pass NVML CC settings by reference in release/1.2.1
Inference runtime<NV>General operational aspects of TRTLLM execution not in other categories.<NV>General operational aspects of TRTLLM execution not in other categories.Status: Open.#18816 In NVIDIA/TensorRT-LLM;[Performance]: Defer dynamic NVFP4 scale finalization for fused GELU MLPs
Customized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.Status: Open.#18787 In NVIDIA/TensorRT-LLM;Docs: legacy performance-tuning guide has broken image links and dead anchor; kv_cache_manager.md links removed code
Doc<NV>TRTLLM's textual/illustrative materials: API refs, guides, tutorials. Improvement & clarity.<NV>TRTLLM's textual/illustrative materials: API refs, guides, tutorials. Improvement & clarity.Status: Open.#18777 In NVIDIA/TensorRT-LLM;seq_slot_manager.py uses undefined name 'logger'
Disaggregated serving<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.Pytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#18776 In NVIDIA/TensorRT-LLM;[Bug]: [CUDA] cute::make_int_sequence<cute::size(...)> fails in dependent template contexts
bugSomething isn't workingSomething isn't workingCustomized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.Windows<NV>Windows operating system specific issues and compatibility problems<NV>Windows operating system specific issues and compatibility problemsStatus: Open.#18775 In NVIDIA/TensorRT-LLM;[Bug]: Responses history trimming hangs on over-capacity input without a completed turn
Inference runtime<NV>General operational aspects of TRTLLM execution not in other categories.<NV>General operational aspects of TRTLLM execution not in other categories.Status: Open.#18759 In NVIDIA/TensorRT-LLM;[Bug]: Cosmos3 sequence-parallel path attends zero-padded text K/V when CFG prompts have unequal lengths (cfg_size=1, ulysses/CP > 1)
bugSomething isn't workingSomething isn't workingStatus: Open.#18687 In NVIDIA/TensorRT-LLM;[Feature]: Load SGLang-format W4AFP8 checkpoints (quant_method: w4afp8, e.g. PhalaCloud GLM-5.x) natively
Low PrecisionLower-precision formats (INT8/INT4/FP8) for TRTLLM quantization (AWQ, GPTQ).Lower-precision formats (INT8/INT4/FP8) for TRTLLM quantization (AWQ, GPTQ).Status: Open.#18664 In NVIDIA/TensorRT-LLM;[Bug]: trtllm-serve stays alive returning 503 forever after a rank-crash hard kill (no exit for supervisors)
Inference runtime<NV>General operational aspects of TRTLLM execution not in other categories.<NV>General operational aspects of TRTLLM execution not in other categories.Pytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#18663 In NVIDIA/TensorRT-LLM;[Bug]: MTP accepts no draft tokens with use_kv_cache_manager_v2=false on DSA models (acceptance length 1.0)
Speculative Decoding<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafter<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafterStatus: Open.#18662 In NVIDIA/TensorRT-LLM;