-
Notifications
You must be signed in to change notification settings - Fork 2.7k
All issues
Issue creation is restricted in this repository
- #15044 · laikhtewari opened
on Jun 6, 2026 1 - #3148 · juney-nvidia opened
on Mar 29, 2025 5 - #3124 · juney-nvidia opened
on Mar 27, 2025 11
Issues
is:issue state:open
is:issue state:open
Search results
[Bug]: -a 107-real is silently folded to 100, no SM107 code is generated
Infra<NV>automated tests, build checks, github actions, system stability & efficiency.<NV>automated tests, build checks, github actions, system stability & efficiency.Status: Open.#19067 In NVIDIA/TensorRT-LLM;[Bug]: construct_if_true in moe_gemm_tma_ws_launcher.inl fails to compile at C++20
Customized kernels<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.<NV>Specialized/modified CUDA kernels in TRTLLM for LLM ops, beyond standard TRT. Dev & perf.Status: Open.#19063 In NVIDIA/TensorRT-LLM;[Bug]: C++17 build cannot compile torch 2.12 headers with nvcc, but C++20 breaks float2(fp8) functional casts
Pytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#19059 In NVIDIA/TensorRT-LLM;[Feature]: Remove the FA4 install workaround after the PyTorch 26.05 upgrade
feature requestNew feature or request. This includes new model, dtype, functionality supportNew feature or request. This includes new model, dtype, functionality supportInfra<NV>automated tests, build checks, github actions, system stability & efficiency.<NV>automated tests, build checks, github actions, system stability & efficiency.Status: Open.#19033 In NVIDIA/TensorRT-LLM;[Bug]: Multimodal metadata failures terminate the PyTorch executor instead of failing the request
MultimodalLabel for issues & PRs regarding Multimodal related objectsLabel for issues & PRs regarding Multimodal related objectsPytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#18971 In NVIDIA/TensorRT-LLM;[Bug]: EAGLE3 dynamic-tree decoding emits tokens outside the requested grammar
Speculative Decoding<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafter<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafterStatus: Open.#18892 In NVIDIA/TensorRT-LLM;- Status: Open.#18891 In NVIDIA/TensorRT-LLM;
[Bug] Hybrid-Mamba KV cache manager is constructed without vocab_size; the first multimodal request hard-kills every rank
KV-Cache Managementkv-cache management for efficient LLM inferencekv-cache management for efficient LLM inferenceMultimodalLabel for issues & PRs regarding Multimodal related objectsLabel for issues & PRs regarding Multimodal related objectsPytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#18849 In NVIDIA/TensorRT-LLM;[Bug] configure_cpu_affinity() is not idempotent: a second call un-pins every thread
Inference runtime<NV>General operational aspects of TRTLLM execution not in other categories.<NV>General operational aspects of TRTLLM execution not in other categories.Status: Open.#18848 In NVIDIA/TensorRT-LLM;[Bug] Worker CPU affinity is applied process-wide in shared-process deployments and never restored
Inference runtime<NV>General operational aspects of TRTLLM execution not in other categories.<NV>General operational aspects of TRTLLM execution not in other categories.Status: Open.#18847 In NVIDIA/TensorRT-LLM;fix: Pass NVML CC settings by reference in release/1.2.1
Inference runtime<NV>General operational aspects of TRTLLM execution not in other categories.<NV>General operational aspects of TRTLLM execution not in other categories.Status: Open.#18816 In NVIDIA/TensorRT-LLM;Docs: legacy performance-tuning guide has broken image links and dead anchor; kv_cache_manager.md links removed code
Doc<NV>TRTLLM's textual/illustrative materials: API refs, guides, tutorials. Improvement & clarity.<NV>TRTLLM's textual/illustrative materials: API refs, guides, tutorials. Improvement & clarity.Status: Open.#18777 In NVIDIA/TensorRT-LLM;