-
Notifications
You must be signed in to change notification settings - Fork 817
Pull requests: NVIDIA/TransformerEngine
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[PyT] Guard THD learnable dSink on older cuDNN
#3470
opened Sep 3, 2026 by
KshitijLakhani
Collaborator
Loading…
1 of 13 tasks
[PyTorch] Fix FP8 illegal memory access in single-process multi-GPU execution
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3469
opened Sep 3, 2026 by
SuperGoodGame
Loading…
Add opt-in rowwise-only quantized primary weights
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3468
opened Sep 3, 2026 by
xiuhu17
Contributor
Loading…
8 of 13 tasks
[Common] row-scaled nvfp4 path: add single-launch group fused amax
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3467
opened Sep 3, 2026 by
cael-ling
Contributor
Loading…
2 of 13 tasks
[PyTorch] Fix: Resolve EP symm-mem window offset for both old and new torch version
#3466
opened Sep 2, 2026 by
phu0ngng
Collaborator
Loading…
7 of 13 tasks
[JAX] Make MoEBlock aware of Dense TP axes
#3462
opened Sep 2, 2026 by
jberchtold-nvidia
Collaborator
•
Draft
8 of 13 tasks
[Common] Group NVFP4 Quantize Kernels
#3458
opened Sep 1, 2026 by
Oleg-Goncharov
Collaborator
Loading…
9 of 13 tasks
[JAX] Add SiTU-GLU support to JAX
#3457
opened Sep 1, 2026 by
jberchtold-nvidia
Collaborator
•
Draft
8 of 13 tasks
[PyTorch] Schedule delayed-scaling updates after backward
#3456
opened Sep 1, 2026 by
pggPL
Collaborator
Loading…
9 of 13 tasks
[Common] row-scaled nvfp4 path: fuse row/col amax into a single TMA-tiled kernel
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3454
opened Sep 1, 2026 by
cael-ling
Contributor
Loading…
2 of 13 tasks
[JAX] Add sqrtsoftplus router score function
#3448
opened Aug 31, 2026 by
jberchtold-nvidia
Collaborator
Loading…
8 of 13 tasks
[JAX] Attention support for explicit sharding
#3446
opened Aug 31, 2026 by
jberchtold-nvidia
Collaborator
•
Draft
8 of 13 tasks
Relax runtime checks for activation recompute into Warnings
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3436
opened Aug 28, 2026 by
ghadiaravi13
Contributor
•
Draft
1 of 13 tasks
[PyTorch] Reduce CUDA graph memory retention
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3427
opened Aug 26, 2026 by
buptzyb
Contributor
Loading…
[PyTorch] Fix mutable QB bounds in CUDA graphs
org-contribution
#3426
opened Aug 26, 2026 by
harryzhou2000
Contributor
Loading…
Support paged stashing for GroupedLinear activations
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3423
opened Aug 25, 2026 by
lhb8125
Contributor
Loading…
[PyTorch] Allow CP P2P transport group overrides
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3420
opened Aug 24, 2026 by
xiaoyao0115
•
Draft
[Common] Fix TMA synchronization in quantization kernels
#3417
opened Aug 22, 2026 by
Oleg-Goncharov
Collaborator
Loading…
6 of 13 tasks
Ring Attention: free unused kv comm buffers
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3411
opened Aug 21, 2026 by
francesco-bertolotti
Contributor
Loading…
6 of 13 tasks
Previous Next
ProTip!
Follow long discussions with comments:>50.