-
Notifications
You must be signed in to change notification settings - Fork 780
Pull requests: NVIDIA/TransformerEngine
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[Add] Add FlashAttention 4 context parallel support.
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3259
opened Jul 25, 2026 by
Baibaifan
Contributor
Loading…
test(pytorch): cover QuantizedTensor view NotImplementedError
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3257
opened Jul 25, 2026 by
andrewwhitecdw
Loading…
fix(pytorch): add missing f-string prefixes to error messages
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3256
opened Jul 25, 2026 by
andrewwhitecdw
Loading…
Acquire the GIL in lazy init_extension
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3255
opened Jul 24, 2026 by
xiuhu17
Contributor
Loading…
5 tasks done
[CI] Pin JAX image to 2026-07-21
#3250
opened Jul 24, 2026 by
fheinecke
Collaborator
Loading…
4 of 13 tasks
Opt in trusted FA CI checkpoints to pickle loading
#3247
opened Jul 24, 2026 by
sudhakarsingh27
Member
Loading…
6 of 9 tasks
[PyTorch] Allow multi-version Flash Attention tests to use checkpoints with pickles
2.18
testing
Improvements to tests or testing infrastructure
#3245
opened Jul 23, 2026 by
timmoon10
Member
Loading…
8 of 14 tasks
[Pytorch] Add support for row-wise quanted input for grouped gemm
#3244
opened Jul 23, 2026 by
YangFei1990
Collaborator
•
Draft
8 of 13 tasks
[PyTorch] Fix NCCL communicator init in cuSOLVERMp context creation
#3240
opened Jul 22, 2026 by
vcherepanov-nv
Collaborator
Loading…
5 of 13 tasks
Activation + GroupedLinear Fusion for MOE and other MOE optimizations
2.18
#3238
opened Jul 22, 2026 by
vthumbe1503
Collaborator
Loading…
13 tasks
[All] Bump minimum supported cuDNN version to 9.11
#3236
opened Jul 22, 2026 by
cyanguwa
Collaborator
Loading…
8 of 13 tasks
Add stream-ordered CP gradient return primitive
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3235
opened Jul 22, 2026 by
foraxe
Loading…
5 of 13 tasks
[PyTorch] Optionally release columnwise copy of frozen FP8 block-scaled weights after dgrad
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
[CI] Fix old NVCC dependency in GH CI
#3230
opened Jul 21, 2026 by
fheinecke
Collaborator
Loading…
5 of 13 tasks
[PyTorch] NCCL EP eager mode and drop-on-overflow policy
#3229
opened Jul 21, 2026 by
phu0ngng
Collaborator
Loading…
8 of 13 tasks
[CI] Minor CMake refactoring
#3227
opened Jul 21, 2026 by
fheinecke
Collaborator
Loading…
5 of 13 tasks
Single grouped weight fixes
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3225
opened Jul 21, 2026 by
CarlosGomes98
Contributor
Loading…
13 tasks
Improve device-init grouped linear module with single grouped weight support
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3224
opened Jul 20, 2026 by
zhongbozhu
Collaborator
Loading…
13 tasks
[Common] Experimental CuTeDSL MXFP4 backend
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
[PyTorch] Add experimental packed-contiguous THD all-gather CP
#3221
opened Jul 18, 2026 by
sudhakarsingh27
Member
•
Draft
Work around intermittent SM120 FP8 gradient corruption in RTC cast-transpose
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3215
opened Jul 15, 2026 by
AlbertYang514
Loading…
Previous Next
ProTip!
Mix and match filters to narrow down what you’re looking for.