Repository navigation
mlir: add gpu_block and gpu_thread, gpu_lane, gpu_warp primitive - #100
Merged
Merged
Conversation
yadej
force-pushed
the
dev/rcesista/add-gpu-primitive
branch
from
July 27, 2026 08:08
93e70a1 to
e549085
Compare
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
yadej
force-pushed
the
dev/rcesista/add-gpu-primitive
branch
from
August 4, 2026 07:59
e549085 to
b914e5f
Compare
yadej
force-pushed
the
dev/rcesista/add-gpu-primitive
branch
from
September 8, 2026 07:39
b914e5f to
96d17ee
Compare
yadej
force-pushed
the
dev/rcesista/add-gpu-primitive
branch
2 times, most recently
from
September 21, 2026 15:06
ef7c848 to
9a52f78
Compare
- Usage of TileForAll with gpu mapping - Fusion of gpu mapping for correct IR
And fix some primitive and type problem
…dify some test Fix annotation the same with gpu_block
… gpu and regen some test
yadej
force-pushed
the
dev/rcesista/add-gpu-primitive
branch
from
September 22, 2026 11:13
9a52f78 to
66dd8d1
Compare
Fix some tests on gpu + add schedule to nvgpu test
yadej
force-pushed
the
dev/rcesista/add-gpu-primitive
branch
from
September 22, 2026 11:57
66dd8d1 to
b6973f7
Compare
yadej
marked this pull request as ready for review
September 22, 2026 12:22
…that we don't put loop to several gpu_primitive
guillon
approved these changes
Sep 23, 2026
guillon
left a comment
Member
There was a problem hiding this comment.
Great!
In the description you mention "Also should work with descript", though you implemented it, didn't you?
Contributor
Author
|
Yeah it is implemented since there is descript test on that. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
The current version of appointing gpu block id or thread id is not very good.
So to be more flexible we should be able to choose what GPU block or thread dimension we want a loop to be.
Description
Add 4 primitives gpu_block and gpu_thread, gpu_lane, and gpu_warp to the scheduler.
Each primitive accepts at most a list of 3 elements each mapped to a certain dimension.
It works with memref and tensor. Also work with descript.
Only works on CUDA cores.
Limitation
The mapping of a thread can only be done on a loop that is a tile.
Doesn't work with splitting yet.
TODO:
Next Step: