Repository navigation
tvm: preserve elementwise ranks for fusion - #146
Conversation
bb1f778 to
1f17feb
Compare
|
@liamsemeria I basically implemented what I discussed some time ago, i.e. generating elementwise with full rank dimensions instead of reshaping it, giving axis names i to z, i.e. i, j, k, l, ... up to the rank. This avoids reshapes in the middle which were fused instead of the actual relu op. Is it in line with what you did in MLIR? Should I also look at updating the MLIR side? |
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
@guillon yeah this matches how I did it in MLIR, based off of what I'm seeing in these tvm tests. |
liamsemeria
left a comment
There was a problem hiding this comment.
I really like that you added the test graphs.
LGTM.
One tiny thing that I noticed was for the loop dim names in the relu, the dim names are i0 and i1. If its easy to do could we change the dim names to the ones that the scheduled op uses? (i, j)
Fixes this for relu and pad/unpad. Now axes names appear as declared in TVMOps. |
2df640f to
9881450
Compare
Motivation
TVM elementwise operators were flattened to a single axis by inserting hidden
reshapes. Consumer or producer fusion therefore targeted a reshape rather than
the intended elementwise operation, making fusion ineffective.
Description
Generate TVM ReLU directly with the rank and dimensions propagated from its
input tensor type. This removes the intermediate reshapes, exposes one parallel
scheduling axis per tensor dimension, and allows producer and consumer fusion
to target the ReLU block directly.
Update the matmul and convolution FileCheck tests and add coverage for fusing
ReLU under different matmul loop levels and through the descriptor scheduler.
Additionally, add representative ResNet18 and YOLO9000 multi-node graph
fixtures and allow
loop-exploreto schedule a named graph node with--node.Also, updated TVM Ops genration for relu and pd to use explicit axes names such that IR reflects chooses dims names.
Added tile10d (
PFPCWRPRP) strategy which perform tentative consumer fusion.Commits
tests: fix wrongly ordered conv2d axestvm: implement proper fusion, discarding reshapestests: add multi node test graphs from yolo/resnetexplore: support scheduling a selected graph nodetvm: generate explicit axes names and fix pad2d usagestrategies: add tile10d strategy with consumer fusionTesting
relu -> ...and consumer fusiontargets
... -> relu.pytest -q tests/pytest/tvm— 14 passedDiscussion
Explicit axis coalescing remains possible future scheduling work. This change
intentionally exposes the propagated tensor rank instead of adding implicit
reshapes.