Skip to content

[ISel] Improve clmul fallback implementation for i32 and i64 - #203727

Merged
folkertdev merged 10 commits into
llvm:mainfrom
folkertdev:clmul-better-fallback
Jun 18, 2026
Merged

folkertdev merged 10 commits into
llvm:mainfrom
folkertdev:clmul-better-fallback

Conversation

@folkertdev

@folkertdev folkertdev commented Jun 13, 2026

Copy link
Copy Markdown
Contributor

Improve the clmul fallback implementation for i32..=i64.

The new approach is "multiplication with holes", based on bearssl source, https://www.bearssl.org/constanttime.html#ghash-for-gcm, and the polyval crate. https://www.bearssl.org/constanttime.html#ghash-for-gcm explains the idea.

Future work is

  • extend the fallback to smaller integer types. These can use a smaller hole (the wider hole works but uses more instructions than the naive fallback)
  • investigate how to extend this approach to wider integers. We can use wider holes, but a recursive definition defined in terms of smaller integer types might be better. https://timtaubert.de/blog/2017/06/verified-binary-multiplication-for-ghash/ has some good info.

I've tested this locally with a fuzzer against the current LLVM implementation (via the rust standard library) and pclmulqdq.

CC #203694
CC: @RKSimon, @topperc, @eisenwave, @xarkenz, @artagnon.

@llvmorg-github-actions

llvmorg-github-actions Bot commented Jun 13, 2026

Copy link
Copy Markdown

@llvm/pr-subscribers-llvm-analysis
@llvm/pr-subscribers-backend-risc-v
@llvm/pr-subscribers-backend-amdgpu
@llvm/pr-subscribers-backend-powerpc
@llvm/pr-subscribers-llvm-selectiondag

@llvm/pr-subscribers-backend-x86

Author: Folkert de Vries (folkertdev)

Changes

Well I think I've nerdsniped myself, at least for the simple case that can be based on multiplication with holes (via bearssl source, https://www.bearssl.org/constanttime.html#ghash-for-gcm, and the polyval crate).

That still leaves the widening case though, for which this blog post has some good info https://timtaubert.de/blog/2017/06/verified-binary-multiplication-for-ghash/. Also for 8-bit and 16-bit inputs this approach can be generalized but with smaller holes and fewer multiplications. I'll leave that as future work too.

I've tested this locally with a fuzzer against the current LLVM implementation (via the rust standard library) and pclmulqdq. So, I think it's correct but I'm not sure how to show that in a convincing way.

I know this needs tests updated for more targets, but I'd like CI to tell me which ones.

CC #203694
CC: @RKSimon, @topperc, @eisenwave, @xarkenz, @artagnon.


Patch is 279.90 KiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/203727.diff

4 Files Affected:

  • (modified) llvm/include/llvm/CodeGen/BasicTTIImpl.h (+20-5)
  • (modified) llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp (+51)
  • (modified) llvm/test/CodeGen/AArch64/clmul.ll (+506-2292)
  • (modified) llvm/test/CodeGen/X86/clmul.ll (+969-2739)
diff --git a/llvm/include/llvm/CodeGen/BasicTTIImpl.h b/llvm/include/llvm/CodeGen/BasicTTIImpl.h
index e4b6bf51c7a4e..b11057695c2bd 100644
--- a/llvm/include/llvm/CodeGen/BasicTTIImpl.h
+++ b/llvm/include/llvm/CodeGen/BasicTTIImpl.h
@@ -3090,18 +3090,33 @@ class BasicTTIImplBase : public TargetTransformInfoImplCRTPBase<T> {
     case Intrinsic::clmul: {
       // This cost model should match the expansion in
       // TargetLowering::expandCLMUL.
-      InstructionCost PerBitCostMul =
-          thisT()->getArithmeticInstrCost(Instruction::And, RetTy, CostKind) +
-          thisT()->getArithmeticInstrCost(Instruction::Mul, RetTy, CostKind) +
+      unsigned BW = RetTy->getScalarSizeInBits();
+      InstructionCost AndCost =
+          thisT()->getArithmeticInstrCost(Instruction::And, RetTy, CostKind);
+      InstructionCost OrCost =
+          thisT()->getArithmeticInstrCost(Instruction::Or, RetTy, CostKind);
+      InstructionCost XorCost =
           thisT()->getArithmeticInstrCost(Instruction::Xor, RetTy, CostKind);
+      InstructionCost MulCost =
+          thisT()->getArithmeticInstrCost(Instruction::Mul, RetTy, CostKind);
+
+      // When the multiplication with holes approach is used, that emits 16 MULs, 8
+      // + 4 ANDs, 12 XORs and 3 ORs.
+      if (BW >= 32 && BW <= 64 &&
+          TLI->isOperationLegalOrCustom(ISD::MUL,
+                                        TLI->getValueType(DL, RetTy))) {
+        return 16 * MulCost + 12 * AndCost + 12 * XorCost + 3 * OrCost;
+      }
+
+      InstructionCost PerBitCostMul = AndCost + MulCost + XorCost;
       InstructionCost PerBitCostBittest =
-          thisT()->getArithmeticInstrCost(Instruction::And, RetTy, CostKind) +
+          AndCost +
           thisT()->getCmpSelInstrCost(BinaryOperator::Select, RetTy, RetTy,
                                       ICmpInst::BAD_ICMP_PREDICATE, CostKind) +
           thisT()->getCmpSelInstrCost(Instruction::ICmp, RetTy, RetTy,
                                       ICmpInst::ICMP_NE, CostKind);
       InstructionCost PerBitCost = std::min(PerBitCostMul, PerBitCostBittest);
-      return RetTy->getScalarSizeInBits() * PerBitCost;
+      return BW * PerBitCost;
     }
     default:
       break;
diff --git a/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp b/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
index 95e03a419fb43..175a41aa502dc 100644
--- a/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
+++ b/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
@@ -8899,6 +8899,57 @@ SDValue TargetLowering::expandCLMUL(SDNode *Node, SelectionDAG &DAG) const {
     // calculation in BasicTTIImpl::getTypeBasedIntrinsicInstrCost for
     // Intrinsic::clmul.
 
+    // Strategy 4: multiplication with holes. 
+    //
+    // Uses "holes" (sequences of zeroes) to avoid carry spilling. When carries
+    // do occur, they wind up in a "hole" and are subsequently masked out of the
+    // result.
+    //
+    // A hole of 3 bits is optimal for 32-bit and 64-bit inputs. 128-bit
+    // integers need a larger hole, and for smaller integers the fallback below
+    // is more efficient.
+    //
+    // Based on bmul64 in bearssl and bmul in the rust polyval crate.
+    if (BW >= 32 && BW <= 64 &&
+        isOperationLegalOrCustom(ISD::MUL, getTypeToTransformTo(Ctx, VT))) {
+
+      // Set every fourth bit of each nibble, equivalent to 0b100010001...0001.
+      APInt MaskVal(BW, 0);
+      for (unsigned i = 0; i < BW; i += 4)
+        MaskVal.setBit(i);
+
+      // Create versions of X and Y that keep only the first, second, third, or
+      // fourth bit of every nibble.
+      SDValue M[4], Xp[4], Yp[4];
+      for (unsigned i = 0; i < 4; ++i) {
+        M[i] = DAG.getConstant(MaskVal.shl(i), DL, VT);
+        Xp[i] = DAG.getNode(ISD::AND, DL, VT, X, M[i]);
+        Yp[i] = DAG.getNode(ISD::AND, DL, VT, Y, M[i]);
+      }
+
+      // Codegens these expressions (16 multiplications):
+      //
+      // z0 = (x0 * y0) ^ (x1 * y3) ^ (x2 * y2) ^ (x3 * y1);
+      // z1 = (x0 * y1) ^ (x1 * y0) ^ (x2 * y3) ^ (x3 * y2);
+      // z2 = (x0 * y2) ^ (x1 * y1) ^ (x2 * y0) ^ (x3 * y3);
+      // z3 = (x0 * y3) ^ (x1 * y2) ^ (x2 * y1) ^ (x3 * y0);
+      SDValue Res;
+      for (unsigned i = 0; i < 4; ++i) {
+        SDValue Zi;
+        for (unsigned A = 0; A < 4; ++A) {
+          unsigned B = (i + 4 - A) % 4;
+          SDValue P = DAG.getNode(ISD::MUL, DL, VT, Xp[A], Yp[B]);
+          Zi = Zi ? DAG.getNode(ISD::XOR, DL, VT, Zi, P) : P;
+        }
+
+        // Keep only the bits belonging to this iteration, and bitwise or it all
+        // together.
+        Zi = DAG.getNode(ISD::AND, DL, VT, Zi, M[i]);
+        Res = Res ? DAG.getNode(ISD::OR, DL, VT, Res, Zi) : Zi;
+      }
+      return Res;
+    }
+
     EVT SetCCVT = getSetCCResultType(DAG.getDataLayout(), Ctx, VT);
 
     SDValue Res = DAG.getConstant(0, DL, VT);
diff --git a/llvm/test/CodeGen/AArch64/clmul.ll b/llvm/test/CodeGen/AArch64/clmul.ll
index bb42e411c4f99..bfa4b73b4e677 100644
--- a/llvm/test/CodeGen/AArch64/clmul.ll
+++ b/llvm/test/CodeGen/AArch64/clmul.ll
@@ -80,101 +80,49 @@ define i16 @clmul_i16(i16 %x, i16 %y) {
 define i32 @clmul_i32(i32 %x, i32 %y) {
 ; CHECK-NEON-LABEL: clmul_i32:
 ; CHECK-NEON:       // %bb.0:
-; CHECK-NEON-NEXT:    and w8, w1, #0x2
-; CHECK-NEON-NEXT:    and w9, w1, #0x1
-; CHECK-NEON-NEXT:    and w10, w1, #0x4
-; CHECK-NEON-NEXT:    mul w8, w0, w8
-; CHECK-NEON-NEXT:    and w11, w1, #0x8
-; CHECK-NEON-NEXT:    and w12, w1, #0x10
-; CHECK-NEON-NEXT:    mul w9, w0, w9
-; CHECK-NEON-NEXT:    and w13, w1, #0x20
-; CHECK-NEON-NEXT:    and w14, w1, #0x40
-; CHECK-NEON-NEXT:    mul w10, w0, w10
-; CHECK-NEON-NEXT:    and w2, w1, #0x800
-; CHECK-NEON-NEXT:    and w15, w1, #0x80
-; CHECK-NEON-NEXT:    mul w11, w0, w11
-; CHECK-NEON-NEXT:    and w16, w1, #0x100
-; CHECK-NEON-NEXT:    and w17, w1, #0x200
-; CHECK-NEON-NEXT:    mul w12, w0, w12
-; CHECK-NEON-NEXT:    eor w8, w9, w8
-; CHECK-NEON-NEXT:    and w9, w1, #0x1000
-; CHECK-NEON-NEXT:    mul w13, w0, w13
-; CHECK-NEON-NEXT:    and w18, w1, #0x400
-; CHECK-NEON-NEXT:    mul w14, w0, w14
-; CHECK-NEON-NEXT:    eor w10, w10, w11
-; CHECK-NEON-NEXT:    and w11, w1, #0x2000
-; CHECK-NEON-NEXT:    mul w2, w0, w2
-; CHECK-NEON-NEXT:    eor w8, w8, w10
-; CHECK-NEON-NEXT:    and w10, w1, #0x4000
-; CHECK-NEON-NEXT:    mul w9, w0, w9
-; CHECK-NEON-NEXT:    eor w12, w12, w13
-; CHECK-NEON-NEXT:    and w13, w1, #0x8000
-; CHECK-NEON-NEXT:    mul w15, w0, w15
-; CHECK-NEON-NEXT:    eor w12, w12, w14
-; CHECK-NEON-NEXT:    and w14, w1, #0x10000
-; CHECK-NEON-NEXT:    mul w16, w0, w16
-; CHECK-NEON-NEXT:    eor w8, w8, w12
-; CHECK-NEON-NEXT:    and w12, w1, #0x20000
-; CHECK-NEON-NEXT:    mul w11, w0, w11
-; CHECK-NEON-NEXT:    eor w9, w2, w9
-; CHECK-NEON-NEXT:    and w2, w1, #0x400000
-; CHECK-NEON-NEXT:    mul w17, w0, w17
-; CHECK-NEON-NEXT:    mul w10, w0, w10
-; CHECK-NEON-NEXT:    eor w15, w15, w16
-; CHECK-NEON-NEXT:    and w16, w1, #0x40000
-; CHECK-NEON-NEXT:    mul w13, w0, w13
-; CHECK-NEON-NEXT:    eor w9, w9, w11
-; CHECK-NEON-NEXT:    and w11, w1, #0x800000
-; CHECK-NEON-NEXT:    mul w18, w0, w18
-; CHECK-NEON-NEXT:    eor w15, w15, w17
-; CHECK-NEON-NEXT:    and w17, w1, #0x80000
-; CHECK-NEON-NEXT:    mul w14, w0, w14
-; CHECK-NEON-NEXT:    eor w9, w9, w10
-; CHECK-NEON-NEXT:    and w10, w1, #0x1000000
-; CHECK-NEON-NEXT:    mul w12, w0, w12
-; CHECK-NEON-NEXT:    eor w9, w9, w13
-; CHECK-NEON-NEXT:    and w13, w1, #0x2000000
-; CHECK-NEON-NEXT:    mul w16, w0, w16
-; CHECK-NEON-NEXT:    eor w15, w15, w18
-; CHECK-NEON-NEXT:    and w18, w1, #0x100000
-; CHECK-NEON-NEXT:    mul w2, w0, w2
-; CHECK-NEON-NEXT:    eor w8, w8, w15
-; CHECK-NEON-NEXT:    and w15, w1, #0x200000
-; CHECK-NEON-NEXT:    mul w11, w0, w11
+; CHECK-NEON-NEXT:    and w8, w1, #0x11111111
+; CHECK-NEON-NEXT:    and w9, w0, #0x22222222
+; CHECK-NEON-NEXT:    and w10, w1, #0x22222222
+; CHECK-NEON-NEXT:    and w11, w0, #0x11111111
+; CHECK-NEON-NEXT:    and w13, w1, #0x88888888
+; CHECK-NEON-NEXT:    and w15, w0, #0x44444444
+; CHECK-NEON-NEXT:    and w17, w1, #0x44444444
+; CHECK-NEON-NEXT:    and w18, w0, #0x88888888
+; CHECK-NEON-NEXT:    mul w12, w9, w8
+; CHECK-NEON-NEXT:    mul w14, w11, w10
+; CHECK-NEON-NEXT:    mul w16, w15, w13
+; CHECK-NEON-NEXT:    mul w0, w18, w17
+; CHECK-NEON-NEXT:    mul w1, w9, w13
 ; CHECK-NEON-NEXT:    eor w12, w14, w12
-; CHECK-NEON-NEXT:    and w14, w1, #0x4000000
-; CHECK-NEON-NEXT:    mul w17, w0, w17
-; CHECK-NEON-NEXT:    eor w12, w12, w16
-; CHECK-NEON-NEXT:    and w16, w1, #0x8000000
-; CHECK-NEON-NEXT:    mul w10, w0, w10
-; CHECK-NEON-NEXT:    eor w8, w8, w9
-; CHECK-NEON-NEXT:    mul w13, w0, w13
-; CHECK-NEON-NEXT:    eor w11, w2, w11
-; CHECK-NEON-NEXT:    and w2, w1, #0x20000000
-; CHECK-NEON-NEXT:    mul w18, w0, w18
-; CHECK-NEON-NEXT:    eor w12, w12, w17
-; CHECK-NEON-NEXT:    and w17, w1, #0x10000000
-; CHECK-NEON-NEXT:    mul w14, w0, w14
-; CHECK-NEON-NEXT:    eor w10, w11, w10
-; CHECK-NEON-NEXT:    and w11, w1, #0x40000000
-; CHECK-NEON-NEXT:    mul w15, w0, w15
-; CHECK-NEON-NEXT:    eor w10, w10, w13
-; CHECK-NEON-NEXT:    and w13, w1, #0x80000000
-; CHECK-NEON-NEXT:    mul w16, w0, w16
-; CHECK-NEON-NEXT:    eor w12, w12, w18
-; CHECK-NEON-NEXT:    mul w17, w0, w17
-; CHECK-NEON-NEXT:    eor w10, w10, w14
-; CHECK-NEON-NEXT:    mul w2, w0, w2
-; CHECK-NEON-NEXT:    eor w9, w12, w15
-; CHECK-NEON-NEXT:    mul w11, w0, w11
-; CHECK-NEON-NEXT:    eor w10, w10, w16
-; CHECK-NEON-NEXT:    eor w8, w8, w9
-; CHECK-NEON-NEXT:    mul w13, w0, w13
-; CHECK-NEON-NEXT:    eor w9, w10, w17
-; CHECK-NEON-NEXT:    eor w8, w8, w9
-; CHECK-NEON-NEXT:    eor w10, w2, w11
-; CHECK-NEON-NEXT:    eor w9, w10, w13
-; CHECK-NEON-NEXT:    eor w0, w8, w9
+; CHECK-NEON-NEXT:    mul w2, w11, w8
+; CHECK-NEON-NEXT:    mul w3, w15, w17
+; CHECK-NEON-NEXT:    eor w14, w16, w0
+; CHECK-NEON-NEXT:    mul w4, w18, w10
+; CHECK-NEON-NEXT:    eor w12, w12, w14
+; CHECK-NEON-NEXT:    mul w5, w9, w10
+; CHECK-NEON-NEXT:    eor w14, w2, w1
+; CHECK-NEON-NEXT:    and w12, w12, #0x22222222
+; CHECK-NEON-NEXT:    mul w9, w9, w17
+; CHECK-NEON-NEXT:    mul w17, w11, w17
+; CHECK-NEON-NEXT:    eor w16, w3, w4
+; CHECK-NEON-NEXT:    mul w10, w15, w10
+; CHECK-NEON-NEXT:    mul w15, w15, w8
+; CHECK-NEON-NEXT:    mul w11, w11, w13
+; CHECK-NEON-NEXT:    eor w17, w17, w5
+; CHECK-NEON-NEXT:    mul w13, w18, w13
+; CHECK-NEON-NEXT:    mul w8, w18, w8
+; CHECK-NEON-NEXT:    eor w9, w11, w9
+; CHECK-NEON-NEXT:    eor w13, w15, w13
+; CHECK-NEON-NEXT:    eor w8, w10, w8
+; CHECK-NEON-NEXT:    eor w10, w14, w16
+; CHECK-NEON-NEXT:    eor w11, w17, w13
+; CHECK-NEON-NEXT:    eor w8, w9, w8
+; CHECK-NEON-NEXT:    and w9, w10, #0x11111111
+; CHECK-NEON-NEXT:    and w10, w11, #0x44444444
+; CHECK-NEON-NEXT:    and w8, w8, #0x88888888
+; CHECK-NEON-NEXT:    orr w9, w9, w12
+; CHECK-NEON-NEXT:    orr w8, w10, w8
+; CHECK-NEON-NEXT:    orr w0, w9, w8
 ; CHECK-NEON-NEXT:    ret
 ;
 ; CHECK-AES-LABEL: clmul_i32:
@@ -191,274 +139,49 @@ define i32 @clmul_i32(i32 %x, i32 %y) {
 define i64 @clmul_i64(i64 %x, i64 %y) {
 ; CHECK-NEON-LABEL: clmul_i64:
 ; CHECK-NEON:       // %bb.0:
-; CHECK-NEON-NEXT:    sub sp, sp, #304
-; CHECK-NEON-NEXT:    stp x29, x30, [sp, #208] // 16-byte Folded Spill
-; CHECK-NEON-NEXT:    stp x28, x27, [sp, #224] // 16-byte Folded Spill
-; CHECK-NEON-NEXT:    stp x26, x25, [sp, #240] // 16-byte Folded Spill
-; CHECK-NEON-NEXT:    stp x24, x23, [sp, #256] // 16-byte Folded Spill
-; CHECK-NEON-NEXT:    stp x22, x21, [sp, #272] // 16-byte Folded Spill
-; CHECK-NEON-NEXT:    stp x20, x19, [sp, #288] // 16-byte Folded Spill
-; CHECK-NEON-NEXT:    .cfi_def_cfa_offset 304
-; CHECK-NEON-NEXT:    .cfi_offset w19, -8
-; CHECK-NEON-NEXT:    .cfi_offset w20, -16
-; CHECK-NEON-NEXT:    .cfi_offset w21, -24
-; CHECK-NEON-NEXT:    .cfi_offset w22, -32
-; CHECK-NEON-NEXT:    .cfi_offset w23, -40
-; CHECK-NEON-NEXT:    .cfi_offset w24, -48
-; CHECK-NEON-NEXT:    .cfi_offset w25, -56
-; CHECK-NEON-NEXT:    .cfi_offset w26, -64
-; CHECK-NEON-NEXT:    .cfi_offset w27, -72
-; CHECK-NEON-NEXT:    .cfi_offset w28, -80
-; CHECK-NEON-NEXT:    .cfi_offset w30, -88
-; CHECK-NEON-NEXT:    .cfi_offset w29, -96
-; CHECK-NEON-NEXT:    and x8, x1, #0x2
-; CHECK-NEON-NEXT:    mul x9, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x1
-; CHECK-NEON-NEXT:    mul x10, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x4
-; CHECK-NEON-NEXT:    mul x11, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x8
-; CHECK-NEON-NEXT:    mul x13, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x10
-; CHECK-NEON-NEXT:    eor x9, x10, x9
-; CHECK-NEON-NEXT:    mul x12, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x20
-; CHECK-NEON-NEXT:    mul x14, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x40
-; CHECK-NEON-NEXT:    eor x10, x11, x13
-; CHECK-NEON-NEXT:    and x11, x1, #0x10000000000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #200] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x80
-; CHECK-NEON-NEXT:    mul x15, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x100
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #160] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x200
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #152] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x400
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #184] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x800
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #192] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x1000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #144] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x2000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #136] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x4000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #176] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x8000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #168] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x10000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #120] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x20000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #80] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x40000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #72] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x80000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #104] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x100000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #96] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x200000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #128] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x400000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #112] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x800000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #64] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x1000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #40] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x2000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    ldr x30, [sp, #40] // 8-byte Reload
-; CHECK-NEON-NEXT:    str x8, [sp, #32] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x4000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #56] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x8000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #48] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x10000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #88] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x20000000
-; CHECK-NEON-NEXT:    mul x26, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x40000000
-; CHECK-NEON-NEXT:    mul x22, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x80000000
-; CHECK-NEON-NEXT:    mul x23, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x100000000
-; CHECK-NEON-NEXT:    mul x24, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x200000000
-; CHECK-NEON-NEXT:    eor x22, x26, x22
-; CHECK-NEON-NEXT:    ldr x26, [sp, #32] // 8-byte Reload
-; CHECK-NEON-NEXT:    mul x25, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x400000000
-; CHECK-NEON-NEXT:    eor x22, x22, x23
-; CHECK-NEON-NEXT:    and x23, x1, #0x400000000000000
-; CHECK-NEON-NEXT:    mul x27, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x800000000
-; CHECK-NEON-NEXT:    eor x22, x22, x24
-; CHECK-NEON-NEXT:    ldr x24, [sp, #48] // 8-byte Reload
-; CHECK-NEON-NEXT:    mul x28, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x1000000000
-; CHECK-NEON-NEXT:    eor x22, x22, x25
-; CHECK-NEON-NEXT:    ldr x25, [sp, #88] // 8-byte Reload
-; CHECK-NEON-NEXT:    mul x29, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x2000000000
-; CHECK-NEON-NEXT:    eor x22, x22, x27
-; CHECK-NEON-NEXT:    mul x21, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x4000000000
-; CHECK-NEON-NEXT:    mul x7, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x8000000000
-; CHECK-NEON-NEXT:    mul x19, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x10000000000
-; CHECK-NEON-NEXT:    mul x5, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x20000000000
-; CHECK-NEON-NEXT:    eor x7, x21, x7
-; CHECK-NEON-NEXT:    mul x6, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x40000000000
-; CHECK-NEON-NEXT:    mul x20, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x80000000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    mul x23, x0, x23
-; CHECK-NEON-NEXT:    str x8, [sp, #24] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x100000000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #16] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x200000000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #8] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x400000000000
-; CHECK-NEON-NEXT:    mul x4, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x800000000000
-; CHECK-NEON-NEXT:    mul x17, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x1000000000000
-; CHECK-NEON-NEXT:    mul x18, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x2000000000000
-; CHECK-NEON-NEXT:    mul x3, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x4000000000000
-; CHECK-NEON-NEXT:    eor x17, x4, x17
-; CHECK-NEON-NEXT:    mul x2, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x8000000000000
-; CHECK-NEON-NEXT:    eor x17, x17, x18
-; CHECK-NEON-NEXT:    and x18, x1, #0x4000000000000000
-; CHECK-NEON-NEXT:    mul x16, x0, x8
-; CHECK-NEON-NEXT:    eor x8, x9, x10
-; CHECK-NEON-NEXT:    ldr x9, [sp, #160] // 8-byte Reload
-; CHECK-NEON-NEXT:    eor x10, x12, x14
-; CHECK-NEON-NEXT:    ldr x12, [sp, #80] // 8-byte Reload
-; CHECK-NEON-NEXT:    eor x17, x17, x3
-; CHECK-NEON-NEXT:    eor x9, x15, x9
-; CHECK-NEON-NEXT:    mul x15, x0, x11
-; CHECK-NEON-NEXT:    ldr x11, [sp, #200] // 8-byte Reload
-; CHECK-NEON-NEXT:    eor x17, x17, x2
-; CHECK-NEON-NEXT:    eor x10, x10, x11
-; CHECK-NEON-NEXT:    ldr x11, [sp, #152] // 8-byte Reload
-; CHECK-NEON-NEXT:    mul x18, x0, x18
-; CHECK-NEON-NEXT:    eor x8, x8, x10
-; CHECK-NEON-NEXT:    ldr x10, [sp, #184] // 8-byte Reload
-; CHECK-NEON-NEXT:    eor x16, x17, x16
-; CHECK-NEON-NEXT:    eor x9, x9, x11
-; CHECK-NEON-NEXT:    and x11, x1, #0x20000000000000
-; CHECK-NEON-NEXT:    ldr x17, [sp, #24] // 8-byte Reload
-; CHECK-NEON-NEXT:    eor x9, x9, x10
-; CHECK-NEON-NEXT:    mul x14, x0, x11
-; CHECK-NEON-NEXT:    and x10, x1, #0x40000000000000
-; CHECK-NEON-NEXT:    eor x11, x8, x9
-; CHECK-NEON-NEXT:    ldr x8, [sp, #192] // 8-byte Reload
-; CHECK-NEON-NEXT:    ldr x9, [sp, #144] // 8-byte Reload
-; CHECK-NEON-NEXT:    mul x13, x0, x10
-; CHECK-NEON-NEXT:    ldr x10, [sp, #136] // 8-byte Reload
-; CHECK-NEON-NEXT:    eor x15, x16, x15
-; CHECK-NEON-NEXT:    eor x8, x8, x9
-; CHECK-NEON-NEXT:    ldr x9, [sp, #120] // 8-byte Reload
-; CHECK-NEON-NEXT:    ldr x16, [sp, #16] // 8-byte Reload
-; CHECK-NEON-NEXT:    eor x8, x8, x10
-; CHECK-NEON-NEXT:    ldr x10, [sp, #72] // 8-byte Reload
-; CHECK-NEON-NEXT:    eor x9, x9, x12
-; CHECK-NEON...
[truncated]

@llvmorg-github-actions

Copy link
Copy Markdown

@llvm/pr-subscribers-backend-aarch64

Author: Folkert de Vries (folkertdev)

Changes

Well I think I've nerdsniped myself, at least for the simple case that can be based on multiplication with holes (via bearssl source, https://www.bearssl.org/constanttime.html#ghash-for-gcm, and the polyval crate).

That still leaves the widening case though, for which this blog post has some good info https://timtaubert.de/blog/2017/06/verified-binary-multiplication-for-ghash/. Also for 8-bit and 16-bit inputs this approach can be generalized but with smaller holes and fewer multiplications. I'll leave that as future work too.

I've tested this locally with a fuzzer against the current LLVM implementation (via the rust standard library) and pclmulqdq. So, I think it's correct but I'm not sure how to show that in a convincing way.

I know this needs tests updated for more targets, but I'd like CI to tell me which ones.

CC #203694
CC: @RKSimon, @topperc, @eisenwave, @xarkenz, @artagnon.


Patch is 279.90 KiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/203727.diff

4 Files Affected:

  • (modified) llvm/include/llvm/CodeGen/BasicTTIImpl.h (+20-5)
  • (modified) llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp (+51)
  • (modified) llvm/test/CodeGen/AArch64/clmul.ll (+506-2292)
  • (modified) llvm/test/CodeGen/X86/clmul.ll (+969-2739)
diff --git a/llvm/include/llvm/CodeGen/BasicTTIImpl.h b/llvm/include/llvm/CodeGen/BasicTTIImpl.h
index e4b6bf51c7a4e..b11057695c2bd 100644
--- a/llvm/include/llvm/CodeGen/BasicTTIImpl.h
+++ b/llvm/include/llvm/CodeGen/BasicTTIImpl.h
@@ -3090,18 +3090,33 @@ class BasicTTIImplBase : public TargetTransformInfoImplCRTPBase<T> {
     case Intrinsic::clmul: {
       // This cost model should match the expansion in
       // TargetLowering::expandCLMUL.
-      InstructionCost PerBitCostMul =
-          thisT()->getArithmeticInstrCost(Instruction::And, RetTy, CostKind) +
-          thisT()->getArithmeticInstrCost(Instruction::Mul, RetTy, CostKind) +
+      unsigned BW = RetTy->getScalarSizeInBits();
+      InstructionCost AndCost =
+          thisT()->getArithmeticInstrCost(Instruction::And, RetTy, CostKind);
+      InstructionCost OrCost =
+          thisT()->getArithmeticInstrCost(Instruction::Or, RetTy, CostKind);
+      InstructionCost XorCost =
           thisT()->getArithmeticInstrCost(Instruction::Xor, RetTy, CostKind);
+      InstructionCost MulCost =
+          thisT()->getArithmeticInstrCost(Instruction::Mul, RetTy, CostKind);
+
+      // When the multiplication with holes approach is used, that emits 16 MULs, 8
+      // + 4 ANDs, 12 XORs and 3 ORs.
+      if (BW >= 32 && BW <= 64 &&
+          TLI->isOperationLegalOrCustom(ISD::MUL,
+                                        TLI->getValueType(DL, RetTy))) {
+        return 16 * MulCost + 12 * AndCost + 12 * XorCost + 3 * OrCost;
+      }
+
+      InstructionCost PerBitCostMul = AndCost + MulCost + XorCost;
       InstructionCost PerBitCostBittest =
-          thisT()->getArithmeticInstrCost(Instruction::And, RetTy, CostKind) +
+          AndCost +
           thisT()->getCmpSelInstrCost(BinaryOperator::Select, RetTy, RetTy,
                                       ICmpInst::BAD_ICMP_PREDICATE, CostKind) +
           thisT()->getCmpSelInstrCost(Instruction::ICmp, RetTy, RetTy,
                                       ICmpInst::ICMP_NE, CostKind);
       InstructionCost PerBitCost = std::min(PerBitCostMul, PerBitCostBittest);
-      return RetTy->getScalarSizeInBits() * PerBitCost;
+      return BW * PerBitCost;
     }
     default:
       break;
diff --git a/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp b/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
index 95e03a419fb43..175a41aa502dc 100644
--- a/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
+++ b/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
@@ -8899,6 +8899,57 @@ SDValue TargetLowering::expandCLMUL(SDNode *Node, SelectionDAG &DAG) const {
     // calculation in BasicTTIImpl::getTypeBasedIntrinsicInstrCost for
     // Intrinsic::clmul.
 
+    // Strategy 4: multiplication with holes. 
+    //
+    // Uses "holes" (sequences of zeroes) to avoid carry spilling. When carries
+    // do occur, they wind up in a "hole" and are subsequently masked out of the
+    // result.
+    //
+    // A hole of 3 bits is optimal for 32-bit and 64-bit inputs. 128-bit
+    // integers need a larger hole, and for smaller integers the fallback below
+    // is more efficient.
+    //
+    // Based on bmul64 in bearssl and bmul in the rust polyval crate.
+    if (BW >= 32 && BW <= 64 &&
+        isOperationLegalOrCustom(ISD::MUL, getTypeToTransformTo(Ctx, VT))) {
+
+      // Set every fourth bit of each nibble, equivalent to 0b100010001...0001.
+      APInt MaskVal(BW, 0);
+      for (unsigned i = 0; i < BW; i += 4)
+        MaskVal.setBit(i);
+
+      // Create versions of X and Y that keep only the first, second, third, or
+      // fourth bit of every nibble.
+      SDValue M[4], Xp[4], Yp[4];
+      for (unsigned i = 0; i < 4; ++i) {
+        M[i] = DAG.getConstant(MaskVal.shl(i), DL, VT);
+        Xp[i] = DAG.getNode(ISD::AND, DL, VT, X, M[i]);
+        Yp[i] = DAG.getNode(ISD::AND, DL, VT, Y, M[i]);
+      }
+
+      // Codegens these expressions (16 multiplications):
+      //
+      // z0 = (x0 * y0) ^ (x1 * y3) ^ (x2 * y2) ^ (x3 * y1);
+      // z1 = (x0 * y1) ^ (x1 * y0) ^ (x2 * y3) ^ (x3 * y2);
+      // z2 = (x0 * y2) ^ (x1 * y1) ^ (x2 * y0) ^ (x3 * y3);
+      // z3 = (x0 * y3) ^ (x1 * y2) ^ (x2 * y1) ^ (x3 * y0);
+      SDValue Res;
+      for (unsigned i = 0; i < 4; ++i) {
+        SDValue Zi;
+        for (unsigned A = 0; A < 4; ++A) {
+          unsigned B = (i + 4 - A) % 4;
+          SDValue P = DAG.getNode(ISD::MUL, DL, VT, Xp[A], Yp[B]);
+          Zi = Zi ? DAG.getNode(ISD::XOR, DL, VT, Zi, P) : P;
+        }
+
+        // Keep only the bits belonging to this iteration, and bitwise or it all
+        // together.
+        Zi = DAG.getNode(ISD::AND, DL, VT, Zi, M[i]);
+        Res = Res ? DAG.getNode(ISD::OR, DL, VT, Res, Zi) : Zi;
+      }
+      return Res;
+    }
+
     EVT SetCCVT = getSetCCResultType(DAG.getDataLayout(), Ctx, VT);
 
     SDValue Res = DAG.getConstant(0, DL, VT);
diff --git a/llvm/test/CodeGen/AArch64/clmul.ll b/llvm/test/CodeGen/AArch64/clmul.ll
index bb42e411c4f99..bfa4b73b4e677 100644
--- a/llvm/test/CodeGen/AArch64/clmul.ll
+++ b/llvm/test/CodeGen/AArch64/clmul.ll
@@ -80,101 +80,49 @@ define i16 @clmul_i16(i16 %x, i16 %y) {
 define i32 @clmul_i32(i32 %x, i32 %y) {
 ; CHECK-NEON-LABEL: clmul_i32:
 ; CHECK-NEON:       // %bb.0:
-; CHECK-NEON-NEXT:    and w8, w1, #0x2
-; CHECK-NEON-NEXT:    and w9, w1, #0x1
-; CHECK-NEON-NEXT:    and w10, w1, #0x4
-; CHECK-NEON-NEXT:    mul w8, w0, w8
-; CHECK-NEON-NEXT:    and w11, w1, #0x8
-; CHECK-NEON-NEXT:    and w12, w1, #0x10
-; CHECK-NEON-NEXT:    mul w9, w0, w9
-; CHECK-NEON-NEXT:    and w13, w1, #0x20
-; CHECK-NEON-NEXT:    and w14, w1, #0x40
-; CHECK-NEON-NEXT:    mul w10, w0, w10
-; CHECK-NEON-NEXT:    and w2, w1, #0x800
-; CHECK-NEON-NEXT:    and w15, w1, #0x80
-; CHECK-NEON-NEXT:    mul w11, w0, w11
-; CHECK-NEON-NEXT:    and w16, w1, #0x100
-; CHECK-NEON-NEXT:    and w17, w1, #0x200
-; CHECK-NEON-NEXT:    mul w12, w0, w12
-; CHECK-NEON-NEXT:    eor w8, w9, w8
-; CHECK-NEON-NEXT:    and w9, w1, #0x1000
-; CHECK-NEON-NEXT:    mul w13, w0, w13
-; CHECK-NEON-NEXT:    and w18, w1, #0x400
-; CHECK-NEON-NEXT:    mul w14, w0, w14
-; CHECK-NEON-NEXT:    eor w10, w10, w11
-; CHECK-NEON-NEXT:    and w11, w1, #0x2000
-; CHECK-NEON-NEXT:    mul w2, w0, w2
-; CHECK-NEON-NEXT:    eor w8, w8, w10
-; CHECK-NEON-NEXT:    and w10, w1, #0x4000
-; CHECK-NEON-NEXT:    mul w9, w0, w9
-; CHECK-NEON-NEXT:    eor w12, w12, w13
-; CHECK-NEON-NEXT:    and w13, w1, #0x8000
-; CHECK-NEON-NEXT:    mul w15, w0, w15
-; CHECK-NEON-NEXT:    eor w12, w12, w14
-; CHECK-NEON-NEXT:    and w14, w1, #0x10000
-; CHECK-NEON-NEXT:    mul w16, w0, w16
-; CHECK-NEON-NEXT:    eor w8, w8, w12
-; CHECK-NEON-NEXT:    and w12, w1, #0x20000
-; CHECK-NEON-NEXT:    mul w11, w0, w11
-; CHECK-NEON-NEXT:    eor w9, w2, w9
-; CHECK-NEON-NEXT:    and w2, w1, #0x400000
-; CHECK-NEON-NEXT:    mul w17, w0, w17
-; CHECK-NEON-NEXT:    mul w10, w0, w10
-; CHECK-NEON-NEXT:    eor w15, w15, w16
-; CHECK-NEON-NEXT:    and w16, w1, #0x40000
-; CHECK-NEON-NEXT:    mul w13, w0, w13
-; CHECK-NEON-NEXT:    eor w9, w9, w11
-; CHECK-NEON-NEXT:    and w11, w1, #0x800000
-; CHECK-NEON-NEXT:    mul w18, w0, w18
-; CHECK-NEON-NEXT:    eor w15, w15, w17
-; CHECK-NEON-NEXT:    and w17, w1, #0x80000
-; CHECK-NEON-NEXT:    mul w14, w0, w14
-; CHECK-NEON-NEXT:    eor w9, w9, w10
-; CHECK-NEON-NEXT:    and w10, w1, #0x1000000
-; CHECK-NEON-NEXT:    mul w12, w0, w12
-; CHECK-NEON-NEXT:    eor w9, w9, w13
-; CHECK-NEON-NEXT:    and w13, w1, #0x2000000
-; CHECK-NEON-NEXT:    mul w16, w0, w16
-; CHECK-NEON-NEXT:    eor w15, w15, w18
-; CHECK-NEON-NEXT:    and w18, w1, #0x100000
-; CHECK-NEON-NEXT:    mul w2, w0, w2
-; CHECK-NEON-NEXT:    eor w8, w8, w15
-; CHECK-NEON-NEXT:    and w15, w1, #0x200000
-; CHECK-NEON-NEXT:    mul w11, w0, w11
+; CHECK-NEON-NEXT:    and w8, w1, #0x11111111
+; CHECK-NEON-NEXT:    and w9, w0, #0x22222222
+; CHECK-NEON-NEXT:    and w10, w1, #0x22222222
+; CHECK-NEON-NEXT:    and w11, w0, #0x11111111
+; CHECK-NEON-NEXT:    and w13, w1, #0x88888888
+; CHECK-NEON-NEXT:    and w15, w0, #0x44444444
+; CHECK-NEON-NEXT:    and w17, w1, #0x44444444
+; CHECK-NEON-NEXT:    and w18, w0, #0x88888888
+; CHECK-NEON-NEXT:    mul w12, w9, w8
+; CHECK-NEON-NEXT:    mul w14, w11, w10
+; CHECK-NEON-NEXT:    mul w16, w15, w13
+; CHECK-NEON-NEXT:    mul w0, w18, w17
+; CHECK-NEON-NEXT:    mul w1, w9, w13
 ; CHECK-NEON-NEXT:    eor w12, w14, w12
-; CHECK-NEON-NEXT:    and w14, w1, #0x4000000
-; CHECK-NEON-NEXT:    mul w17, w0, w17
-; CHECK-NEON-NEXT:    eor w12, w12, w16
-; CHECK-NEON-NEXT:    and w16, w1, #0x8000000
-; CHECK-NEON-NEXT:    mul w10, w0, w10
-; CHECK-NEON-NEXT:    eor w8, w8, w9
-; CHECK-NEON-NEXT:    mul w13, w0, w13
-; CHECK-NEON-NEXT:    eor w11, w2, w11
-; CHECK-NEON-NEXT:    and w2, w1, #0x20000000
-; CHECK-NEON-NEXT:    mul w18, w0, w18
-; CHECK-NEON-NEXT:    eor w12, w12, w17
-; CHECK-NEON-NEXT:    and w17, w1, #0x10000000
-; CHECK-NEON-NEXT:    mul w14, w0, w14
-; CHECK-NEON-NEXT:    eor w10, w11, w10
-; CHECK-NEON-NEXT:    and w11, w1, #0x40000000
-; CHECK-NEON-NEXT:    mul w15, w0, w15
-; CHECK-NEON-NEXT:    eor w10, w10, w13
-; CHECK-NEON-NEXT:    and w13, w1, #0x80000000
-; CHECK-NEON-NEXT:    mul w16, w0, w16
-; CHECK-NEON-NEXT:    eor w12, w12, w18
-; CHECK-NEON-NEXT:    mul w17, w0, w17
-; CHECK-NEON-NEXT:    eor w10, w10, w14
-; CHECK-NEON-NEXT:    mul w2, w0, w2
-; CHECK-NEON-NEXT:    eor w9, w12, w15
-; CHECK-NEON-NEXT:    mul w11, w0, w11
-; CHECK-NEON-NEXT:    eor w10, w10, w16
-; CHECK-NEON-NEXT:    eor w8, w8, w9
-; CHECK-NEON-NEXT:    mul w13, w0, w13
-; CHECK-NEON-NEXT:    eor w9, w10, w17
-; CHECK-NEON-NEXT:    eor w8, w8, w9
-; CHECK-NEON-NEXT:    eor w10, w2, w11
-; CHECK-NEON-NEXT:    eor w9, w10, w13
-; CHECK-NEON-NEXT:    eor w0, w8, w9
+; CHECK-NEON-NEXT:    mul w2, w11, w8
+; CHECK-NEON-NEXT:    mul w3, w15, w17
+; CHECK-NEON-NEXT:    eor w14, w16, w0
+; CHECK-NEON-NEXT:    mul w4, w18, w10
+; CHECK-NEON-NEXT:    eor w12, w12, w14
+; CHECK-NEON-NEXT:    mul w5, w9, w10
+; CHECK-NEON-NEXT:    eor w14, w2, w1
+; CHECK-NEON-NEXT:    and w12, w12, #0x22222222
+; CHECK-NEON-NEXT:    mul w9, w9, w17
+; CHECK-NEON-NEXT:    mul w17, w11, w17
+; CHECK-NEON-NEXT:    eor w16, w3, w4
+; CHECK-NEON-NEXT:    mul w10, w15, w10
+; CHECK-NEON-NEXT:    mul w15, w15, w8
+; CHECK-NEON-NEXT:    mul w11, w11, w13
+; CHECK-NEON-NEXT:    eor w17, w17, w5
+; CHECK-NEON-NEXT:    mul w13, w18, w13
+; CHECK-NEON-NEXT:    mul w8, w18, w8
+; CHECK-NEON-NEXT:    eor w9, w11, w9
+; CHECK-NEON-NEXT:    eor w13, w15, w13
+; CHECK-NEON-NEXT:    eor w8, w10, w8
+; CHECK-NEON-NEXT:    eor w10, w14, w16
+; CHECK-NEON-NEXT:    eor w11, w17, w13
+; CHECK-NEON-NEXT:    eor w8, w9, w8
+; CHECK-NEON-NEXT:    and w9, w10, #0x11111111
+; CHECK-NEON-NEXT:    and w10, w11, #0x44444444
+; CHECK-NEON-NEXT:    and w8, w8, #0x88888888
+; CHECK-NEON-NEXT:    orr w9, w9, w12
+; CHECK-NEON-NEXT:    orr w8, w10, w8
+; CHECK-NEON-NEXT:    orr w0, w9, w8
 ; CHECK-NEON-NEXT:    ret
 ;
 ; CHECK-AES-LABEL: clmul_i32:
@@ -191,274 +139,49 @@ define i32 @clmul_i32(i32 %x, i32 %y) {
 define i64 @clmul_i64(i64 %x, i64 %y) {
 ; CHECK-NEON-LABEL: clmul_i64:
 ; CHECK-NEON:       // %bb.0:
-; CHECK-NEON-NEXT:    sub sp, sp, #304
-; CHECK-NEON-NEXT:    stp x29, x30, [sp, #208] // 16-byte Folded Spill
-; CHECK-NEON-NEXT:    stp x28, x27, [sp, #224] // 16-byte Folded Spill
-; CHECK-NEON-NEXT:    stp x26, x25, [sp, #240] // 16-byte Folded Spill
-; CHECK-NEON-NEXT:    stp x24, x23, [sp, #256] // 16-byte Folded Spill
-; CHECK-NEON-NEXT:    stp x22, x21, [sp, #272] // 16-byte Folded Spill
-; CHECK-NEON-NEXT:    stp x20, x19, [sp, #288] // 16-byte Folded Spill
-; CHECK-NEON-NEXT:    .cfi_def_cfa_offset 304
-; CHECK-NEON-NEXT:    .cfi_offset w19, -8
-; CHECK-NEON-NEXT:    .cfi_offset w20, -16
-; CHECK-NEON-NEXT:    .cfi_offset w21, -24
-; CHECK-NEON-NEXT:    .cfi_offset w22, -32
-; CHECK-NEON-NEXT:    .cfi_offset w23, -40
-; CHECK-NEON-NEXT:    .cfi_offset w24, -48
-; CHECK-NEON-NEXT:    .cfi_offset w25, -56
-; CHECK-NEON-NEXT:    .cfi_offset w26, -64
-; CHECK-NEON-NEXT:    .cfi_offset w27, -72
-; CHECK-NEON-NEXT:    .cfi_offset w28, -80
-; CHECK-NEON-NEXT:    .cfi_offset w30, -88
-; CHECK-NEON-NEXT:    .cfi_offset w29, -96
-; CHECK-NEON-NEXT:    and x8, x1, #0x2
-; CHECK-NEON-NEXT:    mul x9, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x1
-; CHECK-NEON-NEXT:    mul x10, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x4
-; CHECK-NEON-NEXT:    mul x11, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x8
-; CHECK-NEON-NEXT:    mul x13, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x10
-; CHECK-NEON-NEXT:    eor x9, x10, x9
-; CHECK-NEON-NEXT:    mul x12, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x20
-; CHECK-NEON-NEXT:    mul x14, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x40
-; CHECK-NEON-NEXT:    eor x10, x11, x13
-; CHECK-NEON-NEXT:    and x11, x1, #0x10000000000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #200] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x80
-; CHECK-NEON-NEXT:    mul x15, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x100
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #160] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x200
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #152] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x400
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #184] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x800
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #192] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x1000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #144] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x2000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #136] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x4000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #176] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x8000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #168] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x10000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #120] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x20000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #80] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x40000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #72] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x80000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #104] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x100000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #96] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x200000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #128] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x400000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #112] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x800000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #64] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x1000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #40] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x2000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    ldr x30, [sp, #40] // 8-byte Reload
-; CHECK-NEON-NEXT:    str x8, [sp, #32] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x4000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #56] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x8000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #48] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x10000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #88] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x20000000
-; CHECK-NEON-NEXT:    mul x26, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x40000000
-; CHECK-NEON-NEXT:    mul x22, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x80000000
-; CHECK-NEON-NEXT:    mul x23, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x100000000
-; CHECK-NEON-NEXT:    mul x24, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x200000000
-; CHECK-NEON-NEXT:    eor x22, x26, x22
-; CHECK-NEON-NEXT:    ldr x26, [sp, #32] // 8-byte Reload
-; CHECK-NEON-NEXT:    mul x25, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x400000000
-; CHECK-NEON-NEXT:    eor x22, x22, x23
-; CHECK-NEON-NEXT:    and x23, x1, #0x400000000000000
-; CHECK-NEON-NEXT:    mul x27, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x800000000
-; CHECK-NEON-NEXT:    eor x22, x22, x24
-; CHECK-NEON-NEXT:    ldr x24, [sp, #48] // 8-byte Reload
-; CHECK-NEON-NEXT:    mul x28, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x1000000000
-; CHECK-NEON-NEXT:    eor x22, x22, x25
-; CHECK-NEON-NEXT:    ldr x25, [sp, #88] // 8-byte Reload
-; CHECK-NEON-NEXT:    mul x29, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x2000000000
-; CHECK-NEON-NEXT:    eor x22, x22, x27
-; CHECK-NEON-NEXT:    mul x21, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x4000000000
-; CHECK-NEON-NEXT:    mul x7, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x8000000000
-; CHECK-NEON-NEXT:    mul x19, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x10000000000
-; CHECK-NEON-NEXT:    mul x5, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x20000000000
-; CHECK-NEON-NEXT:    eor x7, x21, x7
-; CHECK-NEON-NEXT:    mul x6, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x40000000000
-; CHECK-NEON-NEXT:    mul x20, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x80000000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    mul x23, x0, x23
-; CHECK-NEON-NEXT:    str x8, [sp, #24] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x100000000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #16] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x200000000000
-; CHECK-NEON-NEXT:    mul x8, x0, x8
-; CHECK-NEON-NEXT:    str x8, [sp, #8] // 8-byte Spill
-; CHECK-NEON-NEXT:    and x8, x1, #0x400000000000
-; CHECK-NEON-NEXT:    mul x4, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x800000000000
-; CHECK-NEON-NEXT:    mul x17, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x1000000000000
-; CHECK-NEON-NEXT:    mul x18, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x2000000000000
-; CHECK-NEON-NEXT:    mul x3, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x4000000000000
-; CHECK-NEON-NEXT:    eor x17, x4, x17
-; CHECK-NEON-NEXT:    mul x2, x0, x8
-; CHECK-NEON-NEXT:    and x8, x1, #0x8000000000000
-; CHECK-NEON-NEXT:    eor x17, x17, x18
-; CHECK-NEON-NEXT:    and x18, x1, #0x4000000000000000
-; CHECK-NEON-NEXT:    mul x16, x0, x8
-; CHECK-NEON-NEXT:    eor x8, x9, x10
-; CHECK-NEON-NEXT:    ldr x9, [sp, #160] // 8-byte Reload
-; CHECK-NEON-NEXT:    eor x10, x12, x14
-; CHECK-NEON-NEXT:    ldr x12, [sp, #80] // 8-byte Reload
-; CHECK-NEON-NEXT:    eor x17, x17, x3
-; CHECK-NEON-NEXT:    eor x9, x15, x9
-; CHECK-NEON-NEXT:    mul x15, x0, x11
-; CHECK-NEON-NEXT:    ldr x11, [sp, #200] // 8-byte Reload
-; CHECK-NEON-NEXT:    eor x17, x17, x2
-; CHECK-NEON-NEXT:    eor x10, x10, x11
-; CHECK-NEON-NEXT:    ldr x11, [sp, #152] // 8-byte Reload
-; CHECK-NEON-NEXT:    mul x18, x0, x18
-; CHECK-NEON-NEXT:    eor x8, x8, x10
-; CHECK-NEON-NEXT:    ldr x10, [sp, #184] // 8-byte Reload
-; CHECK-NEON-NEXT:    eor x16, x17, x16
-; CHECK-NEON-NEXT:    eor x9, x9, x11
-; CHECK-NEON-NEXT:    and x11, x1, #0x20000000000000
-; CHECK-NEON-NEXT:    ldr x17, [sp, #24] // 8-byte Reload
-; CHECK-NEON-NEXT:    eor x9, x9, x10
-; CHECK-NEON-NEXT:    mul x14, x0, x11
-; CHECK-NEON-NEXT:    and x10, x1, #0x40000000000000
-; CHECK-NEON-NEXT:    eor x11, x8, x9
-; CHECK-NEON-NEXT:    ldr x8, [sp, #192] // 8-byte Reload
-; CHECK-NEON-NEXT:    ldr x9, [sp, #144] // 8-byte Reload
-; CHECK-NEON-NEXT:    mul x13, x0, x10
-; CHECK-NEON-NEXT:    ldr x10, [sp, #136] // 8-byte Reload
-; CHECK-NEON-NEXT:    eor x15, x16, x15
-; CHECK-NEON-NEXT:    eor x8, x8, x9
-; CHECK-NEON-NEXT:    ldr x9, [sp, #120] // 8-byte Reload
-; CHECK-NEON-NEXT:    ldr x16, [sp, #16] // 8-byte Reload
-; CHECK-NEON-NEXT:    eor x8, x8, x10
-; CHECK-NEON-NEXT:    ldr x10, [sp, #72] // 8-byte Reload
-; CHECK-NEON-NEXT:    eor x9, x9, x12
-; CHECK-NEON...
[truncated]

@github-actions

github-actions Bot commented Jun 13, 2026

Copy link
Copy Markdown

✅ With the latest revision this PR passed the C/C++ code formatter.

@folkertdev
folkertdev force-pushed the clmul-better-fallback branch from 039cf96 to d8dc723 Compare June 13, 2026 21:03
@github-actions

github-actions Bot commented Jun 13, 2026

Copy link
Copy Markdown

🪟 Windows x64 Test Results

  • 135997 tests passed
  • 3458 tests skipped

✅ The build succeeded and all tests passed.

@eisenwave

eisenwave commented Jun 13, 2026

Copy link
Copy Markdown
Contributor

Do you know how this performs compared to the suggested approach outlined in

You can find some quick-bench there and an implementation that works for a _BitInt of any size, so it shouldn't be too hard to compare.

I'm mainly concerned about having too many different implementations, especially if what you're adding only covers between 32 and 64 bits. That still doesn't get rid of the worst-case linear implementation.

Thank you for working on this!

Comment thread llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
Comment thread llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp Outdated
Comment thread llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp Outdated
Comment thread llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp Outdated
Comment thread llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
Comment thread llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp Outdated
Comment thread llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp Outdated

@folkertdev folkertdev left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the comments, I'll look at the technical stuff tomorrow.

Re #203694 I'm not sure, I'll have to look into it. On the one hand I'd assume projects like bearssl would use the most efficient implementation, on the other hand occasionally good results just don't see wide adoption.

Just based on the cost model though, maybe you have an intuition for whether that other implementation beats it?

When the multiplication with holes approach is used, that emits 16 MULs, 8 + 4 ANDs, 12 XORs and 3 ORs.

Comment thread llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
Comment thread llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp Outdated
Comment thread llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp Outdated
Comment thread llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp Outdated
@eisenwave

Copy link
Copy Markdown
Contributor

Just based on the cost model though, maybe you have an intuition for whether that other implementation beats it?

When the multiplication with holes approach is used, that emits 16 MULs, 8 + 4 ANDs, 12 XORs and 3 ORs.

I checked what a narrowing form of my approach looks like for 64-bit, and it generates 21 ANDs, 41 SHLs/SHRs, 24 XORs, and a large amount of memory access gunk because the 8-element lookup table is creating too much register pressure (?). It does not generate any multiplication.

Just when eyeballing it, https://godbolt.org/z/sM914e4fo from rust-lang/rust#157831 looks much better than my https://godbolt.org/z/5eMKjndvj

@folkertdev

Copy link
Copy Markdown
Contributor Author

The comparison should be against the multiply with holes approach right? so really

It should be possible to extend this approach, I just didn't yet because

  • this algorithm is complicated as it is
  • I'm not actually that familiar with C++ (and the dialect used by LLVM)
  • I'm not that familiar with this part of LLVM

But the stride (which is currently hardcoded to 4) can be lowered to generate good code for i8 and i16, and for wider integers maybe a stride of 5 can work for i128, or maybe at that point a karatsuba approach is better (i.e. the 128-bit implementation is just defined in terms of 64-bit, and so on for 256-bit and up).

So I think the naive fallback can eventually be removed, it just seemed a bit much to take on in a single PR.

@folkertdev folkertdev left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lol I updated the riscv tests locally and

5 files changed, 55626 insertions(+), 89048 deletions(-)

So I'll hold off on adding all of that until we're happy with the implementation.

Comment thread llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp Outdated
Comment thread llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
@folkertdev
folkertdev requested a review from artagnon June 14, 2026 10:32
Comment thread llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp Outdated
Comment thread llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
Comment thread llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp Outdated
@github-actions

github-actions Bot commented Jun 17, 2026

Copy link
Copy Markdown

🐧 Linux x64 Test Results

  • 197077 tests passed
  • 5400 tests skipped

✅ The build succeeded and all tests passed.

Comment thread llvm/include/llvm/CodeGen/BasicTTIImpl.h Outdated
it does one shift and xor per bit, the holey approach is 43 operations
@llvmorg-github-actions llvmorg-github-actions Bot added the llvm:analysis Includes value tracking, cost tables and constant folding label Jun 17, 2026
@folkertdev
folkertdev requested review from artagnon and topperc June 18, 2026 12:20

@topperc topperc left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@artagnon artagnon left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Kindly clean up the title/description of the PR before landing, as it will be used as the commit message.

@folkertdev folkertdev changed the title use multiplication with holes for 32-bit/64-bit clmul fallback improve clmul fallback implementation for i32 and i64 Jun 18, 2026
@folkertdev
folkertdev enabled auto-merge (squash) June 18, 2026 15:49
@artagnon artagnon changed the title improve clmul fallback implementation for i32 and i64 [ISel] improve clmul fallback implementation for i32 and i64 Jun 18, 2026
@artagnon artagnon changed the title [ISel] improve clmul fallback implementation for i32 and i64 [ISel] Improve clmul fallback implementation for i32 and i64 Jun 18, 2026
@folkertdev
folkertdev disabled auto-merge June 18, 2026 17:18
@folkertdev
folkertdev enabled auto-merge (squash) June 18, 2026 17:18
@folkertdev
folkertdev merged commit 77f7abc into llvm:main Jun 18, 2026
14 of 15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backend:AArch64 backend:AMDGPU backend:PowerPC backend:RISC-V backend:X86 llvm:analysis Includes value tracking, cost tables and constant folding llvm:SelectionDAG SelectionDAGISel as well

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants