[ISel] Improve clmul fallback implementation for i32 and i64 - #203727
Conversation
|
@llvm/pr-subscribers-llvm-analysis @llvm/pr-subscribers-backend-x86 Author: Folkert de Vries (folkertdev) ChangesWell I think I've nerdsniped myself, at least for the simple case that can be based on multiplication with holes (via bearssl source, https://www.bearssl.org/constanttime.html#ghash-for-gcm, and the polyval crate). That still leaves the widening case though, for which this blog post has some good info https://timtaubert.de/blog/2017/06/verified-binary-multiplication-for-ghash/. Also for 8-bit and 16-bit inputs this approach can be generalized but with smaller holes and fewer multiplications. I'll leave that as future work too. I've tested this locally with a fuzzer against the current LLVM implementation (via the rust standard library) and I know this needs tests updated for more targets, but I'd like CI to tell me which ones. CC #203694 Patch is 279.90 KiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/203727.diff 4 Files Affected:
diff --git a/llvm/include/llvm/CodeGen/BasicTTIImpl.h b/llvm/include/llvm/CodeGen/BasicTTIImpl.h
index e4b6bf51c7a4e..b11057695c2bd 100644
--- a/llvm/include/llvm/CodeGen/BasicTTIImpl.h
+++ b/llvm/include/llvm/CodeGen/BasicTTIImpl.h
@@ -3090,18 +3090,33 @@ class BasicTTIImplBase : public TargetTransformInfoImplCRTPBase<T> {
case Intrinsic::clmul: {
// This cost model should match the expansion in
// TargetLowering::expandCLMUL.
- InstructionCost PerBitCostMul =
- thisT()->getArithmeticInstrCost(Instruction::And, RetTy, CostKind) +
- thisT()->getArithmeticInstrCost(Instruction::Mul, RetTy, CostKind) +
+ unsigned BW = RetTy->getScalarSizeInBits();
+ InstructionCost AndCost =
+ thisT()->getArithmeticInstrCost(Instruction::And, RetTy, CostKind);
+ InstructionCost OrCost =
+ thisT()->getArithmeticInstrCost(Instruction::Or, RetTy, CostKind);
+ InstructionCost XorCost =
thisT()->getArithmeticInstrCost(Instruction::Xor, RetTy, CostKind);
+ InstructionCost MulCost =
+ thisT()->getArithmeticInstrCost(Instruction::Mul, RetTy, CostKind);
+
+ // When the multiplication with holes approach is used, that emits 16 MULs, 8
+ // + 4 ANDs, 12 XORs and 3 ORs.
+ if (BW >= 32 && BW <= 64 &&
+ TLI->isOperationLegalOrCustom(ISD::MUL,
+ TLI->getValueType(DL, RetTy))) {
+ return 16 * MulCost + 12 * AndCost + 12 * XorCost + 3 * OrCost;
+ }
+
+ InstructionCost PerBitCostMul = AndCost + MulCost + XorCost;
InstructionCost PerBitCostBittest =
- thisT()->getArithmeticInstrCost(Instruction::And, RetTy, CostKind) +
+ AndCost +
thisT()->getCmpSelInstrCost(BinaryOperator::Select, RetTy, RetTy,
ICmpInst::BAD_ICMP_PREDICATE, CostKind) +
thisT()->getCmpSelInstrCost(Instruction::ICmp, RetTy, RetTy,
ICmpInst::ICMP_NE, CostKind);
InstructionCost PerBitCost = std::min(PerBitCostMul, PerBitCostBittest);
- return RetTy->getScalarSizeInBits() * PerBitCost;
+ return BW * PerBitCost;
}
default:
break;
diff --git a/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp b/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
index 95e03a419fb43..175a41aa502dc 100644
--- a/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
+++ b/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
@@ -8899,6 +8899,57 @@ SDValue TargetLowering::expandCLMUL(SDNode *Node, SelectionDAG &DAG) const {
// calculation in BasicTTIImpl::getTypeBasedIntrinsicInstrCost for
// Intrinsic::clmul.
+ // Strategy 4: multiplication with holes.
+ //
+ // Uses "holes" (sequences of zeroes) to avoid carry spilling. When carries
+ // do occur, they wind up in a "hole" and are subsequently masked out of the
+ // result.
+ //
+ // A hole of 3 bits is optimal for 32-bit and 64-bit inputs. 128-bit
+ // integers need a larger hole, and for smaller integers the fallback below
+ // is more efficient.
+ //
+ // Based on bmul64 in bearssl and bmul in the rust polyval crate.
+ if (BW >= 32 && BW <= 64 &&
+ isOperationLegalOrCustom(ISD::MUL, getTypeToTransformTo(Ctx, VT))) {
+
+ // Set every fourth bit of each nibble, equivalent to 0b100010001...0001.
+ APInt MaskVal(BW, 0);
+ for (unsigned i = 0; i < BW; i += 4)
+ MaskVal.setBit(i);
+
+ // Create versions of X and Y that keep only the first, second, third, or
+ // fourth bit of every nibble.
+ SDValue M[4], Xp[4], Yp[4];
+ for (unsigned i = 0; i < 4; ++i) {
+ M[i] = DAG.getConstant(MaskVal.shl(i), DL, VT);
+ Xp[i] = DAG.getNode(ISD::AND, DL, VT, X, M[i]);
+ Yp[i] = DAG.getNode(ISD::AND, DL, VT, Y, M[i]);
+ }
+
+ // Codegens these expressions (16 multiplications):
+ //
+ // z0 = (x0 * y0) ^ (x1 * y3) ^ (x2 * y2) ^ (x3 * y1);
+ // z1 = (x0 * y1) ^ (x1 * y0) ^ (x2 * y3) ^ (x3 * y2);
+ // z2 = (x0 * y2) ^ (x1 * y1) ^ (x2 * y0) ^ (x3 * y3);
+ // z3 = (x0 * y3) ^ (x1 * y2) ^ (x2 * y1) ^ (x3 * y0);
+ SDValue Res;
+ for (unsigned i = 0; i < 4; ++i) {
+ SDValue Zi;
+ for (unsigned A = 0; A < 4; ++A) {
+ unsigned B = (i + 4 - A) % 4;
+ SDValue P = DAG.getNode(ISD::MUL, DL, VT, Xp[A], Yp[B]);
+ Zi = Zi ? DAG.getNode(ISD::XOR, DL, VT, Zi, P) : P;
+ }
+
+ // Keep only the bits belonging to this iteration, and bitwise or it all
+ // together.
+ Zi = DAG.getNode(ISD::AND, DL, VT, Zi, M[i]);
+ Res = Res ? DAG.getNode(ISD::OR, DL, VT, Res, Zi) : Zi;
+ }
+ return Res;
+ }
+
EVT SetCCVT = getSetCCResultType(DAG.getDataLayout(), Ctx, VT);
SDValue Res = DAG.getConstant(0, DL, VT);
diff --git a/llvm/test/CodeGen/AArch64/clmul.ll b/llvm/test/CodeGen/AArch64/clmul.ll
index bb42e411c4f99..bfa4b73b4e677 100644
--- a/llvm/test/CodeGen/AArch64/clmul.ll
+++ b/llvm/test/CodeGen/AArch64/clmul.ll
@@ -80,101 +80,49 @@ define i16 @clmul_i16(i16 %x, i16 %y) {
define i32 @clmul_i32(i32 %x, i32 %y) {
; CHECK-NEON-LABEL: clmul_i32:
; CHECK-NEON: // %bb.0:
-; CHECK-NEON-NEXT: and w8, w1, #0x2
-; CHECK-NEON-NEXT: and w9, w1, #0x1
-; CHECK-NEON-NEXT: and w10, w1, #0x4
-; CHECK-NEON-NEXT: mul w8, w0, w8
-; CHECK-NEON-NEXT: and w11, w1, #0x8
-; CHECK-NEON-NEXT: and w12, w1, #0x10
-; CHECK-NEON-NEXT: mul w9, w0, w9
-; CHECK-NEON-NEXT: and w13, w1, #0x20
-; CHECK-NEON-NEXT: and w14, w1, #0x40
-; CHECK-NEON-NEXT: mul w10, w0, w10
-; CHECK-NEON-NEXT: and w2, w1, #0x800
-; CHECK-NEON-NEXT: and w15, w1, #0x80
-; CHECK-NEON-NEXT: mul w11, w0, w11
-; CHECK-NEON-NEXT: and w16, w1, #0x100
-; CHECK-NEON-NEXT: and w17, w1, #0x200
-; CHECK-NEON-NEXT: mul w12, w0, w12
-; CHECK-NEON-NEXT: eor w8, w9, w8
-; CHECK-NEON-NEXT: and w9, w1, #0x1000
-; CHECK-NEON-NEXT: mul w13, w0, w13
-; CHECK-NEON-NEXT: and w18, w1, #0x400
-; CHECK-NEON-NEXT: mul w14, w0, w14
-; CHECK-NEON-NEXT: eor w10, w10, w11
-; CHECK-NEON-NEXT: and w11, w1, #0x2000
-; CHECK-NEON-NEXT: mul w2, w0, w2
-; CHECK-NEON-NEXT: eor w8, w8, w10
-; CHECK-NEON-NEXT: and w10, w1, #0x4000
-; CHECK-NEON-NEXT: mul w9, w0, w9
-; CHECK-NEON-NEXT: eor w12, w12, w13
-; CHECK-NEON-NEXT: and w13, w1, #0x8000
-; CHECK-NEON-NEXT: mul w15, w0, w15
-; CHECK-NEON-NEXT: eor w12, w12, w14
-; CHECK-NEON-NEXT: and w14, w1, #0x10000
-; CHECK-NEON-NEXT: mul w16, w0, w16
-; CHECK-NEON-NEXT: eor w8, w8, w12
-; CHECK-NEON-NEXT: and w12, w1, #0x20000
-; CHECK-NEON-NEXT: mul w11, w0, w11
-; CHECK-NEON-NEXT: eor w9, w2, w9
-; CHECK-NEON-NEXT: and w2, w1, #0x400000
-; CHECK-NEON-NEXT: mul w17, w0, w17
-; CHECK-NEON-NEXT: mul w10, w0, w10
-; CHECK-NEON-NEXT: eor w15, w15, w16
-; CHECK-NEON-NEXT: and w16, w1, #0x40000
-; CHECK-NEON-NEXT: mul w13, w0, w13
-; CHECK-NEON-NEXT: eor w9, w9, w11
-; CHECK-NEON-NEXT: and w11, w1, #0x800000
-; CHECK-NEON-NEXT: mul w18, w0, w18
-; CHECK-NEON-NEXT: eor w15, w15, w17
-; CHECK-NEON-NEXT: and w17, w1, #0x80000
-; CHECK-NEON-NEXT: mul w14, w0, w14
-; CHECK-NEON-NEXT: eor w9, w9, w10
-; CHECK-NEON-NEXT: and w10, w1, #0x1000000
-; CHECK-NEON-NEXT: mul w12, w0, w12
-; CHECK-NEON-NEXT: eor w9, w9, w13
-; CHECK-NEON-NEXT: and w13, w1, #0x2000000
-; CHECK-NEON-NEXT: mul w16, w0, w16
-; CHECK-NEON-NEXT: eor w15, w15, w18
-; CHECK-NEON-NEXT: and w18, w1, #0x100000
-; CHECK-NEON-NEXT: mul w2, w0, w2
-; CHECK-NEON-NEXT: eor w8, w8, w15
-; CHECK-NEON-NEXT: and w15, w1, #0x200000
-; CHECK-NEON-NEXT: mul w11, w0, w11
+; CHECK-NEON-NEXT: and w8, w1, #0x11111111
+; CHECK-NEON-NEXT: and w9, w0, #0x22222222
+; CHECK-NEON-NEXT: and w10, w1, #0x22222222
+; CHECK-NEON-NEXT: and w11, w0, #0x11111111
+; CHECK-NEON-NEXT: and w13, w1, #0x88888888
+; CHECK-NEON-NEXT: and w15, w0, #0x44444444
+; CHECK-NEON-NEXT: and w17, w1, #0x44444444
+; CHECK-NEON-NEXT: and w18, w0, #0x88888888
+; CHECK-NEON-NEXT: mul w12, w9, w8
+; CHECK-NEON-NEXT: mul w14, w11, w10
+; CHECK-NEON-NEXT: mul w16, w15, w13
+; CHECK-NEON-NEXT: mul w0, w18, w17
+; CHECK-NEON-NEXT: mul w1, w9, w13
; CHECK-NEON-NEXT: eor w12, w14, w12
-; CHECK-NEON-NEXT: and w14, w1, #0x4000000
-; CHECK-NEON-NEXT: mul w17, w0, w17
-; CHECK-NEON-NEXT: eor w12, w12, w16
-; CHECK-NEON-NEXT: and w16, w1, #0x8000000
-; CHECK-NEON-NEXT: mul w10, w0, w10
-; CHECK-NEON-NEXT: eor w8, w8, w9
-; CHECK-NEON-NEXT: mul w13, w0, w13
-; CHECK-NEON-NEXT: eor w11, w2, w11
-; CHECK-NEON-NEXT: and w2, w1, #0x20000000
-; CHECK-NEON-NEXT: mul w18, w0, w18
-; CHECK-NEON-NEXT: eor w12, w12, w17
-; CHECK-NEON-NEXT: and w17, w1, #0x10000000
-; CHECK-NEON-NEXT: mul w14, w0, w14
-; CHECK-NEON-NEXT: eor w10, w11, w10
-; CHECK-NEON-NEXT: and w11, w1, #0x40000000
-; CHECK-NEON-NEXT: mul w15, w0, w15
-; CHECK-NEON-NEXT: eor w10, w10, w13
-; CHECK-NEON-NEXT: and w13, w1, #0x80000000
-; CHECK-NEON-NEXT: mul w16, w0, w16
-; CHECK-NEON-NEXT: eor w12, w12, w18
-; CHECK-NEON-NEXT: mul w17, w0, w17
-; CHECK-NEON-NEXT: eor w10, w10, w14
-; CHECK-NEON-NEXT: mul w2, w0, w2
-; CHECK-NEON-NEXT: eor w9, w12, w15
-; CHECK-NEON-NEXT: mul w11, w0, w11
-; CHECK-NEON-NEXT: eor w10, w10, w16
-; CHECK-NEON-NEXT: eor w8, w8, w9
-; CHECK-NEON-NEXT: mul w13, w0, w13
-; CHECK-NEON-NEXT: eor w9, w10, w17
-; CHECK-NEON-NEXT: eor w8, w8, w9
-; CHECK-NEON-NEXT: eor w10, w2, w11
-; CHECK-NEON-NEXT: eor w9, w10, w13
-; CHECK-NEON-NEXT: eor w0, w8, w9
+; CHECK-NEON-NEXT: mul w2, w11, w8
+; CHECK-NEON-NEXT: mul w3, w15, w17
+; CHECK-NEON-NEXT: eor w14, w16, w0
+; CHECK-NEON-NEXT: mul w4, w18, w10
+; CHECK-NEON-NEXT: eor w12, w12, w14
+; CHECK-NEON-NEXT: mul w5, w9, w10
+; CHECK-NEON-NEXT: eor w14, w2, w1
+; CHECK-NEON-NEXT: and w12, w12, #0x22222222
+; CHECK-NEON-NEXT: mul w9, w9, w17
+; CHECK-NEON-NEXT: mul w17, w11, w17
+; CHECK-NEON-NEXT: eor w16, w3, w4
+; CHECK-NEON-NEXT: mul w10, w15, w10
+; CHECK-NEON-NEXT: mul w15, w15, w8
+; CHECK-NEON-NEXT: mul w11, w11, w13
+; CHECK-NEON-NEXT: eor w17, w17, w5
+; CHECK-NEON-NEXT: mul w13, w18, w13
+; CHECK-NEON-NEXT: mul w8, w18, w8
+; CHECK-NEON-NEXT: eor w9, w11, w9
+; CHECK-NEON-NEXT: eor w13, w15, w13
+; CHECK-NEON-NEXT: eor w8, w10, w8
+; CHECK-NEON-NEXT: eor w10, w14, w16
+; CHECK-NEON-NEXT: eor w11, w17, w13
+; CHECK-NEON-NEXT: eor w8, w9, w8
+; CHECK-NEON-NEXT: and w9, w10, #0x11111111
+; CHECK-NEON-NEXT: and w10, w11, #0x44444444
+; CHECK-NEON-NEXT: and w8, w8, #0x88888888
+; CHECK-NEON-NEXT: orr w9, w9, w12
+; CHECK-NEON-NEXT: orr w8, w10, w8
+; CHECK-NEON-NEXT: orr w0, w9, w8
; CHECK-NEON-NEXT: ret
;
; CHECK-AES-LABEL: clmul_i32:
@@ -191,274 +139,49 @@ define i32 @clmul_i32(i32 %x, i32 %y) {
define i64 @clmul_i64(i64 %x, i64 %y) {
; CHECK-NEON-LABEL: clmul_i64:
; CHECK-NEON: // %bb.0:
-; CHECK-NEON-NEXT: sub sp, sp, #304
-; CHECK-NEON-NEXT: stp x29, x30, [sp, #208] // 16-byte Folded Spill
-; CHECK-NEON-NEXT: stp x28, x27, [sp, #224] // 16-byte Folded Spill
-; CHECK-NEON-NEXT: stp x26, x25, [sp, #240] // 16-byte Folded Spill
-; CHECK-NEON-NEXT: stp x24, x23, [sp, #256] // 16-byte Folded Spill
-; CHECK-NEON-NEXT: stp x22, x21, [sp, #272] // 16-byte Folded Spill
-; CHECK-NEON-NEXT: stp x20, x19, [sp, #288] // 16-byte Folded Spill
-; CHECK-NEON-NEXT: .cfi_def_cfa_offset 304
-; CHECK-NEON-NEXT: .cfi_offset w19, -8
-; CHECK-NEON-NEXT: .cfi_offset w20, -16
-; CHECK-NEON-NEXT: .cfi_offset w21, -24
-; CHECK-NEON-NEXT: .cfi_offset w22, -32
-; CHECK-NEON-NEXT: .cfi_offset w23, -40
-; CHECK-NEON-NEXT: .cfi_offset w24, -48
-; CHECK-NEON-NEXT: .cfi_offset w25, -56
-; CHECK-NEON-NEXT: .cfi_offset w26, -64
-; CHECK-NEON-NEXT: .cfi_offset w27, -72
-; CHECK-NEON-NEXT: .cfi_offset w28, -80
-; CHECK-NEON-NEXT: .cfi_offset w30, -88
-; CHECK-NEON-NEXT: .cfi_offset w29, -96
-; CHECK-NEON-NEXT: and x8, x1, #0x2
-; CHECK-NEON-NEXT: mul x9, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x1
-; CHECK-NEON-NEXT: mul x10, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x4
-; CHECK-NEON-NEXT: mul x11, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x8
-; CHECK-NEON-NEXT: mul x13, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x10
-; CHECK-NEON-NEXT: eor x9, x10, x9
-; CHECK-NEON-NEXT: mul x12, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x20
-; CHECK-NEON-NEXT: mul x14, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x40
-; CHECK-NEON-NEXT: eor x10, x11, x13
-; CHECK-NEON-NEXT: and x11, x1, #0x10000000000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #200] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x80
-; CHECK-NEON-NEXT: mul x15, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x100
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #160] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x200
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #152] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x400
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #184] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x800
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #192] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x1000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #144] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x2000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #136] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x4000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #176] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x8000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #168] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x10000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #120] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x20000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #80] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x40000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #72] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x80000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #104] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x100000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #96] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x200000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #128] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x400000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #112] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x800000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #64] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x1000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #40] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x2000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: ldr x30, [sp, #40] // 8-byte Reload
-; CHECK-NEON-NEXT: str x8, [sp, #32] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x4000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #56] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x8000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #48] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x10000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #88] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x20000000
-; CHECK-NEON-NEXT: mul x26, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x40000000
-; CHECK-NEON-NEXT: mul x22, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x80000000
-; CHECK-NEON-NEXT: mul x23, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x100000000
-; CHECK-NEON-NEXT: mul x24, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x200000000
-; CHECK-NEON-NEXT: eor x22, x26, x22
-; CHECK-NEON-NEXT: ldr x26, [sp, #32] // 8-byte Reload
-; CHECK-NEON-NEXT: mul x25, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x400000000
-; CHECK-NEON-NEXT: eor x22, x22, x23
-; CHECK-NEON-NEXT: and x23, x1, #0x400000000000000
-; CHECK-NEON-NEXT: mul x27, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x800000000
-; CHECK-NEON-NEXT: eor x22, x22, x24
-; CHECK-NEON-NEXT: ldr x24, [sp, #48] // 8-byte Reload
-; CHECK-NEON-NEXT: mul x28, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x1000000000
-; CHECK-NEON-NEXT: eor x22, x22, x25
-; CHECK-NEON-NEXT: ldr x25, [sp, #88] // 8-byte Reload
-; CHECK-NEON-NEXT: mul x29, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x2000000000
-; CHECK-NEON-NEXT: eor x22, x22, x27
-; CHECK-NEON-NEXT: mul x21, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x4000000000
-; CHECK-NEON-NEXT: mul x7, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x8000000000
-; CHECK-NEON-NEXT: mul x19, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x10000000000
-; CHECK-NEON-NEXT: mul x5, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x20000000000
-; CHECK-NEON-NEXT: eor x7, x21, x7
-; CHECK-NEON-NEXT: mul x6, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x40000000000
-; CHECK-NEON-NEXT: mul x20, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x80000000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: mul x23, x0, x23
-; CHECK-NEON-NEXT: str x8, [sp, #24] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x100000000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #16] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x200000000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #8] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x400000000000
-; CHECK-NEON-NEXT: mul x4, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x800000000000
-; CHECK-NEON-NEXT: mul x17, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x1000000000000
-; CHECK-NEON-NEXT: mul x18, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x2000000000000
-; CHECK-NEON-NEXT: mul x3, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x4000000000000
-; CHECK-NEON-NEXT: eor x17, x4, x17
-; CHECK-NEON-NEXT: mul x2, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x8000000000000
-; CHECK-NEON-NEXT: eor x17, x17, x18
-; CHECK-NEON-NEXT: and x18, x1, #0x4000000000000000
-; CHECK-NEON-NEXT: mul x16, x0, x8
-; CHECK-NEON-NEXT: eor x8, x9, x10
-; CHECK-NEON-NEXT: ldr x9, [sp, #160] // 8-byte Reload
-; CHECK-NEON-NEXT: eor x10, x12, x14
-; CHECK-NEON-NEXT: ldr x12, [sp, #80] // 8-byte Reload
-; CHECK-NEON-NEXT: eor x17, x17, x3
-; CHECK-NEON-NEXT: eor x9, x15, x9
-; CHECK-NEON-NEXT: mul x15, x0, x11
-; CHECK-NEON-NEXT: ldr x11, [sp, #200] // 8-byte Reload
-; CHECK-NEON-NEXT: eor x17, x17, x2
-; CHECK-NEON-NEXT: eor x10, x10, x11
-; CHECK-NEON-NEXT: ldr x11, [sp, #152] // 8-byte Reload
-; CHECK-NEON-NEXT: mul x18, x0, x18
-; CHECK-NEON-NEXT: eor x8, x8, x10
-; CHECK-NEON-NEXT: ldr x10, [sp, #184] // 8-byte Reload
-; CHECK-NEON-NEXT: eor x16, x17, x16
-; CHECK-NEON-NEXT: eor x9, x9, x11
-; CHECK-NEON-NEXT: and x11, x1, #0x20000000000000
-; CHECK-NEON-NEXT: ldr x17, [sp, #24] // 8-byte Reload
-; CHECK-NEON-NEXT: eor x9, x9, x10
-; CHECK-NEON-NEXT: mul x14, x0, x11
-; CHECK-NEON-NEXT: and x10, x1, #0x40000000000000
-; CHECK-NEON-NEXT: eor x11, x8, x9
-; CHECK-NEON-NEXT: ldr x8, [sp, #192] // 8-byte Reload
-; CHECK-NEON-NEXT: ldr x9, [sp, #144] // 8-byte Reload
-; CHECK-NEON-NEXT: mul x13, x0, x10
-; CHECK-NEON-NEXT: ldr x10, [sp, #136] // 8-byte Reload
-; CHECK-NEON-NEXT: eor x15, x16, x15
-; CHECK-NEON-NEXT: eor x8, x8, x9
-; CHECK-NEON-NEXT: ldr x9, [sp, #120] // 8-byte Reload
-; CHECK-NEON-NEXT: ldr x16, [sp, #16] // 8-byte Reload
-; CHECK-NEON-NEXT: eor x8, x8, x10
-; CHECK-NEON-NEXT: ldr x10, [sp, #72] // 8-byte Reload
-; CHECK-NEON-NEXT: eor x9, x9, x12
-; CHECK-NEON...
[truncated]
|
|
@llvm/pr-subscribers-backend-aarch64 Author: Folkert de Vries (folkertdev) ChangesWell I think I've nerdsniped myself, at least for the simple case that can be based on multiplication with holes (via bearssl source, https://www.bearssl.org/constanttime.html#ghash-for-gcm, and the polyval crate). That still leaves the widening case though, for which this blog post has some good info https://timtaubert.de/blog/2017/06/verified-binary-multiplication-for-ghash/. Also for 8-bit and 16-bit inputs this approach can be generalized but with smaller holes and fewer multiplications. I'll leave that as future work too. I've tested this locally with a fuzzer against the current LLVM implementation (via the rust standard library) and I know this needs tests updated for more targets, but I'd like CI to tell me which ones. CC #203694 Patch is 279.90 KiB, truncated to 20.00 KiB below, full version: https://github.com/llvm/llvm-project/pull/203727.diff 4 Files Affected:
diff --git a/llvm/include/llvm/CodeGen/BasicTTIImpl.h b/llvm/include/llvm/CodeGen/BasicTTIImpl.h
index e4b6bf51c7a4e..b11057695c2bd 100644
--- a/llvm/include/llvm/CodeGen/BasicTTIImpl.h
+++ b/llvm/include/llvm/CodeGen/BasicTTIImpl.h
@@ -3090,18 +3090,33 @@ class BasicTTIImplBase : public TargetTransformInfoImplCRTPBase<T> {
case Intrinsic::clmul: {
// This cost model should match the expansion in
// TargetLowering::expandCLMUL.
- InstructionCost PerBitCostMul =
- thisT()->getArithmeticInstrCost(Instruction::And, RetTy, CostKind) +
- thisT()->getArithmeticInstrCost(Instruction::Mul, RetTy, CostKind) +
+ unsigned BW = RetTy->getScalarSizeInBits();
+ InstructionCost AndCost =
+ thisT()->getArithmeticInstrCost(Instruction::And, RetTy, CostKind);
+ InstructionCost OrCost =
+ thisT()->getArithmeticInstrCost(Instruction::Or, RetTy, CostKind);
+ InstructionCost XorCost =
thisT()->getArithmeticInstrCost(Instruction::Xor, RetTy, CostKind);
+ InstructionCost MulCost =
+ thisT()->getArithmeticInstrCost(Instruction::Mul, RetTy, CostKind);
+
+ // When the multiplication with holes approach is used, that emits 16 MULs, 8
+ // + 4 ANDs, 12 XORs and 3 ORs.
+ if (BW >= 32 && BW <= 64 &&
+ TLI->isOperationLegalOrCustom(ISD::MUL,
+ TLI->getValueType(DL, RetTy))) {
+ return 16 * MulCost + 12 * AndCost + 12 * XorCost + 3 * OrCost;
+ }
+
+ InstructionCost PerBitCostMul = AndCost + MulCost + XorCost;
InstructionCost PerBitCostBittest =
- thisT()->getArithmeticInstrCost(Instruction::And, RetTy, CostKind) +
+ AndCost +
thisT()->getCmpSelInstrCost(BinaryOperator::Select, RetTy, RetTy,
ICmpInst::BAD_ICMP_PREDICATE, CostKind) +
thisT()->getCmpSelInstrCost(Instruction::ICmp, RetTy, RetTy,
ICmpInst::ICMP_NE, CostKind);
InstructionCost PerBitCost = std::min(PerBitCostMul, PerBitCostBittest);
- return RetTy->getScalarSizeInBits() * PerBitCost;
+ return BW * PerBitCost;
}
default:
break;
diff --git a/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp b/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
index 95e03a419fb43..175a41aa502dc 100644
--- a/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
+++ b/llvm/lib/CodeGen/SelectionDAG/TargetLowering.cpp
@@ -8899,6 +8899,57 @@ SDValue TargetLowering::expandCLMUL(SDNode *Node, SelectionDAG &DAG) const {
// calculation in BasicTTIImpl::getTypeBasedIntrinsicInstrCost for
// Intrinsic::clmul.
+ // Strategy 4: multiplication with holes.
+ //
+ // Uses "holes" (sequences of zeroes) to avoid carry spilling. When carries
+ // do occur, they wind up in a "hole" and are subsequently masked out of the
+ // result.
+ //
+ // A hole of 3 bits is optimal for 32-bit and 64-bit inputs. 128-bit
+ // integers need a larger hole, and for smaller integers the fallback below
+ // is more efficient.
+ //
+ // Based on bmul64 in bearssl and bmul in the rust polyval crate.
+ if (BW >= 32 && BW <= 64 &&
+ isOperationLegalOrCustom(ISD::MUL, getTypeToTransformTo(Ctx, VT))) {
+
+ // Set every fourth bit of each nibble, equivalent to 0b100010001...0001.
+ APInt MaskVal(BW, 0);
+ for (unsigned i = 0; i < BW; i += 4)
+ MaskVal.setBit(i);
+
+ // Create versions of X and Y that keep only the first, second, third, or
+ // fourth bit of every nibble.
+ SDValue M[4], Xp[4], Yp[4];
+ for (unsigned i = 0; i < 4; ++i) {
+ M[i] = DAG.getConstant(MaskVal.shl(i), DL, VT);
+ Xp[i] = DAG.getNode(ISD::AND, DL, VT, X, M[i]);
+ Yp[i] = DAG.getNode(ISD::AND, DL, VT, Y, M[i]);
+ }
+
+ // Codegens these expressions (16 multiplications):
+ //
+ // z0 = (x0 * y0) ^ (x1 * y3) ^ (x2 * y2) ^ (x3 * y1);
+ // z1 = (x0 * y1) ^ (x1 * y0) ^ (x2 * y3) ^ (x3 * y2);
+ // z2 = (x0 * y2) ^ (x1 * y1) ^ (x2 * y0) ^ (x3 * y3);
+ // z3 = (x0 * y3) ^ (x1 * y2) ^ (x2 * y1) ^ (x3 * y0);
+ SDValue Res;
+ for (unsigned i = 0; i < 4; ++i) {
+ SDValue Zi;
+ for (unsigned A = 0; A < 4; ++A) {
+ unsigned B = (i + 4 - A) % 4;
+ SDValue P = DAG.getNode(ISD::MUL, DL, VT, Xp[A], Yp[B]);
+ Zi = Zi ? DAG.getNode(ISD::XOR, DL, VT, Zi, P) : P;
+ }
+
+ // Keep only the bits belonging to this iteration, and bitwise or it all
+ // together.
+ Zi = DAG.getNode(ISD::AND, DL, VT, Zi, M[i]);
+ Res = Res ? DAG.getNode(ISD::OR, DL, VT, Res, Zi) : Zi;
+ }
+ return Res;
+ }
+
EVT SetCCVT = getSetCCResultType(DAG.getDataLayout(), Ctx, VT);
SDValue Res = DAG.getConstant(0, DL, VT);
diff --git a/llvm/test/CodeGen/AArch64/clmul.ll b/llvm/test/CodeGen/AArch64/clmul.ll
index bb42e411c4f99..bfa4b73b4e677 100644
--- a/llvm/test/CodeGen/AArch64/clmul.ll
+++ b/llvm/test/CodeGen/AArch64/clmul.ll
@@ -80,101 +80,49 @@ define i16 @clmul_i16(i16 %x, i16 %y) {
define i32 @clmul_i32(i32 %x, i32 %y) {
; CHECK-NEON-LABEL: clmul_i32:
; CHECK-NEON: // %bb.0:
-; CHECK-NEON-NEXT: and w8, w1, #0x2
-; CHECK-NEON-NEXT: and w9, w1, #0x1
-; CHECK-NEON-NEXT: and w10, w1, #0x4
-; CHECK-NEON-NEXT: mul w8, w0, w8
-; CHECK-NEON-NEXT: and w11, w1, #0x8
-; CHECK-NEON-NEXT: and w12, w1, #0x10
-; CHECK-NEON-NEXT: mul w9, w0, w9
-; CHECK-NEON-NEXT: and w13, w1, #0x20
-; CHECK-NEON-NEXT: and w14, w1, #0x40
-; CHECK-NEON-NEXT: mul w10, w0, w10
-; CHECK-NEON-NEXT: and w2, w1, #0x800
-; CHECK-NEON-NEXT: and w15, w1, #0x80
-; CHECK-NEON-NEXT: mul w11, w0, w11
-; CHECK-NEON-NEXT: and w16, w1, #0x100
-; CHECK-NEON-NEXT: and w17, w1, #0x200
-; CHECK-NEON-NEXT: mul w12, w0, w12
-; CHECK-NEON-NEXT: eor w8, w9, w8
-; CHECK-NEON-NEXT: and w9, w1, #0x1000
-; CHECK-NEON-NEXT: mul w13, w0, w13
-; CHECK-NEON-NEXT: and w18, w1, #0x400
-; CHECK-NEON-NEXT: mul w14, w0, w14
-; CHECK-NEON-NEXT: eor w10, w10, w11
-; CHECK-NEON-NEXT: and w11, w1, #0x2000
-; CHECK-NEON-NEXT: mul w2, w0, w2
-; CHECK-NEON-NEXT: eor w8, w8, w10
-; CHECK-NEON-NEXT: and w10, w1, #0x4000
-; CHECK-NEON-NEXT: mul w9, w0, w9
-; CHECK-NEON-NEXT: eor w12, w12, w13
-; CHECK-NEON-NEXT: and w13, w1, #0x8000
-; CHECK-NEON-NEXT: mul w15, w0, w15
-; CHECK-NEON-NEXT: eor w12, w12, w14
-; CHECK-NEON-NEXT: and w14, w1, #0x10000
-; CHECK-NEON-NEXT: mul w16, w0, w16
-; CHECK-NEON-NEXT: eor w8, w8, w12
-; CHECK-NEON-NEXT: and w12, w1, #0x20000
-; CHECK-NEON-NEXT: mul w11, w0, w11
-; CHECK-NEON-NEXT: eor w9, w2, w9
-; CHECK-NEON-NEXT: and w2, w1, #0x400000
-; CHECK-NEON-NEXT: mul w17, w0, w17
-; CHECK-NEON-NEXT: mul w10, w0, w10
-; CHECK-NEON-NEXT: eor w15, w15, w16
-; CHECK-NEON-NEXT: and w16, w1, #0x40000
-; CHECK-NEON-NEXT: mul w13, w0, w13
-; CHECK-NEON-NEXT: eor w9, w9, w11
-; CHECK-NEON-NEXT: and w11, w1, #0x800000
-; CHECK-NEON-NEXT: mul w18, w0, w18
-; CHECK-NEON-NEXT: eor w15, w15, w17
-; CHECK-NEON-NEXT: and w17, w1, #0x80000
-; CHECK-NEON-NEXT: mul w14, w0, w14
-; CHECK-NEON-NEXT: eor w9, w9, w10
-; CHECK-NEON-NEXT: and w10, w1, #0x1000000
-; CHECK-NEON-NEXT: mul w12, w0, w12
-; CHECK-NEON-NEXT: eor w9, w9, w13
-; CHECK-NEON-NEXT: and w13, w1, #0x2000000
-; CHECK-NEON-NEXT: mul w16, w0, w16
-; CHECK-NEON-NEXT: eor w15, w15, w18
-; CHECK-NEON-NEXT: and w18, w1, #0x100000
-; CHECK-NEON-NEXT: mul w2, w0, w2
-; CHECK-NEON-NEXT: eor w8, w8, w15
-; CHECK-NEON-NEXT: and w15, w1, #0x200000
-; CHECK-NEON-NEXT: mul w11, w0, w11
+; CHECK-NEON-NEXT: and w8, w1, #0x11111111
+; CHECK-NEON-NEXT: and w9, w0, #0x22222222
+; CHECK-NEON-NEXT: and w10, w1, #0x22222222
+; CHECK-NEON-NEXT: and w11, w0, #0x11111111
+; CHECK-NEON-NEXT: and w13, w1, #0x88888888
+; CHECK-NEON-NEXT: and w15, w0, #0x44444444
+; CHECK-NEON-NEXT: and w17, w1, #0x44444444
+; CHECK-NEON-NEXT: and w18, w0, #0x88888888
+; CHECK-NEON-NEXT: mul w12, w9, w8
+; CHECK-NEON-NEXT: mul w14, w11, w10
+; CHECK-NEON-NEXT: mul w16, w15, w13
+; CHECK-NEON-NEXT: mul w0, w18, w17
+; CHECK-NEON-NEXT: mul w1, w9, w13
; CHECK-NEON-NEXT: eor w12, w14, w12
-; CHECK-NEON-NEXT: and w14, w1, #0x4000000
-; CHECK-NEON-NEXT: mul w17, w0, w17
-; CHECK-NEON-NEXT: eor w12, w12, w16
-; CHECK-NEON-NEXT: and w16, w1, #0x8000000
-; CHECK-NEON-NEXT: mul w10, w0, w10
-; CHECK-NEON-NEXT: eor w8, w8, w9
-; CHECK-NEON-NEXT: mul w13, w0, w13
-; CHECK-NEON-NEXT: eor w11, w2, w11
-; CHECK-NEON-NEXT: and w2, w1, #0x20000000
-; CHECK-NEON-NEXT: mul w18, w0, w18
-; CHECK-NEON-NEXT: eor w12, w12, w17
-; CHECK-NEON-NEXT: and w17, w1, #0x10000000
-; CHECK-NEON-NEXT: mul w14, w0, w14
-; CHECK-NEON-NEXT: eor w10, w11, w10
-; CHECK-NEON-NEXT: and w11, w1, #0x40000000
-; CHECK-NEON-NEXT: mul w15, w0, w15
-; CHECK-NEON-NEXT: eor w10, w10, w13
-; CHECK-NEON-NEXT: and w13, w1, #0x80000000
-; CHECK-NEON-NEXT: mul w16, w0, w16
-; CHECK-NEON-NEXT: eor w12, w12, w18
-; CHECK-NEON-NEXT: mul w17, w0, w17
-; CHECK-NEON-NEXT: eor w10, w10, w14
-; CHECK-NEON-NEXT: mul w2, w0, w2
-; CHECK-NEON-NEXT: eor w9, w12, w15
-; CHECK-NEON-NEXT: mul w11, w0, w11
-; CHECK-NEON-NEXT: eor w10, w10, w16
-; CHECK-NEON-NEXT: eor w8, w8, w9
-; CHECK-NEON-NEXT: mul w13, w0, w13
-; CHECK-NEON-NEXT: eor w9, w10, w17
-; CHECK-NEON-NEXT: eor w8, w8, w9
-; CHECK-NEON-NEXT: eor w10, w2, w11
-; CHECK-NEON-NEXT: eor w9, w10, w13
-; CHECK-NEON-NEXT: eor w0, w8, w9
+; CHECK-NEON-NEXT: mul w2, w11, w8
+; CHECK-NEON-NEXT: mul w3, w15, w17
+; CHECK-NEON-NEXT: eor w14, w16, w0
+; CHECK-NEON-NEXT: mul w4, w18, w10
+; CHECK-NEON-NEXT: eor w12, w12, w14
+; CHECK-NEON-NEXT: mul w5, w9, w10
+; CHECK-NEON-NEXT: eor w14, w2, w1
+; CHECK-NEON-NEXT: and w12, w12, #0x22222222
+; CHECK-NEON-NEXT: mul w9, w9, w17
+; CHECK-NEON-NEXT: mul w17, w11, w17
+; CHECK-NEON-NEXT: eor w16, w3, w4
+; CHECK-NEON-NEXT: mul w10, w15, w10
+; CHECK-NEON-NEXT: mul w15, w15, w8
+; CHECK-NEON-NEXT: mul w11, w11, w13
+; CHECK-NEON-NEXT: eor w17, w17, w5
+; CHECK-NEON-NEXT: mul w13, w18, w13
+; CHECK-NEON-NEXT: mul w8, w18, w8
+; CHECK-NEON-NEXT: eor w9, w11, w9
+; CHECK-NEON-NEXT: eor w13, w15, w13
+; CHECK-NEON-NEXT: eor w8, w10, w8
+; CHECK-NEON-NEXT: eor w10, w14, w16
+; CHECK-NEON-NEXT: eor w11, w17, w13
+; CHECK-NEON-NEXT: eor w8, w9, w8
+; CHECK-NEON-NEXT: and w9, w10, #0x11111111
+; CHECK-NEON-NEXT: and w10, w11, #0x44444444
+; CHECK-NEON-NEXT: and w8, w8, #0x88888888
+; CHECK-NEON-NEXT: orr w9, w9, w12
+; CHECK-NEON-NEXT: orr w8, w10, w8
+; CHECK-NEON-NEXT: orr w0, w9, w8
; CHECK-NEON-NEXT: ret
;
; CHECK-AES-LABEL: clmul_i32:
@@ -191,274 +139,49 @@ define i32 @clmul_i32(i32 %x, i32 %y) {
define i64 @clmul_i64(i64 %x, i64 %y) {
; CHECK-NEON-LABEL: clmul_i64:
; CHECK-NEON: // %bb.0:
-; CHECK-NEON-NEXT: sub sp, sp, #304
-; CHECK-NEON-NEXT: stp x29, x30, [sp, #208] // 16-byte Folded Spill
-; CHECK-NEON-NEXT: stp x28, x27, [sp, #224] // 16-byte Folded Spill
-; CHECK-NEON-NEXT: stp x26, x25, [sp, #240] // 16-byte Folded Spill
-; CHECK-NEON-NEXT: stp x24, x23, [sp, #256] // 16-byte Folded Spill
-; CHECK-NEON-NEXT: stp x22, x21, [sp, #272] // 16-byte Folded Spill
-; CHECK-NEON-NEXT: stp x20, x19, [sp, #288] // 16-byte Folded Spill
-; CHECK-NEON-NEXT: .cfi_def_cfa_offset 304
-; CHECK-NEON-NEXT: .cfi_offset w19, -8
-; CHECK-NEON-NEXT: .cfi_offset w20, -16
-; CHECK-NEON-NEXT: .cfi_offset w21, -24
-; CHECK-NEON-NEXT: .cfi_offset w22, -32
-; CHECK-NEON-NEXT: .cfi_offset w23, -40
-; CHECK-NEON-NEXT: .cfi_offset w24, -48
-; CHECK-NEON-NEXT: .cfi_offset w25, -56
-; CHECK-NEON-NEXT: .cfi_offset w26, -64
-; CHECK-NEON-NEXT: .cfi_offset w27, -72
-; CHECK-NEON-NEXT: .cfi_offset w28, -80
-; CHECK-NEON-NEXT: .cfi_offset w30, -88
-; CHECK-NEON-NEXT: .cfi_offset w29, -96
-; CHECK-NEON-NEXT: and x8, x1, #0x2
-; CHECK-NEON-NEXT: mul x9, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x1
-; CHECK-NEON-NEXT: mul x10, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x4
-; CHECK-NEON-NEXT: mul x11, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x8
-; CHECK-NEON-NEXT: mul x13, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x10
-; CHECK-NEON-NEXT: eor x9, x10, x9
-; CHECK-NEON-NEXT: mul x12, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x20
-; CHECK-NEON-NEXT: mul x14, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x40
-; CHECK-NEON-NEXT: eor x10, x11, x13
-; CHECK-NEON-NEXT: and x11, x1, #0x10000000000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #200] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x80
-; CHECK-NEON-NEXT: mul x15, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x100
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #160] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x200
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #152] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x400
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #184] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x800
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #192] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x1000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #144] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x2000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #136] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x4000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #176] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x8000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #168] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x10000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #120] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x20000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #80] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x40000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #72] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x80000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #104] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x100000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #96] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x200000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #128] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x400000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #112] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x800000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #64] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x1000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #40] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x2000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: ldr x30, [sp, #40] // 8-byte Reload
-; CHECK-NEON-NEXT: str x8, [sp, #32] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x4000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #56] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x8000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #48] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x10000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #88] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x20000000
-; CHECK-NEON-NEXT: mul x26, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x40000000
-; CHECK-NEON-NEXT: mul x22, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x80000000
-; CHECK-NEON-NEXT: mul x23, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x100000000
-; CHECK-NEON-NEXT: mul x24, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x200000000
-; CHECK-NEON-NEXT: eor x22, x26, x22
-; CHECK-NEON-NEXT: ldr x26, [sp, #32] // 8-byte Reload
-; CHECK-NEON-NEXT: mul x25, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x400000000
-; CHECK-NEON-NEXT: eor x22, x22, x23
-; CHECK-NEON-NEXT: and x23, x1, #0x400000000000000
-; CHECK-NEON-NEXT: mul x27, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x800000000
-; CHECK-NEON-NEXT: eor x22, x22, x24
-; CHECK-NEON-NEXT: ldr x24, [sp, #48] // 8-byte Reload
-; CHECK-NEON-NEXT: mul x28, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x1000000000
-; CHECK-NEON-NEXT: eor x22, x22, x25
-; CHECK-NEON-NEXT: ldr x25, [sp, #88] // 8-byte Reload
-; CHECK-NEON-NEXT: mul x29, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x2000000000
-; CHECK-NEON-NEXT: eor x22, x22, x27
-; CHECK-NEON-NEXT: mul x21, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x4000000000
-; CHECK-NEON-NEXT: mul x7, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x8000000000
-; CHECK-NEON-NEXT: mul x19, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x10000000000
-; CHECK-NEON-NEXT: mul x5, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x20000000000
-; CHECK-NEON-NEXT: eor x7, x21, x7
-; CHECK-NEON-NEXT: mul x6, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x40000000000
-; CHECK-NEON-NEXT: mul x20, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x80000000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: mul x23, x0, x23
-; CHECK-NEON-NEXT: str x8, [sp, #24] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x100000000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #16] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x200000000000
-; CHECK-NEON-NEXT: mul x8, x0, x8
-; CHECK-NEON-NEXT: str x8, [sp, #8] // 8-byte Spill
-; CHECK-NEON-NEXT: and x8, x1, #0x400000000000
-; CHECK-NEON-NEXT: mul x4, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x800000000000
-; CHECK-NEON-NEXT: mul x17, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x1000000000000
-; CHECK-NEON-NEXT: mul x18, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x2000000000000
-; CHECK-NEON-NEXT: mul x3, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x4000000000000
-; CHECK-NEON-NEXT: eor x17, x4, x17
-; CHECK-NEON-NEXT: mul x2, x0, x8
-; CHECK-NEON-NEXT: and x8, x1, #0x8000000000000
-; CHECK-NEON-NEXT: eor x17, x17, x18
-; CHECK-NEON-NEXT: and x18, x1, #0x4000000000000000
-; CHECK-NEON-NEXT: mul x16, x0, x8
-; CHECK-NEON-NEXT: eor x8, x9, x10
-; CHECK-NEON-NEXT: ldr x9, [sp, #160] // 8-byte Reload
-; CHECK-NEON-NEXT: eor x10, x12, x14
-; CHECK-NEON-NEXT: ldr x12, [sp, #80] // 8-byte Reload
-; CHECK-NEON-NEXT: eor x17, x17, x3
-; CHECK-NEON-NEXT: eor x9, x15, x9
-; CHECK-NEON-NEXT: mul x15, x0, x11
-; CHECK-NEON-NEXT: ldr x11, [sp, #200] // 8-byte Reload
-; CHECK-NEON-NEXT: eor x17, x17, x2
-; CHECK-NEON-NEXT: eor x10, x10, x11
-; CHECK-NEON-NEXT: ldr x11, [sp, #152] // 8-byte Reload
-; CHECK-NEON-NEXT: mul x18, x0, x18
-; CHECK-NEON-NEXT: eor x8, x8, x10
-; CHECK-NEON-NEXT: ldr x10, [sp, #184] // 8-byte Reload
-; CHECK-NEON-NEXT: eor x16, x17, x16
-; CHECK-NEON-NEXT: eor x9, x9, x11
-; CHECK-NEON-NEXT: and x11, x1, #0x20000000000000
-; CHECK-NEON-NEXT: ldr x17, [sp, #24] // 8-byte Reload
-; CHECK-NEON-NEXT: eor x9, x9, x10
-; CHECK-NEON-NEXT: mul x14, x0, x11
-; CHECK-NEON-NEXT: and x10, x1, #0x40000000000000
-; CHECK-NEON-NEXT: eor x11, x8, x9
-; CHECK-NEON-NEXT: ldr x8, [sp, #192] // 8-byte Reload
-; CHECK-NEON-NEXT: ldr x9, [sp, #144] // 8-byte Reload
-; CHECK-NEON-NEXT: mul x13, x0, x10
-; CHECK-NEON-NEXT: ldr x10, [sp, #136] // 8-byte Reload
-; CHECK-NEON-NEXT: eor x15, x16, x15
-; CHECK-NEON-NEXT: eor x8, x8, x9
-; CHECK-NEON-NEXT: ldr x9, [sp, #120] // 8-byte Reload
-; CHECK-NEON-NEXT: ldr x16, [sp, #16] // 8-byte Reload
-; CHECK-NEON-NEXT: eor x8, x8, x10
-; CHECK-NEON-NEXT: ldr x10, [sp, #72] // 8-byte Reload
-; CHECK-NEON-NEXT: eor x9, x9, x12
-; CHECK-NEON...
[truncated]
|
|
✅ With the latest revision this PR passed the C/C++ code formatter. |
039cf96 to
d8dc723
Compare
🪟 Windows x64 Test Results
✅ The build succeeded and all tests passed. |
|
Do you know how this performs compared to the suggested approach outlined in You can find some quick-bench there and an implementation that works for a I'm mainly concerned about having too many different implementations, especially if what you're adding only covers between 32 and 64 bits. That still doesn't get rid of the worst-case linear implementation. Thank you for working on this! |
folkertdev
left a comment
There was a problem hiding this comment.
Thanks for the comments, I'll look at the technical stuff tomorrow.
Re #203694 I'm not sure, I'll have to look into it. On the one hand I'd assume projects like bearssl would use the most efficient implementation, on the other hand occasionally good results just don't see wide adoption.
Just based on the cost model though, maybe you have an intuition for whether that other implementation beats it?
When the multiplication with holes approach is used, that emits 16 MULs, 8 + 4 ANDs, 12 XORs and 3 ORs.
I checked what a narrowing form of my approach looks like for 64-bit, and it generates 21 ANDs, 41 SHLs/SHRs, 24 XORs, and a large amount of memory access gunk because the 8-element lookup table is creating too much register pressure (?). It does not generate any multiplication. Just when eyeballing it, https://godbolt.org/z/sM914e4fo from rust-lang/rust#157831 looks much better than my https://godbolt.org/z/5eMKjndvj |
|
The comparison should be against the multiply with holes approach right? so really
It should be possible to extend this approach, I just didn't yet because
But the stride (which is currently hardcoded to 4) can be lowered to generate good code for So I think the naive fallback can eventually be removed, it just seemed a bit much to take on in a single PR. |
folkertdev
left a comment
There was a problem hiding this comment.
Lol I updated the riscv tests locally and
5 files changed, 55626 insertions(+), 89048 deletions(-)
So I'll hold off on adding all of that until we're happy with the implementation.
🐧 Linux x64 Test Results
✅ The build succeeded and all tests passed. |
it does one shift and xor per bit, the holey approach is 43 operations
artagnon
left a comment
There was a problem hiding this comment.
Kindly clean up the title/description of the PR before landing, as it will be used as the commit message.
clmul fallback implementation for i32 and i64
clmul fallback implementation for i32 and i64clmul fallback implementation for i32 and i64
clmul fallback implementation for i32 and i64clmul fallback implementation for i32 and i64
Improve the
clmulfallback implementation fori32..=i64.The new approach is "multiplication with holes", based on bearssl source, https://www.bearssl.org/constanttime.html#ghash-for-gcm, and the polyval crate. https://www.bearssl.org/constanttime.html#ghash-for-gcm explains the idea.
Future work is
I've tested this locally with a fuzzer against the current LLVM implementation (via the rust standard library) and
pclmulqdq.CC #203694
CC: @RKSimon, @topperc, @eisenwave, @xarkenz, @artagnon.