Skip to content

Add SimdBluesteins and SimdRaders - #199

Open
HEnquist wants to merge 2 commits into
ejmahler:masterfrom
HEnquist:simd_bluesteins_raders
Open

HEnquist wants to merge 2 commits into
ejmahler:masterfrom
HEnquist:simd_bluesteins_raders

Conversation

@HEnquist

@HEnquist HEnquist commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Implements #190: one Bluestein's and one Rader's shared by all four SIMD backends, written against SimdVector.

  • SimdBluesteins and SimdRaders in src/simd, which the SIMD planners now build instead of wrapping the scalar versions. The scalar BluesteinsAlgorithm and RadersAlgorithm become aliases for the single-complex-per-vector case, the way RadixN already is.
  • New mul_complex_conjugated on SimdVector: (a * b).conj() is a.conj() * b.conj(), so storing the precomputed side pre-conjugated turns every conjugated product into one operation. On FCMA it costs the same as a plain multiply, vcmla_rot270 instead of vcmla_rot90.
  • New FUSED_COMPLEX_MULTIPLY, because with a single complex per vector and a shuffle-based multiply these loops lose to the scalar loop LLVM autovectorizes into deinterleaving loads. NEON f64 was 3 to 5% slower without it. NEON, SSE and wasm f64 take the scalar path, FCMA f64 keeps the vector path.
  • boilerplate_simd_fft! beside boilerplate_fft!, for the Fft impl of the SimdVector-generic algorithms.

Apple M1, planned FFT at lengths 107 to 65579 against the old path: Bluestein's is 1.6 to 5.2% faster on NEON f32 and 1.8 to 13.0% on FCMA, Rader's 1 to 5% on FCMA. NEON f64 and the scalar planner are unchanged. Rader's gains less because its reorderings are permuted copies with nothing for a vector to do.

Every SIMD planner wrapped the scalar BluesteinsAlgorithm around a SIMD inner FFT. SimdBluesteins
is written against SimdVector instead, so there is one implementation for all four backends rather
than four, and the scalar BluesteinsAlgorithm becomes an alias for the single-complex-per-vector
case of it.

The three pairwise multiply loops are the only SIMD part. They go through new SimdVector methods so
that they land inside the backend's target feature, and use a new mul_complex_conjugated primitive:
(a * b).conj() is a.conj() * b.conj(), so storing the precomputed side pre-conjugated turns every
conjugated product into one operation. FCMA reaches it in the same instruction count as a plain
complex multiply, by swapping vcmla_rot90 for vcmla_rot270.

With one complex number per vector, interleaved data and a shuffle-based multiply, those loops lose
to the plain scalar loop LLVM autovectorizes into deinterleaving loads: NEON f64 came out 3 to 5%
slower at the larger sizes. The new FUSED_COMPLEX_MULTIPLY says whether a backend has native
complex multiply instructions, and the loops take the scalar path when it is false and a vector
holds a single complex number. That covers NEON, SSE and wasm f64, while FCMA f64 keeps the vector
path it wins on.

Measured on an Apple M1, full planned FFT against the scalar Bluestein's around the same inner FFT,
at lengths 107 to 65579: NEON f32 1.6 to 5.2% faster, FCMA f32 2.0 to 12.5%, FCMA f64 1.8 to 13.0%,
NEON f64 and the scalar planner unchanged.
The same treatment as SimdBluesteins, reusing its machinery: SimdRaders is written against
SimdVector, so the four SIMD planners share one implementation instead of wrapping the scalar
RadersAlgorithm, which becomes an alias for the single-complex-per-vector case. No new per-backend
code was needed, only the pairwise multiply loops Bluestein's already added.

Only the multiply between the two inner FFTs is SIMD. The input and output reorderings stay scalar
because both are permuted copies, and without a gather instruction there is nothing for a vector to
do. That caps the win: measured on an Apple M1 at lengths 101 to 64153, FCMA is 1 to 5% faster and
NEON is unchanged, where Bluestein's reached 13%. The gather and scatter dominate the work outside
the inner FFTs.

The immutable and in-place paths differ only in where the inner FFT's scratch comes from, so they
share one body now rather than repeating it as the scalar version did.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant