Conversation
Every SIMD planner wrapped the scalar BluesteinsAlgorithm around a SIMD inner FFT. SimdBluesteins is written against SimdVector instead, so there is one implementation for all four backends rather than four, and the scalar BluesteinsAlgorithm becomes an alias for the single-complex-per-vector case of it. The three pairwise multiply loops are the only SIMD part. They go through new SimdVector methods so that they land inside the backend's target feature, and use a new mul_complex_conjugated primitive: (a * b).conj() is a.conj() * b.conj(), so storing the precomputed side pre-conjugated turns every conjugated product into one operation. FCMA reaches it in the same instruction count as a plain complex multiply, by swapping vcmla_rot90 for vcmla_rot270. With one complex number per vector, interleaved data and a shuffle-based multiply, those loops lose to the plain scalar loop LLVM autovectorizes into deinterleaving loads: NEON f64 came out 3 to 5% slower at the larger sizes. The new FUSED_COMPLEX_MULTIPLY says whether a backend has native complex multiply instructions, and the loops take the scalar path when it is false and a vector holds a single complex number. That covers NEON, SSE and wasm f64, while FCMA f64 keeps the vector path it wins on. Measured on an Apple M1, full planned FFT against the scalar Bluestein's around the same inner FFT, at lengths 107 to 65579: NEON f32 1.6 to 5.2% faster, FCMA f32 2.0 to 12.5%, FCMA f64 1.8 to 13.0%, NEON f64 and the scalar planner unchanged.
The same treatment as SimdBluesteins, reusing its machinery: SimdRaders is written against SimdVector, so the four SIMD planners share one implementation instead of wrapping the scalar RadersAlgorithm, which becomes an alias for the single-complex-per-vector case. No new per-backend code was needed, only the pairwise multiply loops Bluestein's already added. Only the multiply between the two inner FFTs is SIMD. The input and output reorderings stay scalar because both are permuted copies, and without a gather instruction there is nothing for a vector to do. That caps the win: measured on an Apple M1 at lengths 101 to 64153, FCMA is 1 to 5% faster and NEON is unchanged, where Bluestein's reached 13%. The gather and scatter dominate the work outside the inner FFTs. The immutable and in-place paths differ only in where the inner FFT's scratch comes from, so they share one body now rather than repeating it as the scalar version did.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements #190: one Bluestein's and one Rader's shared by all four SIMD backends, written against
SimdVector.SimdBluesteinsandSimdRadersinsrc/simd, which the SIMD planners now build instead of wrapping the scalar versions. The scalarBluesteinsAlgorithmandRadersAlgorithmbecome aliases for the single-complex-per-vector case, the wayRadixNalready is.mul_complex_conjugatedonSimdVector:(a * b).conj()isa.conj() * b.conj(), so storing the precomputed side pre-conjugated turns every conjugated product into one operation. On FCMA it costs the same as a plain multiply,vcmla_rot270instead ofvcmla_rot90.FUSED_COMPLEX_MULTIPLY, because with a single complex per vector and a shuffle-based multiply these loops lose to the scalar loop LLVM autovectorizes into deinterleaving loads. NEON f64 was 3 to 5% slower without it. NEON, SSE and wasm f64 take the scalar path, FCMA f64 keeps the vector path.boilerplate_simd_fft!besideboilerplate_fft!, for theFftimpl of theSimdVector-generic algorithms.Apple M1, planned FFT at lengths 107 to 65579 against the old path: Bluestein's is 1.6 to 5.2% faster on NEON f32 and 1.8 to 13.0% on FCMA, Rader's 1 to 5% on FCMA. NEON f64 and the scalar planner are unchanged. Rader's gains less because its reorderings are permuted copies with nothing for a vector to do.