Align loop points at the sample level - #78
Closed
Xehanort88 wants to merge 3 commits into
Closed
Xehanort88 wants to merge 3 commits into
Xehanort88 wants to merge 3 commits into
Conversation
Tests generate their own audio: an intro followed by a repeated 8-note pattern, so the correct loop points are known in advance and loop detection can be checked for correctness, not just for running. Covers audio loading, loop detection and scoring, zero-crossing snapping, split/extend/tag/txt exports, the interactive picker, batch file discovery and every CLI command. Known bugs are marked as strict xfail, so each fix is flagged until its marker is removed. Also adds pytest as a dev dependency and a CI workflow running the tests on Ubuntu and Windows with Python 3.10 and 3.13.
- Mono tracks were played and exported louder than the original: the analysis-only normalization was applied in place to the audio array that is also used for playback, since to_mono returns its input as-is for mono audio. - Longer loops were never preferred among near-identical scores: _prioritize_duration ran before loop_start/loop_end were set, so every duration it compared was 0. - Loops starting in the first seconds of a track were underscored: the truncated look-behind window was zero-padded on the side nearest the loop point, where the weights are heaviest. - Interactive mode discarded the choice made after 'more', 'all' or 'reset' and prompted again. - extend with fade_length=0 crashed (x[-0:] selects the whole array). - export_tags() without output_dir used the file path as the directory. - The txt export message named loop.txt instead of loops.txt.
Loop points were located on STFT frames (512 samples) and each was then moved to its own nearest zero crossing, without checking that the waveforms at the two points line up. Loops were often off by a few hundred samples, which can be heard as a faint flam or phase smear at the jump. Each candidate is now refined by: - moving the loop end (by up to 2 frames) to where the waveform around it best matches the waveform around the loop start, using normalized cross-correlation over 75 ms on each side; - shifting both points by the same amount (keeping the loop length) to where the waveforms differ the least within +/-5 ms, since the difference at the jump is what causes clicks. When the correlation is below 0.3, the audio at the two points does not match well enough for this to be reliable, and each point is moved to its nearest zero crossing as before. The threshold was chosen by comparing both methods on 181 candidates from real tracks. On game tracks with official loop tags, all top detected loops now have exactly the official loop length (or a multiple of it). The added cost is about 0.5 ms per candidate. Also declares scipy as a direct dependency (already required by librosa).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Loop points are found on STFT frames (512 samples, ~11 ms), and each point is then moved to its own nearest zero crossing. Nothing checks that the waveforms at the two points actually line up, so loops are often a few hundred samples off. You can hear it as a faint flam or phase smear at the jump, even when the loop is musically correct.
Change
Each candidate is refined by a new
_refine_loop_points:_align_loop_end): the loop end moves by up to 2 frames (±1024 samples) to where the waveform around it best matches the waveform around the loop start. This uses normalized cross-correlation over 75 ms on each side, computed withscipy.signal.correlate(FFT) and a cumulative-sum energy normalization._best_seam): both points then shift by the same amount, keeping the loop length, to where the waveforms differ the least within ±5 ms. The difference at the jump is what causes clicks.scipyis declared as a direct dependency. It's already required by librosa (>=1.6.0), so nothing new gets installed.Results
Official loop tags as ground truth. On two game tracks that ship with official loop tags, all 20 top detected loops now have exactly the official loop length (or twice it), to the sample. Before, the best loop on one of them was 117 samples short.
Real tracks. Top candidates from 6 real tracks (WAV, FLAC and Ogg; game music and orchestral), compared on the same candidates. Mismatch = relative RMS difference between the audio around the loop end and around the loop start (±20 ms); jump = size of the waveform discontinuity at the seam, relative to the track's RMS. Lower is better for both.
Choosing the fallback threshold. 181 candidates, grouped by correlation: alignment beats the old method on mismatch in 78–100% of cases from 0.3 upward. Below 0.3 it's no better, so the old behaviour is kept there. Moving both points to the best-matching seam without aligning first was worse than the old method, so the gain comes from the alignment step.
Cost. About 0.5 ms per candidate.
Test plan
uv run pytest: 63 passed (11 new)