Skip to content

Align loop points at the sample level - #78

Closed
Xehanort88 wants to merge 3 commits into
arkrow:masterfrom
Xehanort88:sample-level-alignment
Closed

Xehanort88 wants to merge 3 commits into
arkrow:masterfrom
Xehanort88:sample-level-alignment

Conversation

@Xehanort88

Copy link
Copy Markdown

Depends on #76. The first two commits come from that PR (the test suite and bug fixes); only the last commit, Align loop points at the sample level, is new here. Once #76 is merged, this PR will show only that commit. It's independent of #77.

Problem

Loop points are found on STFT frames (512 samples, ~11 ms), and each point is then moved to its own nearest zero crossing. Nothing checks that the waveforms at the two points actually line up, so loops are often a few hundred samples off. You can hear it as a faint flam or phase smear at the jump, even when the loop is musically correct.

Change

Each candidate is refined by a new _refine_loop_points:

  1. Line up (_align_loop_end): the loop end moves by up to 2 frames (±1024 samples) to where the waveform around it best matches the waveform around the loop start. This uses normalized cross-correlation over 75 ms on each side, computed with scipy.signal.correlate (FFT) and a cumulative-sum energy normalization.
  2. Choose the seam (_best_seam): both points then shift by the same amount, keeping the loop length, to where the waveforms differ the least within ±5 ms. The difference at the jump is what causes clicks.
  3. Fall back: if the best correlation is below 0.3, the audio at the two points doesn't match well enough for this to be reliable, so each point is moved to its nearest zero crossing as before.

scipy is declared as a direct dependency. It's already required by librosa (>=1.6.0), so nothing new gets installed.

Results

Official loop tags as ground truth. On two game tracks that ship with official loop tags, all 20 top detected loops now have exactly the official loop length (or twice it), to the sample. Before, the best loop on one of them was 117 samples short.

Real tracks. Top candidates from 6 real tracks (WAV, FLAC and Ogg; game music and orchestral), compared on the same candidates. Mismatch = relative RMS difference between the audio around the loop end and around the loop start (±20 ms); jump = size of the waveform discontinuity at the seam, relative to the track's RMS. Lower is better for both.

Track type Mismatch, old → new Jump, old → new
Exactly repeating (game music, synthetic) ~1.2–1.4 → 0.00–0.02 ~0.2–0.5 → 0.00
Orchestral recordings (never repeat exactly) ~1.2–1.5 → 0.8–1.2 mostly lower

Choosing the fallback threshold. 181 candidates, grouped by correlation: alignment beats the old method on mismatch in 78–100% of cases from 0.3 upward. Below 0.3 it's no better, so the old behaviour is kept there. Moving both points to the best-matching seam without aligning first was worse than the old method, so the gain comes from the alignment step.

Cost. About 0.5 ms per candidate.

Test plan

  • uv run pytest: 63 passed (11 new)
  • On the synthetic test track, the top loops are now exactly a whole number of patterns long, and the audio after the jump is identical sample for sample to what would have played. Before, they were ~200–300 samples off.
  • Unit tests: alignment with known offsets (including the search limit), low correlation on unrelated audio, loop end kept after loop start on very short loops, seam selection, and the zero-crossing fallback

Tests generate their own audio: an intro followed by a repeated 8-note
pattern, so the correct loop points are known in advance and loop
detection can be checked for correctness, not just for running.

Covers audio loading, loop detection and scoring, zero-crossing
snapping, split/extend/tag/txt exports, the interactive picker, batch
file discovery and every CLI command.

Known bugs are marked as strict xfail, so each fix is flagged until its
marker is removed.

Also adds pytest as a dev dependency and a CI workflow running the tests
on Ubuntu and Windows with Python 3.10 and 3.13.
- Mono tracks were played and exported louder than the original: the
  analysis-only normalization was applied in place to the audio array
  that is also used for playback, since to_mono returns its input as-is
  for mono audio.
- Longer loops were never preferred among near-identical scores:
  _prioritize_duration ran before loop_start/loop_end were set, so every
  duration it compared was 0.
- Loops starting in the first seconds of a track were underscored: the
  truncated look-behind window was zero-padded on the side nearest the
  loop point, where the weights are heaviest.
- Interactive mode discarded the choice made after 'more', 'all' or
  'reset' and prompted again.
- extend with fade_length=0 crashed (x[-0:] selects the whole array).
- export_tags() without output_dir used the file path as the directory.
- The txt export message named loop.txt instead of loops.txt.
Loop points were located on STFT frames (512 samples) and each was then
moved to its own nearest zero crossing, without checking that the
waveforms at the two points line up. Loops were often off by a few
hundred samples, which can be heard as a faint flam or phase smear at
the jump.

Each candidate is now refined by:
- moving the loop end (by up to 2 frames) to where the waveform around
  it best matches the waveform around the loop start, using normalized
  cross-correlation over 75 ms on each side;
- shifting both points by the same amount (keeping the loop length) to
  where the waveforms differ the least within +/-5 ms, since the
  difference at the jump is what causes clicks.

When the correlation is below 0.3, the audio at the two points does not
match well enough for this to be reliable, and each point is moved to
its nearest zero crossing as before. The threshold was chosen by
comparing both methods on 181 candidates from real tracks.

On game tracks with official loop tags, all top detected loops now have
exactly the official loop length (or a multiple of it). The added cost
is about 0.5 ms per candidate.

Also declares scipy as a direct dependency (already required by librosa).
@Xehanort88 Xehanort88 closed this Sep 27, 2026
@Xehanort88
Xehanort88 deleted the sample-level-alignment branch September 27, 2026 12:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant