Skip to content

Nwgrad - #14

Merged
mciach merged 69 commits into
mainfrom
nwgrad
Oct 7, 2026
Merged

mciach merged 69 commits into
mainfrom
nwgrad

Conversation

@mciach

@mciach mciach commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Implement nwgrad backend

mciach and others added 30 commits September 28, 2026 12:10
test_subgradient_alternating_gaps_are_all_opens is a strict xfail: a gap
in one sequence does not reset the other's in-gap flag, so gap-A, gap-B,
gap-A is counted as 2 opens + 1 extend instead of 3 opens.
Initial estimators are checked on fake alignments with prescribed count
features drawn from a known logistic model, which pins down the mapping
from regression coefficients to parameter names and matrix orientation.
Checks that every alignment score equals its count vector dotted with the
parameters, and that the model gradient built from logit_subgradient
matches central finite differences of the log-likelihood in all twelve
mode combinations. The reference log-likelihood is computed from logits,
since logit_logL loses precision on confident predictions.
The core check: one iteration from known parameters equals theta0 plus
step times the model gradient, recomputed independently from Biopython
alignments at the returned alpha, in all twelve mode combinations. Also
covers result structure and consistency, zero-step and zero-iteration
runs, input immutability, thread-count and subgradient_scale invariance,
seeded noise, alphabet inference and default baselines.

test_default_stepfunction_runs is a strict xfail: stepfunction defaults
to None and is called unconditionally.
The learning tests fit every mode combination on simulated homologs, check
that the fit improves and generalizes to held-out pairs, and check that a
planted A<->G substitution preference is recovered.
logit_subgradient walks each alignment column by column and classifies
every gap column as an open or an extend, using one in-gap flag per
sequence. Only a match/mismatch column cleared the flags. A gap in one
sequence did not clear the other sequence's flag, so a run such as

    seq1:  - A - B
    seq2:  C - D -

was counted as 2 opens + 2 extends. The third column was taken to
continue the gap that started in column 1, although a gap in seq2 had
closed it in between. Biopython's PairwiseAligner scores every switch
between the two sequences' gaps as a new open, so this alignment scores
4 x open. The subgradient therefore was not the count vector of the path
that produced the score, and the open/extend components of the gradient
were wrong.

The situation is unreachable while extend_gap_score >= open_gap_score.
Any gapX, gapY, gapX run can be reordered to gapX, gapX, gapY with the
same residues and substitution columns, at the cost of open + extend +
open instead of three opens, so the optimum never needs to alternate.
Nothing in the subgradient iteration constrains the two scores, though,
and once the fit drifted to extend < open the gradient reported phantom
extensions and pushed extend_gap_score in response to gaps that no
alignment contained.

Linear gap mode was unaffected, since it only uses opens + extends, the
number of gap columns. The initial estimate was also unaffected, since
it uses Biopython's own aln.counts(), which was already correct.

The fix resets the other sequence's flag in each gap branch, so a gap
column continues a gap only if the previous column was a gap in the same
sequence. The two strict-xfail regression tests now pass, and their
markers are removed.
…timate

In symmetric and general substitution mode, get_initial_estimate() fitted
the logistic regression on two gap features, the numbers of gap opens and
gap extends, whatever the gap model. For linear gaps it then reported
gap_score = open coefficient + extend coefficient. A linear model gives
every gap column the same score, so the right feature is the total number
of gap columns, with one coefficient. Summing the two coefficients of an
affine fit is not that quantity. In this regression each coefficient
already prices a whole gap column (an open column, or an extend column),
so their sum double-counts. On synthetic data labelled with a per-column
gap coefficient of -0.6, the old estimator returned -1.26 and the new one
returns -0.64. Simple substitution mode already did this correctly: it fits
one coefficient on counts.gaps.

The fix had been parked because it did not explain the positive initial
gap scores seen on one fixture. There the alignments had no gap extensions
at all, so both estimators agreed exactly, and the bonus was a genuine
fit. It is applied now because the nwgrad backend needs it. The initial
estimate is moving onto nwgrad's per-pair gradients, which are the count
vectors of the alignments. In linear mode nwgrad reports only the number
of gap columns, since a linear model has no open/extend distinction to
differentiate. So the open/extend split that the old estimator required
cannot be computed from nwgrad at all, while the total gap count can be
computed identically by both backends. Without this change the two
backends could not start from the same initial estimate in these two
modes, even on tie-free data.

This changes the Biopython backend's initial estimate in the
symmetric/general + linear modes, by design. Affine modes, simple mode,
and every run given initial_parameters are unaffected.
test_initial_estimate_full_linear_merges_gap_coefficients, which pinned
the old sum, is replaced by a test that recovers a known per-column gap
coefficient.
michalsta and others added 29 commits October 2, 2026 08:32
… fit_alpha, fewer temporaries in the log-likelihood (bit-identical)
…r on short pairs); CI builds nwgrad from its perf-dp branch
…ing, raises on non-finite; both backends use it
…teration, numpy path in logit_link.logistic_step
@mciach
mciach merged commit bcbf944 into main Oct 7, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants