Repository navigation
Conversation
test_subgradient_alternating_gaps_are_all_opens is a strict xfail: a gap in one sequence does not reset the other's in-gap flag, so gap-A, gap-B, gap-A is counted as 2 opens + 1 extend instead of 3 opens.
Initial estimators are checked on fake alignments with prescribed count features drawn from a known logistic model, which pins down the mapping from regression coefficients to parameter names and matrix orientation.
Checks that every alignment score equals its count vector dotted with the parameters, and that the model gradient built from logit_subgradient matches central finite differences of the log-likelihood in all twelve mode combinations. The reference log-likelihood is computed from logits, since logit_logL loses precision on confident predictions.
The core check: one iteration from known parameters equals theta0 plus step times the model gradient, recomputed independently from Biopython alignments at the returned alpha, in all twelve mode combinations. Also covers result structure and consistency, zero-step and zero-iteration runs, input immutability, thread-count and subgradient_scale invariance, seeded noise, alphabet inference and default baselines. test_default_stepfunction_runs is a strict xfail: stepfunction defaults to None and is called unconditionally.
The learning tests fit every mode combination on simulated homologs, check that the fit improves and generalizes to held-out pairs, and check that a planted A<->G substitution preference is recovered.
logit_subgradient walks each alignment column by column and classifies
every gap column as an open or an extend, using one in-gap flag per
sequence. Only a match/mismatch column cleared the flags. A gap in one
sequence did not clear the other sequence's flag, so a run such as
seq1: - A - B
seq2: C - D -
was counted as 2 opens + 2 extends. The third column was taken to
continue the gap that started in column 1, although a gap in seq2 had
closed it in between. Biopython's PairwiseAligner scores every switch
between the two sequences' gaps as a new open, so this alignment scores
4 x open. The subgradient therefore was not the count vector of the path
that produced the score, and the open/extend components of the gradient
were wrong.
The situation is unreachable while extend_gap_score >= open_gap_score.
Any gapX, gapY, gapX run can be reordered to gapX, gapX, gapY with the
same residues and substitution columns, at the cost of open + extend +
open instead of three opens, so the optimum never needs to alternate.
Nothing in the subgradient iteration constrains the two scores, though,
and once the fit drifted to extend < open the gradient reported phantom
extensions and pushed extend_gap_score in response to gaps that no
alignment contained.
Linear gap mode was unaffected, since it only uses opens + extends, the
number of gap columns. The initial estimate was also unaffected, since
it uses Biopython's own aln.counts(), which was already correct.
The fix resets the other sequence's flag in each gap branch, so a gap
column continues a gap only if the previous column was a gap in the same
sequence. The two strict-xfail regression tests now pass, and their
markers are removed.
…timate In symmetric and general substitution mode, get_initial_estimate() fitted the logistic regression on two gap features, the numbers of gap opens and gap extends, whatever the gap model. For linear gaps it then reported gap_score = open coefficient + extend coefficient. A linear model gives every gap column the same score, so the right feature is the total number of gap columns, with one coefficient. Summing the two coefficients of an affine fit is not that quantity. In this regression each coefficient already prices a whole gap column (an open column, or an extend column), so their sum double-counts. On synthetic data labelled with a per-column gap coefficient of -0.6, the old estimator returned -1.26 and the new one returns -0.64. Simple substitution mode already did this correctly: it fits one coefficient on counts.gaps. The fix had been parked because it did not explain the positive initial gap scores seen on one fixture. There the alignments had no gap extensions at all, so both estimators agreed exactly, and the bonus was a genuine fit. It is applied now because the nwgrad backend needs it. The initial estimate is moving onto nwgrad's per-pair gradients, which are the count vectors of the alignments. In linear mode nwgrad reports only the number of gap columns, since a linear model has no open/extend distinction to differentiate. So the open/extend split that the old estimator required cannot be computed from nwgrad at all, while the total gap count can be computed identically by both backends. Without this change the two backends could not start from the same initial estimate in these two modes, even on tie-free data. This changes the Biopython backend's initial estimate in the symmetric/general + linear modes, by design. Affine modes, simple mode, and every run given initial_parameters are unaffected. test_initial_estimate_full_linear_merges_gap_coefficients, which pinned the old sum, is replaced by a test that recovers a known per-column gap coefficient.
…imate_from_counts
…lpha_solver='bfgs'
… fit_alpha, fewer temporaries in the log-likelihood (bit-identical)
…wgrad backend when available
…to nwgrad-backend
…r on short pairs); CI builds nwgrad from its perf-dp branch
…vector), same fit
… of realigning with Biopython
…h DISCRIMALIGN_RUN_MIRBENCH=1
…ing, raises on non-finite; both backends use it
…teration, numpy path in logit_link.logistic_step
…ture and results['alphabet']
Implement nwgrad backend
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implement nwgrad backend