Conversation
proteopy/utils/copf.py → new proteopy/utils/_matrix_wrangling.py with reconstruct_symmetric_matrix_from_long(df, var_a_col, var_b_col, value_col, *, diagonal=1.0) — generic naming/docstring, diagonal param, and a raise ValueError (with np.isclose) replacing the bare assert. Not exported anywhere (private). Updated the 3 production call sites in tl/copf.py and the test imports; removed the stray test that lived in tests/utils/helpers.py. New tests/utils/test_matrix_wrangling.py — 9 short one-concern tests (basic reconstruction, symmetry, mirroring, diagonal, int-vs-str selectors, generic labels, single pair, missing-pair raise, conflict raise)
pr.pl.pairwise_peptide_correlations_heatmap(adata, protein, ...) in pl/copf.py, registered in pl/init.py. Single protein; reads adata.uns["pairwise_peptide_correlations"], reconstructs the per-protein matrix, draws a sns.clustermap with symmetric row/col clustering (method, unified linkage shadow, cluster toggle) and one-or-more margin_color annotation strips (default cluster_id). Tick labels from var["peptide_id"]; 1e6 NOISE left as an ordinary category; full numpydoc docstring.
New tests/pl/test_pairwise_peptide_correlations_heatmap.py — 16 tests (signature lock, return semantics, symmetry, linkage pass-through, cluster=False, str-vs-list strips, peptide_id labels, validation, save/non-mutation).
…ourhood_union the pandas test now sets observed=False, and the peptide test now represents multiple protein mappings without duplicate variable names
Silently wrong output:
A table with several correlations per peptide pair (e.g. …_by_batch)
plotted whichever batch came last. It now raises.
A precomputed linkage of the wrong size drew a wrong tree. It now raises
with the expected shape.
Crashes:
Missing values in an annotation column: drawn light gray, with an "NA"
legend entry.
Colour dicts: color_scheme dicts keyed by the column's real values (e.g.
{0.0: …}) now work.
NaN correlations (missing intensities, constant peptides): cells are
drawn hatched with a legend entry, and are treated as r = 0 for
clustering, with a warning. NaN no longer blocks the default
cluster=True. The helper gained allow_missing=False, which the heatmap
sets to True.
Peptides filtered out after the correlations were computed: left out
with a warning.
print_stats without a cluster_id column: it now prints stats per
annotation column, in legend order.
pdf/svg backends: the plot now works when saved as vector graphics.
Layout and readability:
One legend per annotation column, placed to the right of all labels. The
figure grows if needed, and plt.tight_layout() was removed because it
left a gap under the annotation strip.
Several annotation strips never share a colour.
The title sits on the dendrogram.
Peptides that have no correlations at all are shown as empty rows
instead of disappearing.
Tick labels default to "auto" so labels on large proteins don't overlap.
Defaults, validation and AGENTS.md compliance:
margin_color defaults to "proteoform_id".
With cluster=False, a categorical peptide_id order is respected.
Clear errors for an empty or repeated margin_color and for missing COPF
columns (with a hint to run pr.tl.peptide_clusters_from_dendograms).
An invalid save type is rejected before drawing.
A warning when there are fewer than 3 samples.
The docstring now has Raises, Warns and Notes sections (data scale, NaN
behaviour of the upstream correlation).
Tests: heatmap tests went from 16 to 47 and helper tests from 9 to 10.
Each new test that targets a fix fails on the earlier version.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add
pr.pl.pairwise_peptide_correlations_heatmapPurpose
COPF assigns peptides to proteoforms based on their pairwise correlations, but ProteoPy had no way to look at those correlations. This PR adds a per-protein heatmap, so users can check whether the proteoforms COPF reports actually form correlated peptide blocks.
Key changes
pr.pl.pairwise_peptide_correlations_heatmap(adata, protein, ...)inpl/copf.pyadata.uns["pairwise_peptide_correlations"](output ofpr.tl.pairwise_peptide_correlations()).methodsets the linkage method,linkageaccepts a precomputed tree, andcluster=Falseturns clustering off.margin_coloradds annotation strips from any.varcolumn(s). The default isproteoform_idfrompr.tl.peptide_clusters_from_dendograms(). Strips never share a colour, and each strip gets its own legend, placed clear of all labels.linkageof the wrong size, and an invalidmargin_colororsave.var["peptide_id"], the default ordering rule applies whencluster=False, andprint_stats/show/savebehave as in other plots.utils.copf.reconstruct_corrs_df_symmetric_from_long_dfbecameutils._matrix_wrangling.reconstruct_symmetric_matrix_from_long.assert, and has anallow_missingoption. It is not part of the public API.tl/copf.pyandtest_copro.py, only the rename and black formatting changed; an AST comparison confirms this.observed=False, and the multi-protein-mapping test no longer uses duplicate.varnames.Tests / lint
pytest tests/gives 547 passed.flake8is clean;pylint(E/F) scores 10.00/10.