Skip to content

Feature/pairwise peptide correlations heatmap - #39

Open
leventetn wants to merge 5 commits into
mainfrom
feature/pairwise_peptide_correlations_heatmap
Open

leventetn wants to merge 5 commits into
mainfrom
feature/pairwise_peptide_correlations_heatmap

Conversation

@leventetn

Copy link
Copy Markdown
Collaborator

Add pr.pl.pairwise_peptide_correlations_heatmap

Purpose

COPF assigns peptides to proteoforms based on their pairwise correlations, but ProteoPy had no way to look at those correlations. This PR adds a per-protein heatmap, so users can check whether the proteoforms COPF reports actually form correlated peptide blocks.

Key changes

  • New plot: pr.pl.pairwise_peptide_correlations_heatmap(adata, protein, ...) in pl/copf.py
    • Draws one protein per call from adata.uns["pairwise_peptide_correlations"] (output of pr.tl.pairwise_peptide_correlations()).
    • Clusters rows and columns with the same tree, built from 1 − PCC. method sets the linkage method, linkage accepts a precomputed tree, and cluster=False turns clustering off.
    • margin_color adds annotation strips from any .var column(s). The default is proteoform_id from pr.tl.peptide_clusters_from_dendograms(). Strips never share a colour, and each strip gets its own legend, placed clear of all labels.
    • Missing data:
      • NaN correlations are drawn hatched, treated as r = 0 when clustering, and reported in a warning.
      • Peptides with no correlations appear as empty rows and columns.
      • Peptides filtered out after the correlations were computed are dropped, with a warning.
    • Input checks:
      • Clear errors for a per-batch correlation table, a precomputed linkage of the wrong size, and an invalid margin_color or save.
      • A warning when there are fewer than 3 samples.
    • Follows AGENTS.md: labels come from var["peptide_id"], the default ordering rule applies when cluster=False, and print_stats / show / save behave as in other plots.
  • Refactor (separate commit): the private helper utils.copf.reconstruct_corrs_df_symmetric_from_long_df became utils._matrix_wrangling.reconstruct_symmetric_matrix_from_long.
    • It is now generic, raises instead of using assert, and has an allow_missing option. It is not part of the public API.
    • In tl/copf.py and test_copro.py, only the rename and black formatting changed; an AST comparison confirms this.
  • Test fixes: the Karayel test now passes observed=False, and the multi-protein-mapping test no longer uses duplicate .var names.

Tests / lint

  • New tests: 47 for the heatmap and 10 for the helper.
  • Full suite: pytest tests/ gives 547 passed.
  • flake8 is clean; pylint (E/F) scores 10.00/10.

proteopy/utils/copf.py → new proteopy/utils/_matrix_wrangling.py with
reconstruct_symmetric_matrix_from_long(df, var_a_col, var_b_col,
value_col, *, diagonal=1.0) — generic naming/docstring, diagonal param,
and a raise ValueError (with np.isclose) replacing the bare assert. Not
exported anywhere (private).
Updated the 3 production call sites in tl/copf.py and the test imports;
removed the stray test that lived in tests/utils/helpers.py.
New tests/utils/test_matrix_wrangling.py — 9 short one-concern tests
(basic reconstruction, symmetry, mirroring, diagonal, int-vs-str
selectors, generic labels, single pair, missing-pair raise, conflict
raise)
pr.pl.pairwise_peptide_correlations_heatmap(adata, protein, ...) in
pl/copf.py, registered in pl/init.py. Single protein; reads
adata.uns["pairwise_peptide_correlations"], reconstructs the per-protein
matrix, draws a sns.clustermap with symmetric row/col clustering
(method, unified linkage shadow, cluster toggle) and one-or-more
margin_color annotation strips (default cluster_id). Tick labels from
var["peptide_id"]; 1e6 NOISE left as an ordinary category; full numpydoc
docstring.
New tests/pl/test_pairwise_peptide_correlations_heatmap.py — 16 tests
(signature lock, return semantics, symmetry, linkage pass-through,
cluster=False, str-vs-list strips, peptide_id labels, validation,
save/non-mutation).
…ourhood_union

the pandas test now sets observed=False, and the peptide test now
represents multiple protein mappings without duplicate variable names
Silently wrong output:
A table with several correlations per peptide pair (e.g. …_by_batch)
plotted whichever batch came last. It now raises.
A precomputed linkage of the wrong size drew a wrong tree. It now raises
with the expected shape.

Crashes:
Missing values in an annotation column: drawn light gray, with an "NA"
legend entry.
Colour dicts: color_scheme dicts keyed by the column's real values (e.g.
{0.0: …}) now work.
NaN correlations (missing intensities, constant peptides): cells are
drawn hatched with a legend entry, and are treated as r = 0 for
clustering, with a warning. NaN no longer blocks the default
cluster=True. The helper gained allow_missing=False, which the heatmap
sets to True.
Peptides filtered out after the correlations were computed: left out
with a warning.
print_stats without a cluster_id column: it now prints stats per
annotation column, in legend order.
pdf/svg backends: the plot now works when saved as vector graphics.

Layout and readability:
One legend per annotation column, placed to the right of all labels. The
figure grows if needed, and plt.tight_layout() was removed because it
left a gap under the annotation strip.
Several annotation strips never share a colour.
The title sits on the dendrogram.
Peptides that have no correlations at all are shown as empty rows
instead of disappearing.
Tick labels default to "auto" so labels on large proteins don't overlap.

Defaults, validation and AGENTS.md compliance:
margin_color defaults to "proteoform_id".
With cluster=False, a categorical peptide_id order is respected.
Clear errors for an empty or repeated margin_color and for missing COPF
columns (with a hint to run pr.tl.peptide_clusters_from_dendograms).
An invalid save type is rejected before drawing.
A warning when there are fewer than 3 samples.
The docstring now has Raises, Warns and Notes sections (data scale, NaN
behaviour of the upstream correlation).
Tests: heatmap tests went from 16 to 47 and helper tests from 9 to 10.
Each new test that targets a fix fails on the earlier version.
@read-the-docs-community

Copy link
Copy Markdown

Documentation build overview

📚 proteopy | 🛠️ Build #34802534 | 📁 Comparing 6c67611 against latest (3661567)

  🔍 Preview build  

3 files changed
± changelog.html
± _modules/proteopy/pl/copf.html
± _modules/proteopy/tl/copf.html

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant