Skip to content

Feature/var detected by cat upset - #40

Open
leventetn wants to merge 5 commits into
mainfrom
feature/var_detected_by_cat_upset
Open

leventetn wants to merge 5 commits into
mainfrom
feature/var_detected_by_cat_upset

Conversation

@leventetn

Copy link
Copy Markdown
Collaborator

Tests: add spec-guided TDD support conftest

Add tests/conftest.py from the spec-guided-tdd skill, formatted to
ProteoPy's black/flake8 standards. It only affects test modules
marked `pytestmark = pytest.mark.spec_guided` (e.g.
tests/pl/test_upset.py); all other tests run unchanged.

- `pytest` (default, CI) runs the holdout and randomized tests and
  deselects the `_IMP` mirrors
- `pytest --implementer` runs the `_IMP` and randomized tests with
  obscured failure output, for implementation sessions
- `--seed N` reproduces randomized inputs (seed printed in the header)
- provides the `rng` and `report_input` fixtures
- checks test naming and holdout/_IMP pairing
Tests: pl.var_detected_by_cat_upset()

Add a spec-guided test suite for var_detected_by_cat_upset()
(module marked `spec_guided`, run by tests/conftest.py):

- TestVarDetectedByCatUpset: 75 holdout tests plus 9 randomized
  property and metamorphic tests
- TestVarDetectedByCatUpsetIMP: 75 _IMP mirrors on different inputs,
  run only with `pytest --implementer`

Coverage: intersection counts and thresholds (min_count /
min_fraction, inclusivity, per-category denominators, empty
categories, float edge cases), zero/NaN detection, category order
and str coercion, sparse input, non-mutation (incl. views), UpSet
construction and rendered axes, print_stats tables (incl. column-name
collisions), verbose output, save/show, and argument validation.

UpSet calls are observed through spies on the upsetplot class;
randomized tests use the seeded rng fixture (reproduce with
`--seed N`).
Feature: pl.var_detected_by_cat_upset()

Add an UpSet plot of which features (.var) are detected in which
categories of an .obs column, e.g. which proteins are found in which
tissues.

- A feature is a member of a category when it is detected (non-NaN in
  .X; with zero_to_na=True also non-zero) in enough of that
  category's observations: min_count (default 1) or min_fraction,
  inclusive, per category. Setting both raises ValueError.
- Features that are members of no category form a "No category" set,
  always included, also with a count of 0.
- Categories follow the default pl order: category order for a
  Categorical column (unused categories kept), else lexicographic
  order of the str-coerced values. Values colliding after str
  coercion raise ValueError.
- print_stats prints global, per-intersection and per-category
  tables; verbose reports the input matrix, threshold and sizes.
- Sparse .X is densified with a UserWarning; adata is never modified.
- Returns the axes dict of upsetplot's UpSet.plot().

Add upsetplot>=0.9.0,<0.10 as a core dependency. upsetplot 0.9.0 is
broken on pandas 3 / numpy >= 2.4, so proteopy.pl.upset patches it on
import (dot style defaults, single-category aggregation, count-label
positions, all-empty totals warning). The upper bound keeps these
patches from hitting an untested upsetplot release.

Export via pr.pl, list it under Quality Control in the plotting API
docs, and add a HISTORY.md entry.
Docs: plotting tutorial

Add a protein-level plotting tutorial on the Karayel et al. (2020)
erythropoiesis data, walking through the pr.pl quality-control and
exploration plots, each framed by the question it answers:

- conventions: ordering samples and groups via ordered Categoricals,
  missing values vs measured zeros, working with a supplied or
  returned Axes
- coverage and completeness: n_samples_per_category,
  n_cat1_per_cat2_hist, n_proteins_per_sample,
  completeness_per_sample, completeness_per_var,
  var_detected_by_cat_upset (default and min_fraction=1.0
  thresholds)
- abundance: intensity_hist, intensity_box_per_sample,
  abundance_rank, box, cv_by_group
- structure: sample_correlation_matrix, binary_heatmap,
  hclustv_profiles_heatmap

The notebook ends by checking that the raw intensities are unchanged
by the plotting calls. The peptide-level section is a placeholder for
a later installment.

Outputs are stored in the notebook, as Sphinx does not execute
notebooks (nbsphinx_execute = 'never'). The notebook lives in
docs/tutorials/ and is linked into the Sphinx source by a symlink;
it is listed in tutorials/index.rst.
@read-the-docs-community

Copy link
Copy Markdown

Documentation build overview

📚 proteopy | 🛠️ Build #34816013 | 📁 Comparing 2e62835 against latest (3661567)

  🔍 Preview build  

7 files changed · + 3 added · ± 4 modified

+ Added

± Modified

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant