Skip to content

Python scanner adds 'Python code was detected' NOTE to reports for deposits with no Python code #106

Description

@larsvilhuber

Problem

The Python scanner adds this NOTE to REPLICATION.md (via generated/software-warnings.md) even when a deposit contains no Python code at all:

[NOTE] Python code was detected and scanned. Please compare the identified packages against the requirements stated in the README. See Appendix: Candidate Python packages.

The Python appendix then shows No data.. Observed on AEAREP-10180, a Stata-only deposit whose generated/manifest.txt lists 0 .py/.ipynb files. The scripts in that case repo are identical to current master for all three files below.

Root cause

The pipeline treats "the scan's output file exists" as "Python was detected". pipreqs always creates that file.

  1. bitbucket-pipelines.yml, "Run Python parser" step: runs automations/15_run_python_scanner.sh on every case. It skips only when SkipProcessing=yes or ProcessPython=no.
  2. automations/15_run_python_scanner.sh line 30: runs pipreqs . --savepath ../generated/requirements-generated.txt without checking for Python files first. With no .py files, pipreqs still writes the file, containing only a single \n.
  3. tools/filter_requirements.py line 126: scan_ran = os.path.isfile(args.scanned) is therefore True. With no author requirements.txt, lines 140–142 write a header-only python-deps.csv (just Packages) and call write_warning(escalated=False). That writes generated/software-warnings-python.md containing the NOTE.
  4. automations/24_amend_report.sh lines 73–80: concatenates any existing fragment into software-warnings.md, so the NOTE lands in the report. csv2md.py turns the header-only CSV into No data. in the appendix.

The R scanner is not affected, because it only writes its fragment when generated/r-deps-summary.csv exists (automations/14_run_r_scanner.sh line 32).

Reproduce:

printf '\n' > /tmp/scan.txt
python3 tools/filter_requirements.py --author /nonexistent/requirements.txt \
  --scanned /tmp/scan.txt --deps-csv /tmp/deps.csv --warnings /tmp/warn.md
cat /tmp/warn.md   # -> "Python code was detected and scanned..."

Proposed fix

1. Skip the scan when there are no Python files (automations/15_run_python_scanner.sh)

This is shell plumbing, so it belongs in the .sh. Insert it right after projectID=$1, before the pip install:

# Skip entirely when the deposit contains no Python code; pipreqs would
# otherwise emit an empty requirements file that downstream reads as "scanned".
if [ -z "$(find "$projectID" -type f \( -name '*.py' -o -name '*.ipynb' \) -print -quit)" ]
then
  echo "No Python files found in $projectID; skipping Python scan."
  rm -f generated/software-warnings-python.md generated/requirements-generated.txt \
        generated/python-deps.csv generated/python-deps.md
  exit 0
fi

The rm -f is needed because generated/ is committed and persists across pipeline runs. Without it, a fragment left by an earlier run would keep being picked up by 24_amend_report.sh.

2. Handle an empty scan result (tools/filter_requirements.py)

This covers Python files that exist but import only standard-library modules, where pipreqs again emits a lone newline. Replace lines 137–138:

    scanned_lines = read_lines(args.scanned)
    scanned_entries, _ = parse_requirements(args.scanned)

with:

    scanned_lines = read_lines(args.scanned)
    if not scanned_lines:
        # pipreqs always writes --savepath, even with no .py files or only
        # stdlib imports (a lone newline). Nothing to compare: emit no warning.
        if os.path.exists(args.warnings):
            os.remove(args.warnings)
        if author_exists:
            write_deps_csv(args.deps_csv, read_lines(args.author))
        print("pipreqs found no third-party imports; no Python warning written.")
        return
    scanned_entries, _ = parse_requirements(args.scanned)

This keeps an author-provided requirements.txt in the appendix when one exists, and drops only the misleading note. Also update the module docstring: it currently says a fragment is written "whenever a scan was run".

3. Optional: notebooks

pipreqs ignores .ipynb files unless given --scan-notebooks. Once guard 1 counts notebooks as Python, a notebook-only deposit would run the scan and find nothing. Consider adding --scan-notebooks to the pipreqs call on line 30 (check that the pinned pipreqs version supports it).

Tests

Add tests/test_filter_requirements.py, following repo conventions: top-level tests/, importing from tools/ via sys.path.insert(...), runnable with python3 tests/test_filter_requirements.py. Cover:

  • empty scan (\n only) with no author file: no warnings file written, and a pre-existing warnings file is removed;
  • empty scan with a curated author requirements.txt: CSV reflects the author file, no warnings file;
  • non-empty scan with no author file: current behavior (CSV from the scan, non-escalated NOTE);
  • non-empty scan with a conda-dump author file: current escalated behavior is unchanged.

Acceptance criteria

  • On a deposit with no .py/.ipynb files, generated/software-warnings.md contains no Python NOTE, and no requirements-generated.txt, python-deps.csv, or python-deps.md is left in generated/.
  • On a deposit with Python files that import third-party packages, the behavior is unchanged.
  • No inline python3 -c is added to the YAML or shell scripts.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions