Offline document analysis for LLM pipelines.
Core library:
complydoc– command line tool and Python library that reports on documents before they enter an LLM pipeline: token cost across models and extraction paths, extraction readiness, personal and financial identifiers (masked), and hidden or instruction-like content. Runs locally, with outbound network access blocked.
Loader and pipeline tooling, in the same package:
inspect_documentsandcompare_loaders– what LangChain and LlamaIndex loaders extract, the metadata they attach and the connections they attempt, compared side by side with expected factsinspect_chunksandcomplydoc chunks– chunk sizes, sentences and tables cut at a boundary, identifiers repeated across chunksMaskIdentifiers,DropHiddenPassages,StripPathMetadata– pipeline steps for LangChain and LlamaIndex documentsdiff_reports,expectandcomplydoc diff– baselines and assertions for tests and CI
Examples:
playground– runnable command line and Python examples on sample documents, with a CI workflow that fails when a report regresses against a baseline
Learn more:
- Documentation – guides, command line and Python API reference, report JSON schema (source)
- PyPI –
uv tool install complydocorpip install complydoc - Changelog
- Issues – bug reports and feature requests