This project handles voiceprints — biometric data — and the recordings and transcripts they come from. Under most privacy regimes that is a sensitive category (PIPL art. 28, GDPR art. 9). A privacy bug here is therefore higher severity than an ordinary correctness bug, not lower.
- Any real recording, transcript, name, email address or embedding that is committed to this repository, or still reachable through its history
- A code path that writes real data into a tracked path (
audio/,data.js,speakers.json) without warning the user - An export that discloses more than the user asked for — a format that embeds absolute paths or hostnames, or that exports the whole corpus when a single recording was selected
- Anything that stops the bundled demo from being fully synthetic
- A dependency or build step that introduces network access at runtime
Use Security → Report a vulnerability (a private advisory), not a public issue — a public issue would repeat the exposure you are reporting.
If the problem is already public (a real name sitting in the git history, for instance), say so explicitly. That is worth knowing, and the fix is a history rewrite rather than a pull request — which only the repository owner can do.
Please include: the path or commit, what is disclosed, and whether it is still
reachable (git log vs. a direct SHA fetch are different answers —
see Removing sensitive data in the GitHub docs for why force-push alone is not
enough).
- The bundled demo is generated by
tools/make_demo.py: invented people, formant-synthesised speech, fabricated transcripts. - Every public asset — diagrams, screenshots, placeholder text, test fixtures, example names — comes from that generator. Nothing is drawn from a real corpus.
- No network access at runtime. The browser UI loads zero external resources
(the favicon is an inline
data:URI). audio/,data.jsandspeakers.jsonare tracked, because the demo ships committed. Read CONTRIBUTING before pointing the pipeline at your own corpus — those paths get overwritten in place.