Skip to content

feat: add jev support - #53

Merged
joshlarsen merged 2 commits into
mainfrom
261004-add-jev-support
Oct 5, 2026
Merged

joshlarsen merged 2 commits into
mainfrom
261004-add-jev-support

Conversation

@joshlarsen

@joshlarsen joshlarsen commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Summary

Adds opt-in advisory secret classification through TypeSafe's Jev model to both the CLI and Go library. After regex detection, Poltergeist can enrich findings with a likelihood that the matched value is an authentic secret, using its exact location and nearby source as evidence. Classification never suppresses findings, changes entropy filtering, or changes finding-based exit codes: a value labeled likely_dummy is still reported as a finding.

CLI and report changes

  • Adds -classify, using TYPESAFE_API_KEY for authentication. Classification is disabled by default; setting the environment variable alone does not enable it.
  • Adds -classify-timeout (default 10s), -classify-max-candidates (default 1000), and opt-in -classify-cache-dir.
  • Normally classifies entropy-passing matches. -low-entropy also makes reported low-entropy matches eligible; -classify-all, when used with -classify, includes low-entropy candidates even when the report hides them.
  • Adds per-finding classification to text, Markdown, and JSON output, plus enrichment metrics for eligible/scored/uncertain/skipped/failed candidates, cache hits, requests, retries, token usage, and duration. JSON metrics also include reread bytes.
  • Moves startup and scan-progress messages to stderr so stdout can contain a parseable JSON report.
  • Uses an interrupt-aware context for scanning and enrichment.
TYPESAFE_API_KEY='your-api-key' poltergeist -classify -format json ./source

Enabling classification sends raw matched values, bounded surrounding source, relative paths, and rule identifiers/names to TypeSafe, including when report redaction is enabled.

Library and classification semantics

  • Adds Scanner.Classification, ClassificationOptions, an injectable CandidateClassifier interface, classification batch/result types, and NewJevClassifier with configurable credentials, versioned model, endpoint, HTTP client, and optional cache.
  • Adds ScanDirectoryContext; the existing ScanDirectory API remains available and enriches findings when configured. Detection cancellation returns an error; an enrichment deadline returns findings with explicit unscored statuses.
  • Pins the default model to jev-1.13.0 and the policy to secret-authenticity-v1; unversioned model aliases are rejected.
  • Records real_secret_probability in [0, 1], labeled likely_real at or above 0.90, likely_dummy at or below 0.10, and uncertain otherwise. Results carry model, policy, source (live, cache, or memory), and context-truncation metadata.
  • Distinguishes scored, skipped, and error, with reason codes for unscored findings. Missing probabilities are omitted rather than treated as zero.
  • Defines authenticity to include development credentials and expired/revoked keys. Scores do not establish current credential validity and are not calibrated accuracy claims or an automatic suppression policy.

Bounded context and execution

  • Selects candidates deterministically by relative path, line, byte span, and rule ID while preserving report order. HTTP enrichment runs after detection, outside the original scan workers.
  • Records exact match spans and matched-line digests, then rereads only selected files once per file. Checks file identity, size, modification time, and line content before submission; changed or unavailable files retain their findings with skipped classifications.
  • Extracts up to five neighboring lines on either side, capped at 8 KiB. Skips targets over 2 KiB and crops long target lines around the complete match using UTF-8 boundaries and byte origins.
  • Shares overlapping local context within a file, bounded to 16 targets and 16 KiB of serialized request JSON per batch.
  • Uses two reread workers, four inference workers, an eight-batch queue, and a 64 MiB aggregate reread budget. The Jev client limits requests to ten per second and request bodies to 64 KiB per second, including retries.
  • Fixes Go regex line matches to return actual byte offsets and fixes Hyperscan's regex capture refinement to use capture indices, avoiding the wrong occurrence when values repeat.

Jev client and caching

  • Sends per-candidate Noul questions with an explicit authenticity rubric that treats source, paths, comments, and rule names as untrusted data.
  • Validates the returned model, candidate answer mapping, answer types, probability range, response size, and token counts. Errors expose fixed reason codes without raw values or source.
  • Applies a two-second timeout per HTTP attempt and up to two retries for connection failures, HTTP 408/429, and 5xx responses, with exponential backoff, jitter, and retry-header handling inside the overall deadline. Authentication failures stop further submissions, and redirects are rejected.
  • Deduplicates concurrent identical requests and reuses successful responses in memory within an enrichment run; memory is reset at scan boundaries.
  • Adds an optional 24-hour persistent cache keyed by the exact canonical request, including endpoint, model, policy, context, paths, ordering, and batch composition. HMAC filenames use a randomly generated directory-local key.
  • Stores only probabilities, model, policy, and creation time in cache entries. Uses private directory/file permissions, atomic key publication and entry writes, serialized pruning across processes, and a 64 MiB storage cap. Corrupt, expired, inaccessible, or unsupported cache entries fall back nonfatally.

Documentation and validation

  • Adds docs/classification.md covering usage, transmitted data, score interpretation, limits, caching, reproducibility, and library integration; links it from the README.
  • Extends benchmark documentation with disabled/no-finding, sparse/dense, cold/warm-cache, and unavailable-provider scenarios. Recorded local results describe integration overhead, not hosted Jev latency or accuracy.
  • Adds tests for opt-in behavior, redaction, deterministic selection, low-entropy eligibility, exact/repeated/Unicode spans, context cropping and batching, file changes, malformed input, deadlines, response validation, retries, authentication, redirects, rate/request limits, concurrent deduplication, and cache identity/expiry/corruption/permissions/storage bounds.
  • Adds CLI formatter tests confirming classification appears in all three formats without changing exit codes, and that disabled classification leaves optional JSON fields absent. Automated provider tests use local fake HTTP servers.
  • CI observed at the time of this update: Go Test (x86), generated rule documentation, and secret scanning passed; lint and platform builds were still pending. No additional local test run was performed for this description-only update.

@joshlarsen
joshlarsen enabled auto-merge (squash) October 5, 2026 11:09
@joshlarsen
joshlarsen merged commit 20d3fa3 into main Oct 5, 2026
10 checks passed
@joshlarsen
joshlarsen deleted the 261004-add-jev-support branch October 5, 2026 11:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants