Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
70 commits
Select commit Hold shift + click to select a range
71c099d
feat: improve Jev workflows and streamline docs
Coding-Dev-Tools Oct 3, 2026
9165f78
fix: pin README assets for PR preview
Coding-Dev-Tools Oct 3, 2026
3bec5bc
fix: bound Jev review data and disclose remote recall
Coding-Dev-Tools Oct 3, 2026
10cb5ff
fix: preserve Jev support probabilities
Coding-Dev-Tools Oct 3, 2026
b227d25
fix: refresh registered offline benchmark evidence
Coding-Dev-Tools Oct 3, 2026
8f4dcbc
docs: pin README links to refreshed evidence
Coding-Dev-Tools Oct 3, 2026
cd5f3b5
fix: require explicit consent for remote Jev eval
Coding-Dev-Tools Oct 4, 2026
fee9d0c
docs: refresh offline benchmark evidence
Coding-Dev-Tools Oct 4, 2026
24f0fd4
docs: pin README links to refreshed chart
Coding-Dev-Tools Oct 4, 2026
cf6b426
test: align contracts with pinned README links
Coding-Dev-Tools Oct 4, 2026
d7058a8
fix: type Jev advisory planner path
Coding-Dev-Tools Oct 4, 2026
d299855
docs: pin README to refreshed benchmark
Coding-Dev-Tools Oct 4, 2026
a1bdb58
test: accept refreshed benchmark image pin
Coding-Dev-Tools Oct 4, 2026
27ab1b3
fix: bind retrieval preview to grounded Jev plan
Coding-Dev-Tools Oct 4, 2026
e2acf47
docs: pin README benchmark snapshot
Coding-Dev-Tools Oct 4, 2026
a0d3e57
docs: clarify managed Jev allowance terms
Coding-Dev-Tools Oct 4, 2026
6ab0475
docs: align Jev plan details with Cloud policy
Coding-Dev-Tools Oct 4, 2026
c3bfe86
fix: screen Jev questions and correct recall fixture
Coding-Dev-Tools Oct 4, 2026
0f6d292
docs: pin README chart to corrected evidence
Coding-Dev-Tools Oct 4, 2026
52f31d6
fix: surface safe Jev fallback reasons
Coding-Dev-Tools Oct 4, 2026
b08b4e2
docs: pin Jev evidence to the updated snapshot
Coding-Dev-Tools Oct 4, 2026
7516f9c
docs: clarify pooled Team Jev allowance
Coding-Dev-Tools Oct 4, 2026
e440bf6
docs: pin pooled Jev plan guidance
Coding-Dev-Tools Oct 4, 2026
3711dc3
docs: pin README to current Jev and benchmark evidence
Coding-Dev-Tools Oct 4, 2026
4eb16d7
fix(jev): preserve review evidence and recall compatibility
Coding-Dev-Tools Oct 4, 2026
cec7326
docs: pin README to refreshed benchmark evidence
Coding-Dev-Tools Oct 4, 2026
a78a09e
Fix Jev ancestor review and positional API compatibility
Coding-Dev-Tools Oct 4, 2026
5f09d1f
fix: validate visible ancestor memories in Jev review
Coding-Dev-Tools Oct 4, 2026
2ebdb8c
Pin corrected offline benchmark snapshot in every README consumer
Coding-Dev-Tools Oct 4, 2026
4802ecc
docs: pin README to refreshed Jev review evidence
Coding-Dev-Tools Oct 4, 2026
d229892
docs: preserve pre-merge Jev evidence snapshot
Coding-Dev-Tools Oct 4, 2026
fe5d5a6
Enforce current Jev evidence visibility and serialized input bounds
Coding-Dev-Tools Oct 4, 2026
fab7441
Pin current temporal and byte-bound benchmark snapshot
Coding-Dev-Tools Oct 4, 2026
9e40e1a
Merge current Jev review fixes and evidence updates
Coding-Dev-Tools Oct 4, 2026
1d2b32c
fix: honor legacy secret classifications before Jev review
Coding-Dev-Tools Oct 4, 2026
815edc9
docs: pin Jev privacy evidence consumers to source snapshot
Coding-Dev-Tools Oct 4, 2026
cc4f894
Merge latest Jev review fixes and evidence
Coding-Dev-Tools Oct 4, 2026
3c64da4
docs: pin README to benchmark v146
Coding-Dev-Tools Oct 4, 2026
bd9b779
fix: normalize Jev advisory fallback reasons for planning off
Coding-Dev-Tools Oct 4, 2026
135bb72
docs: pin normalized Jev advisory evidence snapshot
Coding-Dev-Tools Oct 4, 2026
0c3482d
Merge current Jev privacy review and benchmark evidence
Coding-Dev-Tools Oct 4, 2026
f2f5701
docs: pin README to benchmark v147
Coding-Dev-Tools Oct 4, 2026
0bd6a8e
fix: normalize Jev fallback and refresh benchmark evidence
Coding-Dev-Tools Oct 4, 2026
e0f6ed3
docs: pin README to benchmark v148
Coding-Dev-Tools Oct 4, 2026
58082c2
fix: update vulnerable benchmark PyJWT pin
Coding-Dev-Tools Oct 4, 2026
94b8d24
docs: publish individual managed Jev limits
Coding-Dev-Tools Oct 4, 2026
fd5b060
fix: address Jev review and docs CI failures
Coding-Dev-Tools Oct 4, 2026
062bff7
chore: export fresh offline evidence for Jev review
Coding-Dev-Tools Oct 4, 2026
5834d39
chore: capture immutable offline evidence outputs
Coding-Dev-Tools Oct 4, 2026
dde3344
chore: regenerate benchmark image exports
Coding-Dev-Tools Oct 4, 2026
4a30d30
docs: register fresh offline evidence for Jev review
Coding-Dev-Tools Oct 4, 2026
7fec59d
docs: checksum fresh offline evidence
Coding-Dev-Tools Oct 4, 2026
e99d2de
docs: bind benchmark guide to refreshed artifact
Coding-Dev-Tools Oct 4, 2026
98757ff
test: bind evidence contract to refreshed artifact
Coding-Dev-Tools Oct 4, 2026
47b3362
docs: cite refreshed offline evidence
Coding-Dev-Tools Oct 4, 2026
8f8c279
docs: refresh registered context evidence chart
Coding-Dev-Tools Oct 4, 2026
f576d8b
docs: refresh registered grounded examples
Coding-Dev-Tools Oct 4, 2026
35037b6
chore: remove temporary offline evidence exporter
Coding-Dev-Tools Oct 4, 2026
c76873d
docs: pin benchmark links to refreshed evidence
Coding-Dev-Tools Oct 4, 2026
d34c692
fix(jev): keep local planning available and refresh offline evidence
Coding-Dev-Tools Oct 6, 2026
503d2cb
docs: pin README to the verified offline evidence commit
Coding-Dev-Tools Oct 6, 2026
d01bcdb
docs: give planner evidence an unambiguous artifact identity
Coding-Dev-Tools Oct 6, 2026
725ca92
docs: pin the README to the distinct planner evidence
Coding-Dev-Tools Oct 6, 2026
d1bd37d
fix(pi): pin proxy-addr 2.0.8 for GHSA-jqcg-44mw-7w3h
Coding-Dev-Tools Oct 6, 2026
12e7d48
fix(mcp): reject managed custom decisions before credential access
Coding-Dev-Tools Oct 6, 2026
e21e321
docs: pin guides and visuals to the managed preflight fix
Coding-Dev-Tools Oct 6, 2026
21a6b82
fix(jev): preserve local plans and honor dashboard consent
Coding-Dev-Tools Oct 7, 2026
de04e64
docs: pin benchmark links to verified PR252 evidence
Coding-Dev-Tools Oct 7, 2026
f964429
fix(jev): preserve offline review and enforce advisory deadlines
Coding-Dev-Tools Oct 8, 2026
cc1f2c0
docs: pin benchmark evidence to the verified source checkpoint
Coding-Dev-Tools Oct 8, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude-plugin/skill-assets.sha256
Original file line number Diff line number Diff line change
Expand Up @@ -3,4 +3,4 @@ c5d0c26f28c9ee14092f9deaf24c98dd8bef49d971fef2b7a537ffb1ab9f2887 .claude-plugin
aeee7a94671ceb306fe2d24c5acc9f2d96ad8a8e7410536566799eea6265f080 skills/engraphis-memory/SKILL.md
055655db84af07561d002f0c69744313d8413c39f3e873f941f0fa0b1e76dc66 skills/engraphis-memory/references/CONVENTIONS.md
9d090a03f5b3f36a34d91f66b72c3844591f6915755ac3a6c6ba5f1b16977de5 skills/engraphis-memory/references/SCOPING.md
68cd0add0d1025d495da2163e7c462b76b5fd1a4ff86b527fee49a4d8d6ab29b skills/engraphis-memory/references/TOOLS.md
db1db0df676e4c05fd80b878a1bd1ffc9f2045c0b0e01c0486b69ce0a5541b73 skills/engraphis-memory/references/TOOLS.md
25 changes: 17 additions & 8 deletions .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -1251,12 +1251,8 @@ jobs:
notes_args=()
latest_args=(--latest)
promote_latest=true
if [ "$WAIVE_QUALIFICATION" = "true" ]; then
# A retained older repair must not displace a newer public Latest.
# The candidate comparison uses the existing public release history;
# failed lookups or unknown tag formats stop before any release write.
current_latest="$(gh release view --repo "$GH_REPO" --json tagName --jq .tagName)"
promote_latest="$(python - "$RELEASE_TAG" "$current_latest" <<'PY'
can_promote_latest() {
python - "$RELEASE_TAG" "$1" <<'PY'
import re
import sys

Expand All @@ -1267,7 +1263,13 @@ jobs:

print("true" if version(sys.argv[1]) >= version(sys.argv[2]) else "false")
PY
)"
}
if [ "$WAIVE_QUALIFICATION" = "true" ]; then
# A retained older repair must not displace a newer public Latest.
# The candidate comparison uses the existing public release history;
# failed lookups or unknown tag formats stop before any release write.
current_latest="$(gh release view --repo "$GH_REPO" --json tagName --jq .tagName)"
promote_latest="$(can_promote_latest "$current_latest")"
latest_args=(--latest=false)
if [ "$promote_latest" = "true" ]; then latest_args=(--latest); fi
{
Expand All @@ -1293,7 +1295,14 @@ jobs:
--repo "$GH_REPO" \
--clobber
if [ "$WAIVE_QUALIFICATION" = "true" ] && [ "$promote_latest" = "true" ]; then
gh release edit "$RELEASE_TAG" --repo "$GH_REPO" --latest
# Another publisher may have advanced Latest during asset upload.
# Recheck immediately before promotion; this narrows the race but
# does not provide an atomic lease across external publishers.
current_latest="$(gh release view --repo "$GH_REPO" --json tagName --jq .tagName)"
promote_latest="$(can_promote_latest "$current_latest")"
if [ "$promote_latest" = "true" ]; then
gh release edit "$RELEASE_TAG" --repo "$GH_REPO" --latest
fi
fi
else
gh release create "$RELEASE_TAG" verified-dist/* release-evidence/* \
Expand Down
50 changes: 39 additions & 11 deletions BENCHMARKS.md
Original file line number Diff line number Diff line change
Expand Up @@ -93,15 +93,15 @@ interpretation and do not count as additional benchmark-quality gains.

### Public numeric evidence registry

Every exact public aggregate retained below comes from the checked-in, public-safe
[`offline-fixtures-v129.json`](docs/benchmark-evidence/offline-fixtures-v129.json) artifact. Its
SHA-256 is
`686394aa5d69cd296360defc5170d956c85565a5282c7766dc6466ba4c6aa6a5`, also recorded in the
Every exact public aggregate retained below comes from the checked-in, public-safe
[`offline-fixtures-pr252-four-gap-guards-20261007.json`](docs/benchmark-evidence/offline-fixtures-pr252-four-gap-guards-20261007.json) artifact. Its
SHA-256 is
`c53b82e4a847bf525c695c25a1813e8a591e86e39ca9e652e385e898ecd04010`, also recorded in the
adjacent `.sha256` file. The artifact contains no raw questions, answers, prompts, customer data,
or per-record content fingerprints.

The fixture-suite digest is
`8e2a562ba5d4c785c1e88481f9019446b43608109147cd7cd82597493004af35`. The artifact defines
The fixture-suite digest is
`e137eacdd70bbea7539dec6ed5cd4b59c980352a8e641094b449e474522aa109`. The artifact defines
the digest algorithm and records the SHA-256 of every suite and dataset file. Each evidence ID
also binds its exact command through `sha256(UTF-8 exact command)`:

Expand All @@ -121,12 +121,19 @@ means no number is claimed in this offline registry.
The context-efficiency chart is generated from the registry values and the selected report schema.
Historical LoCoMo, graph, handoff, consolidation, and security figures remain preserved in their
source artifacts but are omitted from the current chart until each has a matching immutable,
public-safe artifact. The chart labels coding outcomes, external datasets, and operational
capacity as pending evaluation tracks rather than implying scores. Regenerate it with
`python scripts/render_benchmark_report.py --report docs/benchmark-evidence/offline-fixtures-v129.json --output docs/images/context-efficiency.svg` after selecting the report to publish.
public-safe artifact. The chart reports context size, candidate and packed retrieval quality, and
grounded checks in separate panels; provider billing and MCP transport are not measured. Regenerate
the SVG and matching PNG with:

```bash
python scripts/render_benchmark_report.py --report docs/benchmark-evidence/offline-fixtures-pr252-four-gap-guards-20261007.json --output docs/images/context-efficiency.svg --png-output docs/images/context-efficiency.png
```

The companion examples are also generated from that artifact with
`python -m scripts.render_benchmark_examples --report docs/benchmark-evidence/offline-fixtures-v129.json --output docs/images/evidence-backed-agent-examples.svg`.
The companion examples are also generated from that artifact with:

```bash
python -m scripts.render_benchmark_examples --report docs/benchmark-evidence/offline-fixtures-pr252-four-gap-guards-20261007.json --output docs/images/evidence-backed-agent-examples.svg --png-output docs/images/evidence-backed-agent-examples.png
```
The historical-to-executable mapping is in
[`docs/BENCHMARK_CHANGE_COVERAGE.md`](docs/BENCHMARK_CHANGE_COVERAGE.md).

Expand Down Expand Up @@ -290,6 +297,27 @@ LoCoMo and LongMemEval retrieval diagnostics are retained as separate public-saf
and checksum boundaries. Those values remain evidence-retrieval metrics, not end-to-end QA accuracy
or an official LoCoMo leaderboard score.

## Jev-assisted route-selection probe (provider-backed, exploratory)

On 2026-10-03, the configured BYOK client ran `eval/jev_recall_quality.py` against 40 generated
public synthetic tasks (10 each for lexical identifiers, graph relationships, temporal changes,
and procedural context). Each call sent only the task query and at most two locally generated
routes; no recalled memory bodies or stored user data were sent. Jev selected a route on all 40
calls, but the baseline and Jev-assisted means were identical: nDCG@5 **0.740773**, Recall@5
**0.75**, and answer-token coverage **0.75**. Every paired 95% interval was zero, so the
improvement/non-inferiority gate did not pass. This synthetic probe does not establish a benefit
on held-out user workloads, and no retrieval-quality improvement is claimed.

The harness caps a run at 40 provider requests, makes at most one request per task, and requires
explicit `--allow-remote` consent. Repeating it can consume BYOK provider usage:

```bash
python -m eval.jev_recall_quality --allow-remote --max-requests 40 --timeout 8
```

This exploratory provider run is separate from the registered offline artifact and the benchmark
graphic below; it is not a general benchmark or customer-workload evaluation.

## What we do NOT yet claim

- **No official end-to-end LLM QA accuracy.** The deterministic productivity agent measures the
Expand Down
Loading
Loading