From 4a8c2a70a1f204d2636025aa5150f779ed33c0d6 Mon Sep 17 00:00:00 2001 From: Raymond Nicholas Date: Tue, 15 Sep 2026 22:58:33 +0100 Subject: [PATCH 01/12] docs: point the docs at CLAUDE.md, the name the file actually has CODEX.md was renamed to CLAUDE.md in 38a0b36 so Claude Code would auto-load it, but the H1 and all 15 cross-references still named the old file. A reader following any of them looked for a file that is not there. The engine-side Go comments already said CLAUDE.md. Co-Authored-By: Claude Code --- CLAUDE.md | 2 +- docs/development/CURRENT-STATE.md | 6 +++--- docs/development/NEXT.md | 8 ++++---- docs/development/PROGRESS.md | 16 ++++++++-------- 4 files changed, 16 insertions(+), 16 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index efe88a7..3519948 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -1,4 +1,4 @@ -# CODEX.md — working conventions for this repo +# CLAUDE.md — working conventions for this repo Read this fully before writing any code. Then read `docs/development/CURRENT-STATE.md` (what exists and why) and `docs/development/NEXT.md` (the ordered queue, diff --git a/docs/development/CURRENT-STATE.md b/docs/development/CURRENT-STATE.md index 18acd56..fb84a9f 100644 --- a/docs/development/CURRENT-STATE.md +++ b/docs/development/CURRENT-STATE.md @@ -38,7 +38,7 @@ still in-memory and single-process either way, the same caveat cryden's own default limiter carries. Response envelope, error codes, and the migration-copying convention -are all established — see `README.md` and `CODEX.md`. +are all established — see `README.md` and `CLAUDE.md`. ## Tier 0 — bump the engine to v2.5.0: DONE @@ -64,7 +64,7 @@ anything else in this repo. ## Tier 0.5 — admin/operator authorization foundation: DONE Console operator status is a concept this repo owns entirely, not -cryden — see `CODEX.md`'s ownership section for why. +cryden — see `CLAUDE.md`'s ownership section for why. - `migrations/003_operators.up.sql` / `.down.sql` — a new `operators` table, `user_id` (references cryden's own `users.id`), `role` (plain @@ -98,7 +98,7 @@ whichever Tier 4/5 endpoint lands first. ## Tier 1 — auth methods: DONE -Built on `feat/tier1-auth-methods` (its own branch, per `CODEX.md`'s +Built on `feat/tier1-auth-methods` (its own branch, per `CLAUDE.md`'s one-branch-per-tier rule), in this order: - `migrations/004`-`008` — cryden's `0003`-`0007` copied in, renumbered diff --git a/docs/development/NEXT.md b/docs/development/NEXT.md index 5a28e13..396d0b2 100644 --- a/docs/development/NEXT.md +++ b/docs/development/NEXT.md @@ -1,12 +1,12 @@ # api — next up Ordered queue. Take the first unfinished item, build it completely, -verify it (see `CODEX.md`), update the three docs +verify it (see `CLAUDE.md`), update the three docs (`CURRENT-STATE.md`/`NEXT.md`/`PROGRESS.md`), then stop for review before starting the next tier. Specs below are deliberately detailed so you don't need to ask anything mid-build — where something is genuinely unspecified, make the most reasonable call consistent with -`CODEX.md`'s ownership rules and note the assumption in `PROGRESS.md`. +`CLAUDE.md`'s ownership rules and note the assumption in `PROGRESS.md`. Tier 0 and Tier 0.5 are done — see `CURRENT-STATE.md`. Tier 1 is done — see the status note under Tier 1 and `PROGRESS.md`'s @@ -29,7 +29,7 @@ the status note under Tier 3 and `PROGRESS.md`'s 2026-09-15 entries. > What is still owed: a first DB-backed smoke-test run (no Postgres in > this sandbox), a live Apple round trip (no Apple credentials here), and > the WebAuthn ceremonies, which need a real browser authenticator. -> `PROGRESS.md` says all of that plainly, per `CODEX.md`'s verification +> `PROGRESS.md` says all of that plainly, per `CLAUDE.md`'s verification > rule, rather than counting green unit tests as end-to-end coverage. Each of these mirrors an existing engine feature that already has a @@ -285,7 +285,7 @@ Two details were decided rather than assumed, and are recorded in ## Tier 4 — AI-assisted admin endpoints (all behind `RequireAdmin`) Every endpoint in this tier stays read-only/surface-only, no -exceptions — see `CODEX.md`'s hard rule at the top. +exceptions — see `CLAUDE.md`'s hard rule at the top. - **Weekly digest**: `GET /v1/admin/digest` → `cryden.WeeklyDigest`/ `DigestSince`. Plus **scheduling and history** (new, this repo's own diff --git a/docs/development/PROGRESS.md b/docs/development/PROGRESS.md index 90a545d..edaf397 100644 --- a/docs/development/PROGRESS.md +++ b/docs/development/PROGRESS.md @@ -35,14 +35,14 @@ Assumptions made, none blocking: operator all get the identical `403 not_operator` — that distinction is not something to expose to the caller. -Next: Tier 1 (auth methods), each on its own branch per `CODEX.md`. +Next: Tier 1 (auth methods), each on its own branch per `CLAUDE.md`. Copying cryden's migrations `0003`-`0007` into this repo (renumbered continuing from `003_operators`) is the first sub-step, before any TOTP/WebAuthn/magic-link/recovery-code endpoint work starts. ## 2026-09-14 — Tier 1 (auth methods) except Apple -Branch `feat/tier1-auth-methods`, per `CODEX.md`'s one-branch-per-tier +Branch `feat/tier1-auth-methods`, per `CLAUDE.md`'s one-branch-per-tier rule. First session in this repo with a working Go toolchain: Go 1.25.0 plus cryden v2.5.0 and every dependency already in the module cache, so the caveat the Tier 0 entry left open is closed — `go mod tidy` @@ -52,7 +52,7 @@ was verified: **the DB-backed smoke test was not run** (no Postgres and no network in this sandbox) and neither were the WebAuthn ceremonies, which need a real browser authenticator. Those still owe a first run against a real database. Saying that plainly here rather than counting -green builds as "verified end to end", per `CODEX.md`. +green builds as "verified end to end", per `CLAUDE.md`. Built, in commit order: @@ -195,12 +195,12 @@ network), the DB-backed smoke test (no Postgres), and the WebAuthn ceremonies (no browser authenticator). Those remain the first things to run on a real deployment. -Next: Tier 2, on its own branch per `CODEX.md` — and before or alongside +Next: Tier 2, on its own branch per `CLAUDE.md` — and before or alongside it, the first DB-backed smoke-test run of everything in Tier 1. ## 2026-09-15 — Tier 2 (config, named sessions, OAuth health) -Branch `feat/tier2-config-and-oauth-health`, per `CODEX.md`'s +Branch `feat/tier2-config-and-oauth-health`, per `CLAUDE.md`'s one-branch-per-tier rule. Three commits, in order: - `feat: wire anomaly detection and the Redis rate limiter from env` — @@ -319,7 +319,7 @@ rather than silently patched): smoketest does not have). An optional operator token/email flag would fix it if that coverage is wanted later. -Next: Tier 3, on its own branch per `CODEX.md`. Still owed from before +Next: Tier 3, on its own branch per `CLAUDE.md`. Still owed from before it: the first DB-backed smoke-test run, now worth doing against a `REDIS_URL`-less and a `REDIS_URL`-set instance so the shared limiter gets its first real exercise. @@ -382,7 +382,7 @@ trailing `// dev stand-in` comments, which align against the longest line in their group. Both fixed with `gofmt -w`. **What was checked before the toolchain was reachable** — since -`CODEX.md`'s rule is to say what was and was not done rather than to +`CLAUDE.md`'s rule is to say what was and was not done rather than to imply a build — every cryden symbol Stage 1 calls was read directly out of the module cache at `…/cryden/v2@v2.5.0`, first-hand, not recalled. Confirmed: @@ -626,7 +626,7 @@ so the field is `Errors`. ### Verification: what this does NOT cover -Said plainly, per `CODEX.md`, rather than implied by a green suite: +Said plainly, per `CLAUDE.md`, rather than implied by a green suite: - **There is no Postgres and no network in this sandbox.** `migrations/009`, `010` and `011` have **never been applied to a real From 21ac94c53efa6ff475fb707b7372f6f5749d6589 Mon Sep 17 00:00:00 2001 From: Raymond Nicholas Date: Tue, 15 Sep 2026 23:23:20 +0100 Subject: [PATCH 02/12] feat: add the weekly digest and its recorded history GET /v1/admin/digest builds cryden's report for a window ending now and records nothing, so asking twice leaves no trace. DIGEST_INTERVAL_HOURS turns on a scheduler that writes each report to this repo's own digest_runs table, which GET /v1/admin/digest/history reads back. Co-Authored-By: Claude Code --- .env.example | 16 + README.md | 19 ++ config/config.go | 24 ++ digest/memory.go | 88 ++++++ digest/schedule.go | 107 +++++++ digest/schedule_test.go | 283 ++++++++++++++++++ digest/store.go | 214 ++++++++++++++ digest/store_test.go | 248 ++++++++++++++++ httpapi/digest_handlers.go | 149 ++++++++++ httpapi/digest_handlers_test.go | 439 ++++++++++++++++++++++++++++ httpapi/router.go | 23 ++ main.go | 51 ++++ migrations/012_digest_runs.down.sql | 4 + migrations/012_digest_runs.up.sql | 50 ++++ openapi/spec.yaml | 147 +++++++++- 15 files changed, 1861 insertions(+), 1 deletion(-) create mode 100644 digest/memory.go create mode 100644 digest/schedule.go create mode 100644 digest/schedule_test.go create mode 100644 digest/store.go create mode 100644 digest/store_test.go create mode 100644 httpapi/digest_handlers.go create mode 100644 httpapi/digest_handlers_test.go create mode 100644 migrations/012_digest_runs.down.sql create mode 100644 migrations/012_digest_runs.up.sql diff --git a/.env.example b/.env.example index 1f4ec72..b65c5c1 100644 --- a/.env.example +++ b/.env.example @@ -154,3 +154,19 @@ WEBHOOK_URL= WEBHOOK_SECRET= WEBHOOK_EVENTS= WEBHOOK_MAX_ATTEMPTS= + +# The weekly digest schedule. Unset (or 0) means no schedule at all: no +# goroutine runs, nothing is written, and GET /v1/admin/digest/history +# answers 404 not_configured. GET /v1/admin/digest works either way — it +# builds a digest on demand and records nothing. +# +# Set it and this repo builds the engine's digest every N hours and stores +# the rendered report in the digest_runs table, which is what the history +# endpoint lists. 168 is weekly. cryden has no scheduling concept, so this +# table, this job and that endpoint are all this repo's own. +# +# The first run happens one full interval after startup, not at boot: a +# process that restarts more often than the interval elapses would +# otherwise write one row per restart, and a history that grows with +# restarts is not a history of anything. +DIGEST_INTERVAL_HOURS= diff --git a/README.md b/README.md index 17de229..bd051fa 100644 --- a/README.md +++ b/README.md @@ -315,6 +315,25 @@ There is no vendor here: this repo ships no SDK, so "shipped" means "recorded in - An unknown `level` is a `400` naming the four valid values, not an empty list — which is indistinguishable from "the engine has been quiet". - The write is **synchronous**, on the goroutine that logged. That is a real cost and is not the shape a busy deployment wants; it is the shape this one can have, because an asynchronous sink needs a flush policy and a shutdown path, and this repo has no graceful shutdown anywhere yet. A buffer that is never flushed on exit is a log that silently drops its last records before a crash, which for a log is the failure that matters most. `LOG_LEVEL` (default `info`) is what keeps the volume sane in the meantime, since the engine's debug records never reach the sink. +## Weekly digest + +Two endpoints, and only one of them depends on any configuration: + +``` +GET /v1/admin/digest?window_days= # built now, records nothing +GET /v1/admin/digest/history?limit= # what the schedule recorded +``` + +`GET /v1/admin/digest` returns `cryden.DigestSince`'s report verbatim — the text is the engine's, and this repo does not reformat a report it does not own — alongside `since` and `until`, because a client should not have to parse English out of a digest to learn what it covers. `window_days` (1–365, default 7) is passed to the engine rather than implemented here; the seven-day default is what makes it a *weekly* digest. **Asking twice leaves no trace**: the endpoint records nothing, and if that ever stopped being true an operator could no longer tell what the schedule produced from what somebody happened to open. + +Setting `DIGEST_INTERVAL_HOURS` (168 is weekly) turns on the schedule: a background job calls the same report every N hours and writes the rendered result to the `digest_runs` table, which `GET /v1/admin/digest/history` lists newest first. Unset means no schedule — no goroutine runs, nothing is written, and the history endpoint answers `404 not_configured` rather than an empty list an operator would read as "nothing has ever happened". + +- **cryden has no scheduling concept.** `WeeklyDigest`/`DigestSince` build a report on demand and return a string; there is no run record and nothing that remembers a digest was ever generated. So the table, the job and the history endpoint are entirely this repo's own. +- **The row is the report, not a recipe for one.** The rendered text is stored rather than the counts behind it, because a digest covers a window that has *ended*: re-running its query later would not reproduce it, since "the last seven days" is anchored to when it was built. +- **The first run is one full interval after startup**, not at boot. A process that restarts more often than the interval elapses — a crashloop, a deploy pipeline, a laptop — would otherwise write one row per restart, and a history that grows with restarts rather than with time is not a history of anything. +- A failed run is **logged and swallowed**. This runs in a goroutine with nobody to hand an error to, and a scheduler that stopped at the first database blip would silently stop producing digests for the rest of the process's life. +- Nothing on the HTTP surface can create a digest run. Only the scheduler writes, and it is a process component rather than a request handler — the read-only rule the whole admin surface follows. + ## Design notes - `CORS_ORIGINS` is required, no wildcard default — an API handling auth tokens should never allow every origin. diff --git a/config/config.go b/config/config.go index c63f35b..debd264 100644 --- a/config/config.go +++ b/config/config.go @@ -233,6 +233,17 @@ type Config struct { // that retries forever is a load generator pointed at a third party — // and the row stays readable afterwards either way. WebhookMaxAttempts int + + // DigestInterval is how often the background job builds a digest and + // records it in digest_runs, which is what + // GET /v1/admin/digest/history reads back. Zero — the default — runs no + // job at all. + // + // Opt-in rather than "weekly by default", for the reason webhooks and + // cloud logging are: a deployment that has not asked for scheduled + // digests should not have a goroutine quietly accumulating rows. The + // on-demand GET /v1/admin/digest works either way, and writes nothing. + DigestInterval time.Duration } // PasswordHasher values. Bcrypt is the engine's own default and what an @@ -525,6 +536,19 @@ func Load() (Config, error) { } } + // Scheduled digests. Unset or 0 is off (see the field comment); a + // negative is refused rather than read as "off", because it can only + // be a typo and silently treating a typo as the default is how a + // setting an operator meant to change does nothing at all. + digestHours, err := envInt("DIGEST_INTERVAL_HOURS", 0) + if err != nil { + return cfg, err + } + if digestHours < 0 { + return cfg, fmt.Errorf("DIGEST_INTERVAL_HOURS cannot be negative — leave it unset to switch scheduled digests off, got %d", digestHours) + } + cfg.DigestInterval = time.Duration(digestHours) * time.Hour + return cfg, nil } diff --git a/digest/memory.go b/digest/memory.go new file mode 100644 index 0000000..c0f7aba --- /dev/null +++ b/digest/memory.go @@ -0,0 +1,88 @@ +package digest + +import ( + "context" + "sort" + "sync" + "time" +) + +// MemoryStore is the in-process Store, for tests and for any embedding +// host that wants the digest history without a database behind it. +// +// It is a faithful double rather than a convenient one where the two +// implementations could quietly disagree: +// +// - List is ordered by GeneratedAt descending with ID as the tiebreak, +// which is the SQL's ORDER BY generated_at DESC, id DESC. A double +// that sorted only on the timestamp would pass every test while the +// real query returned a stable order the double did not have. +// - Insert applies the same zero-value stamping the Postgres store +// does, through the same resolveTimes, so a test asserting "the store +// filled in the clock" is asserting about real behaviour. +// +// What it does not reproduce is the database: there is no TIMESTAMPTZ +// round trip here, so a time that would not survive one is a difference +// this double cannot show. Postgres stores microseconds; Go's time.Time +// carries nanoseconds, and a monotonic reading is dropped on the way in. +// Nothing in flight depends on either — the window columns are compared +// against each other, never against a stored digest — but the gap is +// worth naming rather than assuming away. +type MemoryStore struct { + mu sync.Mutex + runs []Entry + next int64 + + // Clock stamps an entry that does not carry its own GeneratedAt, so a + // test can make a listing's ordering deterministic instead of hoping + // the wall clock separated two inserts. + Clock func() time.Time +} + +func NewMemoryStore() *MemoryStore { + return &MemoryStore{Clock: func() time.Time { return time.Now().UTC() }} +} + +var _ Store = (*MemoryStore)(nil) + +func (s *MemoryStore) Insert(_ context.Context, e Entry) (Entry, error) { + s.mu.Lock() + defer s.mu.Unlock() + + s.next++ + e.ID = s.next + e = resolveTimes(e, s.Clock()) + s.runs = append(s.runs, e) + return e, nil +} + +func (s *MemoryStore) List(_ context.Context, limit int) ([]Entry, error) { + s.mu.Lock() + defer s.mu.Unlock() + + out := make([]Entry, len(s.runs)) + copy(out, s.runs) + // Newest first, ties broken by id descending — the SQL's + // ORDER BY generated_at DESC, id DESC. + sort.SliceStable(out, func(i, j int) bool { + if !out[i].GeneratedAt.Equal(out[j].GeneratedAt) { + return out[i].GeneratedAt.After(out[j].GeneratedAt) + } + return out[i].ID > out[j].ID + }) + // Clamped here as well as by the handler, so a caller reaching the + // store directly gets the same bounded answer the endpoint gives. + limit = ClampLimit(limit) + if len(out) > limit { + out = out[:limit] + } + return out, nil +} + +// Count returns how many runs have been recorded. A test helper: no +// production caller needs a total, and no endpoint reports one. +func (s *MemoryStore) Count() int { + s.mu.Lock() + defer s.mu.Unlock() + return len(s.runs) +} diff --git a/digest/schedule.go b/digest/schedule.go new file mode 100644 index 0000000..473744a --- /dev/null +++ b/digest/schedule.go @@ -0,0 +1,107 @@ +package digest + +import ( + "context" + "log" + "time" +) + +// Builder produces one digest for the scheduler to record. It returns an +// Entry with ID unset — the store assigns one, and the store's own +// zero-value rule fills in whatever timestamp the builder leaves alone. +// +// It is a function rather than a store interface because the thing being +// wrapped is cryden's DigestSince, which takes the engine and returns a +// string. Threading a whole engine through this package to call it would +// mean this package importing the engine to describe a seam the caller +// can close in three lines. +type Builder func(ctx context.Context) (Entry, error) + +// Scheduler builds a digest on an interval and records each one. +// +// It is the only writer in this package, and it is a process component +// rather than anything a request can reach: no endpoint in this repo +// creates a digest run, so an operator cannot manufacture history +// through the API. That is the same shape the admin surface keeps +// everywhere else — see CLAUDE.md's hard rule. +type Scheduler struct { + Store Store + Build Builder + + // Interval is how long to wait between runs. Zero or negative means + // there is no schedule, and Run returns immediately without starting + // anything — main.go only constructs a Scheduler when + // DIGEST_INTERVAL_HOURS asked for one, so a zero here is a wiring + // mistake rather than a setting. + Interval time.Duration + + // Log receives one line per failed run. Optional; a nil Log discards + // them. + // + // Failures are logged and swallowed rather than returned: this runs in + // its own goroutine with nobody to hand an error to, and a scheduler + // that stopped on the first database blip would silently stop + // producing digests for the rest of the process's life — the exact + // failure a schedule exists to avoid. + Log *log.Logger +} + +// Run blocks until ctx is done, building and recording a digest once per +// Interval. +// +// The first run happens after a full Interval, not at startup. That is +// deliberate: the interval is the schedule, and a process that restarts +// more often than the interval elapses — a crashloop, a deploy pipeline, +// a developer's laptop — would otherwise manufacture one digest row per +// start. A history that grows with restarts rather than with time is not +// a history of anything. +// +// main.go hands this context.Background(), because this repo has no +// graceful shutdown yet — the same caveat, and the same reasoning, as the +// webhook worker's goroutine. Nothing here needs stopping today: an +// interrupted run loses at most one digest, and the next interval builds +// another. +func (s *Scheduler) Run(ctx context.Context) { + if s.Interval <= 0 || s.Store == nil || s.Build == nil { + return + } + + ticker := time.NewTicker(s.Interval) + defer ticker.Stop() + + for { + select { + case <-ctx.Done(): + return + case <-ticker.C: + s.runOnce(ctx) + } + } +} + +// runOnce builds one digest and records it, logging rather than returning +// any failure — see the Log field for why the loop must survive one. +func (s *Scheduler) runOnce(ctx context.Context) { + entry, err := s.Build(ctx) + if err != nil { + s.logf("digest: building the scheduled digest failed: %v", err) + return + } + + // The zero-value rule stamps GeneratedAt and WindowEnd if the builder + // left them; a builder that set neither still produces a readable row. + saved, err := s.Store.Insert(ctx, entry) + if err != nil { + s.logf("digest: recording the scheduled digest failed: %v", err) + return + } + s.logf("digest: recorded a scheduled digest covering %s to %s (run %d)", + saved.WindowStart.UTC().Format(time.RFC3339), saved.WindowEnd.UTC().Format(time.RFC3339), saved.ID) +} + +func (s *Scheduler) logf(format string, args ...any) { + if s.Log == nil { + return + } + s.Log.Printf(format, args...) +} diff --git a/digest/schedule_test.go b/digest/schedule_test.go new file mode 100644 index 0000000..de30138 --- /dev/null +++ b/digest/schedule_test.go @@ -0,0 +1,283 @@ +package digest + +import ( + "bytes" + "context" + "errors" + "log" + "strings" + "sync" + "testing" + "time" +) + +// syncBuffer collects log output written from a goroutine the test does not +// control. A bare bytes.Buffer would be a race the -race build is entitled +// to fail on, and the whole point of these tests is a loop running beside +// the assertion. +type syncBuffer struct { + mu sync.Mutex + buf bytes.Buffer +} + +func (b *syncBuffer) Write(p []byte) (int, error) { + b.mu.Lock() + defer b.mu.Unlock() + return b.buf.Write(p) +} + +func (b *syncBuffer) String() string { + b.mu.Lock() + defer b.mu.Unlock() + return b.buf.String() +} + +// builder is a Builder that counts its calls and hands back a fixed entry, +// so a test can ask whether it was called at all rather than infer it from +// a row. +type builder struct { + mu sync.Mutex + calls int + entry Entry + err error +} + +func (b *builder) build(context.Context) (Entry, error) { + b.mu.Lock() + defer b.mu.Unlock() + b.calls++ + return b.entry, b.err +} + +func (b *builder) callCount() int { + b.mu.Lock() + defer b.mu.Unlock() + return b.calls +} + +// A scheduler with nothing to schedule does nothing at all: no goroutine +// left ticking, no row written. This is the state main.go avoids by only +// constructing a Scheduler when DIGEST_INTERVAL_HOURS asked for one, so +// what is asserted here is that a mistake there stays inert. +func TestSchedulerDoesNothingWithoutASchedule(t *testing.T) { + for _, tc := range []struct { + name string + interval time.Duration + withStore bool + withBuild bool + }{ + {name: "no interval", withStore: true, withBuild: true}, + {name: "no interval and nothing else either"}, + {name: "no store", interval: time.Hour, withBuild: true}, + {name: "no builder", interval: time.Hour, withStore: true}, + } { + t.Run(tc.name, func(t *testing.T) { + store := NewMemoryStore() + build := &builder{entry: Entry{Text: "unreachable"}} + + s := &Scheduler{Interval: tc.interval} + if tc.withStore { + s.Store = store + } + if tc.withBuild { + s.Build = build.build + } + + // Run is expected back promptly rather than at the end of an + // interval: an unconfigured scheduler must not hold a goroutine + // open for an hour first. + done := make(chan struct{}) + go func() { + defer close(done) + s.Run(context.Background()) + }() + select { + case <-done: + case <-time.After(2 * time.Second): + t.Fatal("Run is still going with nothing to run") + } + + if n := store.Count(); n != 0 { + t.Errorf("recorded %d runs, want 0", n) + } + if n := build.callCount(); n != 0 { + t.Errorf("the builder was called %d times, want 0", n) + } + }) + } +} + +// The first run is after a full interval, not at startup. Asserted against +// an hour-long interval, so this cannot pass by being slow: if the run +// happened at startup it would have happened within microseconds of Run +// being called, and the check below waits a hundred milliseconds. +// +// It is the property that keeps a crashlooping process from manufacturing +// one row per restart, which is the failure mode a history table is worst +// at showing — the rows look like a busy week. +func TestSchedulerWaitsAFullIntervalBeforeTheFirstRun(t *testing.T) { + store := NewMemoryStore() + build := &builder{entry: Entry{Text: "the first run"}} + s := &Scheduler{Store: store, Build: build.build, Interval: time.Hour} + + ctx, cancel := context.WithCancel(context.Background()) + done := make(chan struct{}) + go func() { + defer close(done) + s.Run(ctx) + }() + + time.Sleep(100 * time.Millisecond) + if n := build.callCount(); n != 0 { + t.Errorf("the builder ran %d times before the first interval elapsed, want 0", n) + } + if n := store.Count(); n != 0 { + t.Errorf("recorded %d runs before the first interval elapsed, want 0", n) + } + + // And it stops when the context is cancelled rather than at the next + // tick, which is what a caller with a shutdown path would need. + cancel() + select { + case <-done: + case <-time.After(2 * time.Second): + t.Fatal("Run did not return after its context was cancelled") + } +} + +// The positive half: one run per interval, each stored with the window the +// builder computed. The interval is deliberately tiny so the test is +// seconds-cheap; what is being asserted is that the loop records, not how +// often it does. +func TestSchedulerRecordsOneRunPerInterval(t *testing.T) { + store := NewMemoryStore() + since := time.Date(2026, 8, 1, 0, 0, 0, 0, time.UTC) + build := &builder{entry: Entry{WindowStart: since, Text: "the weekly report"}} + s := &Scheduler{Store: store, Build: build.build, Interval: 5 * time.Millisecond} + + ctx, cancel := context.WithCancel(context.Background()) + defer cancel() + go s.Run(ctx) + + waitForRuns(t, store, 2) + + rows, err := store.List(context.Background(), 10) + if err != nil { + t.Fatalf("List: %v", err) + } + if rows[0].Text != "the weekly report" { + t.Errorf("text = %q, want the builder's report stored verbatim", rows[0].Text) + } + if !rows[0].WindowStart.Equal(since) { + t.Errorf("window_start = %v, want the window the builder computed %v", rows[0].WindowStart, since) + } + if rows[0].ID == 0 { + t.Error("a recorded run has no id") + } + if rows[0].GeneratedAt.IsZero() || rows[0].WindowEnd.IsZero() { + t.Errorf("run %+v was stored without its times filled in", rows[0]) + } +} + +// A failed build is logged and the loop keeps going. A scheduler that +// stopped at the first database blip would silently stop producing +// digests for the rest of the process's life — the exact failure a +// schedule exists to avoid — and one that retried instantly would spin. +func TestSchedulerSurvivesAFailedBuild(t *testing.T) { + store := NewMemoryStore() + logs := &syncBuffer{} + + var mu sync.Mutex + attempts := 0 + build := func(context.Context) (Entry, error) { + mu.Lock() + defer mu.Unlock() + attempts++ + if attempts == 1 { + return Entry{}, errors.New("the audit table is unreachable") + } + return Entry{Text: "the second attempt"}, nil + } + + s := &Scheduler{Store: store, Build: build, Interval: 5 * time.Millisecond, Log: log.New(logs, "", 0)} + ctx, cancel := context.WithCancel(context.Background()) + defer cancel() + go s.Run(ctx) + + waitForRuns(t, store, 1) + + if got := logs.String(); !strings.Contains(got, "the audit table is unreachable") { + t.Errorf("log = %q, want the build failure recorded", got) + } + rows, err := store.List(context.Background(), 10) + if err != nil { + t.Fatalf("List: %v", err) + } + if rows[0].Text != "the second attempt" { + t.Errorf("text = %q, want the run that succeeded after the failure", rows[0].Text) + } +} + +// A builder that fails every time writes nothing rather than a row of +// empty text: an entry with no report in it is worse than no entry, because +// a history listing cannot tell it apart from a quiet week. +func TestSchedulerRecordsNothingWhenTheStoreRejectsTheRun(t *testing.T) { + store := &refusingStore{} + s := &Scheduler{Store: store, Build: func(context.Context) (Entry, error) { + return Entry{Text: "never stored"}, nil + }, Interval: 5 * time.Millisecond, Log: log.New(&syncBuffer{}, "", 0)} + + ctx, cancel := context.WithCancel(context.Background()) + defer cancel() + go s.Run(ctx) + + deadline := time.Now().Add(2 * time.Second) + for store.insertAttempts() < 3 && time.Now().Before(deadline) { + time.Sleep(time.Millisecond) + } + if n := store.insertAttempts(); n < 3 { + t.Fatalf("the scheduler made %d insert attempts in two seconds, want the loop to keep trying", n) + } + if n := store.count(); n != 0 { + t.Errorf("%d runs were recorded by a store that refused every insert", n) + } +} + +// waitForRuns blocks until the store holds at least n runs, failing the +// test rather than hanging if the loop never gets there. +func waitForRuns(t *testing.T, store *MemoryStore, n int) { + t.Helper() + deadline := time.Now().Add(5 * time.Second) + for store.Count() < n { + if time.Now().After(deadline) { + t.Fatalf("the scheduler recorded %d runs in five seconds, want %d", store.Count(), n) + } + time.Sleep(time.Millisecond) + } +} + +// refusingStore counts insert attempts and rejects all of them, which is +// what a database that is down looks like from the scheduler's side. +type refusingStore struct { + mu sync.Mutex + attempts int +} + +func (s *refusingStore) Insert(context.Context, Entry) (Entry, error) { + s.mu.Lock() + defer s.mu.Unlock() + s.attempts++ + return Entry{}, errors.New("the database is unreachable") +} + +func (s *refusingStore) List(context.Context, int) ([]Entry, error) { return nil, nil } + +func (s *refusingStore) insertAttempts() int { + s.mu.Lock() + defer s.mu.Unlock() + return s.attempts +} + +func (s *refusingStore) count() int { return 0 } + +var _ Store = (*refusingStore)(nil) diff --git a/digest/store.go b/digest/store.go new file mode 100644 index 0000000..7185621 --- /dev/null +++ b/digest/store.go @@ -0,0 +1,214 @@ +// Package digest keeps the history of the reports this deployment's +// weekly digest produced, and runs the schedule that produces them. +// +// # Why this is here and not in the engine +// +// cryden's WeeklyDigest/DigestSince build a report out of the audit table +// and hand back a string. There is no scheduler in the engine, no run +// record, and nothing that remembers a digest was ever generated — that +// is deliberate on cryden's side, because "when should this fire" and +// "where should it be kept" are deployment questions, not authentication +// ones. So the history table, the background job that fills it, and the +// endpoint that reads it back are all this repo's own. +// +// # Read-only, like everything else on the admin surface +// +// GET /v1/admin/digest and GET /v1/admin/digest/history only ever read, +// and neither of them writes a row: the on-demand endpoint deliberately +// does NOT record what it built, so asking for a digest twice does not +// fabricate two entries in a history an operator reads as a record of +// what was scheduled. Only Scheduler writes, and it is a process +// component rather than a request handler — see CLAUDE.md's hard rule. +// +// # The row is the report, not a recipe for one +// +// Text is stored rather than the counts behind it. A digest covers a +// window that has ended, so re-running its query later would not +// reproduce it: "the last seven days" is anchored to when the digest was +// built. Storing what was actually reported is what makes reading a past +// digest the same experience as reading a fresh one. +package digest + +import ( + "context" + "database/sql" + "errors" + "fmt" + "time" +) + +// Entry is one recorded digest. +type Entry struct { + // ID is assigned by the store and is zero on the value handed to + // Insert. + ID int64 + + // WindowStart and WindowEnd bound the report: everything counted + // happened at or after WindowStart, and WindowEnd is the instant the + // digest was built. Carried as fields rather than parsed back out of + // Text, so a listing can sort and describe runs without reading + // English out of a report. + WindowStart time.Time + WindowEnd time.Time + + // GeneratedAt is when the row was written. Distinct from WindowEnd on + // purpose: a run replayed after an outage covers a window that closed + // before the digest was made. + GeneratedAt time.Time + + // Text is the rendered report, exactly as the engine returned it. + // Stored verbatim and never re-rendered — this repo does not format + // cryden's reports, and a copy reformatted here would be a second + // implementation of a report the engine already owns. + Text string +} + +// Store is the persistence seam. Two implementations: PostgresStore and +// MemoryStore, the in-memory double the tests use — the same split every +// repo-owned store in this repo follows. +type Store interface { + // Insert records one run and returns it with its assigned ID. + // + // Unlike webhook deliveries there is no dedupe and no "already + // recorded" case: two runs of the same window are two things that + // happened, and the second one is exactly what an operator wants to + // see when they suspect the schedule fired twice. + Insert(ctx context.Context, e Entry) (Entry, error) + + // List returns runs newest first, at most limit of them. Ordered by + // GeneratedAt with ID as the tiebreak, so the order is total and two + // runs sharing a timestamp do not swap places between requests. + List(ctx context.Context, limit int) ([]Entry, error) +} + +// ErrNotFound is returned when a run is asked for by an ID that is not +// there. Nothing in this repo reads by ID today — the history is listed, +// never fetched — so this exists for a caller that grows one, and for the +// in-memory double to mean the same thing the Postgres store means. +var ErrNotFound = errors.New("digest run not found") + +// DefaultHistoryLimit bounds a history listing that does not ask for a +// size, and MaxHistoryLimit is the ceiling a request may ask for. A +// digest is at most one row per interval — 52 a year on the weekly +// default — so this is a bound against a caller looping with a large +// limit, not against ordinary volume. +const ( + DefaultHistoryLimit = 20 + MaxHistoryLimit = 200 +) + +// Columns is one const so the scan and the query cannot drift apart — the +// same reason shiplog's logColumns is one. +const runColumns = `id, window_start, window_end, generated_at, digest_text` + +// resolveTimes applies the zero-value rule both stores share: an entry +// that does not carry its own timestamps is stamped by whichever clock +// the store owns, so Postgres stamps with the database host's clock and +// the in-memory double stamps with the test's. +// +// WindowEnd falls back to GeneratedAt rather than to the fallback clock +// directly: the two are the same instant for every run this package +// writes, and deriving one from the other keeps them equal even for a +// caller that supplied only a GeneratedAt. +func resolveTimes(e Entry, fallback time.Time) Entry { + if e.GeneratedAt.IsZero() { + e.GeneratedAt = fallback + } + if e.WindowEnd.IsZero() { + e.WindowEnd = e.GeneratedAt + } + return e +} + +// PostgresStore is the durable Store. +type PostgresStore struct { + db *sql.DB +} + +func NewStore(db *sql.DB) *PostgresStore { + return &PostgresStore{db: db} +} + +var _ Store = (*PostgresStore)(nil) + +func (s *PostgresStore) Insert(ctx context.Context, e Entry) (Entry, error) { + e = resolveTimes(e, time.Now().UTC()) + + err := s.db.QueryRowContext(ctx, + `INSERT INTO digest_runs (window_start, window_end, generated_at, digest_text) + VALUES ($1, $2, $3, $4) + RETURNING id`, + e.WindowStart, e.WindowEnd, e.GeneratedAt, e.Text, + ).Scan(&e.ID) + if err != nil { + return Entry{}, err + } + return e, nil +} + +func (s *PostgresStore) List(ctx context.Context, limit int) ([]Entry, error) { + // Clamped inside the store rather than only at the handler, because + // `LIMIT $1` with a zero returns nothing and with a negative is a + // Postgres error — so an unclamped store answers a caller that + // reached it directly with either a broken listing or a driver fault, + // while the in-memory double would have answered sensibly. Applying + // the same rule in both is what keeps them the same store. + limit = ClampLimit(limit) + + rows, err := s.db.QueryContext(ctx, + `SELECT `+runColumns+` + FROM digest_runs + ORDER BY generated_at DESC, id DESC + LIMIT $1`, + limit, + ) + if err != nil { + return nil, err + } + defer rows.Close() + + runs := make([]Entry, 0, limit) + for rows.Next() { + e, err := scanRun(rows) + if err != nil { + return nil, err + } + runs = append(runs, e) + } + return runs, rows.Err() +} + +// scanRun reads one row. Kept as a function so the column list and the +// scan order are read together. +func scanRun(rows *sql.Rows) (Entry, error) { + var e Entry + if err := rows.Scan(&e.ID, &e.WindowStart, &e.WindowEnd, &e.GeneratedAt, &e.Text); err != nil { + return Entry{}, fmt.Errorf("scanning a digest run: %w", err) + } + return e, nil +} + +// ClampLimit narrows a requested history limit into range. Shared by both +// implementations, so the bound is applied in one place rather than +// restated at each of the two. +// +// A non-positive limit means "the default" rather than "none", because +// `LIMIT 0` is a listing that looks broken while a negative limit is a +// driver error — so without this, a caller reaching the store directly +// would get an answer the endpoint would never give. +// +// The endpoint itself does not rely on the clamp: it bounds the parameter +// and answers 400 outside 1..MaxHistoryLimit, so a console asking for +// limit=0 is told its request was wrong rather than handed a default it +// did not ask for (see httpapi's queryInt and the same rule in the logging +// endpoint). This is what a direct caller gets. +func ClampLimit(limit int) int { + switch { + case limit <= 0: + return DefaultHistoryLimit + case limit > MaxHistoryLimit: + return MaxHistoryLimit + default: + return limit + } +} diff --git a/digest/store_test.go b/digest/store_test.go new file mode 100644 index 0000000..64d2eb8 --- /dev/null +++ b/digest/store_test.go @@ -0,0 +1,248 @@ +package digest + +import ( + "context" + "fmt" + "sync" + "testing" + "time" +) + +// testClock is a hand-wound clock, so an ordering assertion is about the +// store's sort rather than about how far apart two time.Now() calls +// happened to land. The same idea as webhook's, in this package's own +// terms because a double is only worth having if it is the double the +// store under test actually reads. +type testClock struct { + mu sync.Mutex + at time.Time +} + +func newTestClock() *testClock { + // A fixed instant, not now: nothing here should depend on when the + // suite runs. + return &testClock{at: time.Date(2026, 8, 1, 12, 0, 0, 0, time.UTC)} +} + +func (c *testClock) now() time.Time { + c.mu.Lock() + defer c.mu.Unlock() + return c.at +} + +func (c *testClock) advance(d time.Duration) { + c.mu.Lock() + defer c.mu.Unlock() + c.at = c.at.Add(d) +} + +// newTestStore is a MemoryStore on a clock the test controls. +func newTestStore() (*MemoryStore, *testClock) { + clock := newTestClock() + s := NewMemoryStore() + s.Clock = clock.now + return s, clock +} + +func insert(t *testing.T, s Store, e Entry) Entry { + t.Helper() + saved, err := s.Insert(context.Background(), e) + if err != nil { + t.Fatalf("Insert: %v", err) + } + return saved +} + +// The zero-value rule, in all three of its cases. The third is the one +// worth stating: WindowEnd falls back to the entry's own GeneratedAt +// rather than to the clock, so a caller that supplied a GeneratedAt gets +// a WindowEnd equal to it instead of one a few microseconds later. +func TestMemoryStoreStampsOnlyTheTimesItIsNotGiven(t *testing.T) { + ctx := context.Background() + s, clock := newTestStore() + + bare := insert(t, s, Entry{Text: "nothing given"}) + if bare.ID != 1 { + t.Errorf("id = %d on the first insert, want 1", bare.ID) + } + if !bare.GeneratedAt.Equal(clock.now()) { + t.Errorf("generated_at = %v, want the store's clock %v", bare.GeneratedAt, clock.now()) + } + if !bare.WindowEnd.Equal(bare.GeneratedAt) { + t.Errorf("window_end = %v with nothing given, want the generated_at %v", bare.WindowEnd, bare.GeneratedAt) + } + + clock.advance(time.Hour) + start := clock.now().Add(-7 * 24 * time.Hour) + end := clock.now() + partial := insert(t, s, Entry{WindowStart: start, WindowEnd: end, Text: "window given"}) + if !partial.WindowStart.Equal(start) || !partial.WindowEnd.Equal(end) { + t.Errorf("window = %v..%v, want the one supplied %v..%v", partial.WindowStart, partial.WindowEnd, start, end) + } + if !partial.GeneratedAt.Equal(clock.now()) { + t.Errorf("generated_at = %v, want the store's clock for a run that did not state one", partial.GeneratedAt) + } + + // A GeneratedAt in the past with no WindowEnd: the fallback is the + // entry's own timestamp, not the clock. + past := clock.now().Add(-30 * time.Hour) + derived := insert(t, s, Entry{GeneratedAt: past, Text: "generated given"}) + if !derived.WindowEnd.Equal(past) { + t.Errorf("window_end = %v, want the supplied generated_at %v rather than the clock", derived.WindowEnd, past) + } + + // And what was written is what List reads back: the stamping is the + // store's, not a field the caller's copy got and the row did not. + rows, err := s.List(ctx, 10) + if err != nil { + t.Fatalf("List: %v", err) + } + if len(rows) != 3 { + t.Fatalf("listed %d runs, want 3", len(rows)) + } + for _, row := range rows { + if row.GeneratedAt.IsZero() || row.WindowEnd.IsZero() { + t.Errorf("run %d read back with a zero timestamp: %+v", row.ID, row) + } + } +} + +// Newest first, with ID as the tiebreak — the SQL's ORDER BY +// generated_at DESC, id DESC. A double that sorted on the timestamp alone +// would pass every other test in this file and still hand two runs +// sharing an instant back in an order the real query does not. +func TestMemoryStoreListsNewestFirstBreakingTiesByID(t *testing.T) { + s, clock := newTestStore() + + for _, text := range []string{"first", "second", "third"} { + insert(t, s, Entry{Text: text}) + clock.advance(time.Hour) + } + // Two runs in the same instant, which is what a replayed schedule or a + // clock with second-granularity storage produces. + same := clock.now() + older := insert(t, s, Entry{GeneratedAt: same, Text: "same instant, lower id"}) + newer := insert(t, s, Entry{GeneratedAt: same, Text: "same instant, higher id"}) + + rows, err := s.List(context.Background(), 10) + if err != nil { + t.Fatalf("List: %v", err) + } + if len(rows) != 5 { + t.Fatalf("listed %d runs, want 5", len(rows)) + } + if rows[0].ID != newer.ID || rows[1].ID != older.ID { + t.Errorf("the tied pair came back as %d then %d, want %d then %d", + rows[0].ID, rows[1].ID, newer.ID, older.ID) + } + if rows[2].Text != "third" { + t.Errorf("third row = %q, want the next-newest timestamp", rows[2].Text) + } + if rows[4].Text != "first" { + t.Errorf("last row = %q, want the oldest", rows[4].Text) + } +} + +// The listing is bounded, and a non-positive limit means the default +// rather than nothing. Both implementations clamp through ClampLimit; the +// Postgres one is not exercised here — this environment has no Postgres — +// so what is asserted is the rule they share. +func TestMemoryStoreBoundsTheListing(t *testing.T) { + s, _ := newTestStore() + for i := 0; i < DefaultHistoryLimit+5; i++ { + insert(t, s, Entry{Text: fmt.Sprintf("run %d", i)}) + } + + rows, err := s.List(context.Background(), 3) + if err != nil { + t.Fatalf("List: %v", err) + } + if len(rows) != 3 { + t.Errorf("listed %d runs with limit 3, want 3", len(rows)) + } + if rows[0].Text != fmt.Sprintf("run %d", DefaultHistoryLimit+4) { + t.Errorf("first row = %q, want the newest of the bounded set", rows[0].Text) + } + + for _, limit := range []int{0, -1} { + rows, err := s.List(context.Background(), limit) + if err != nil { + t.Fatalf("List(%d): %v", limit, err) + } + if len(rows) != DefaultHistoryLimit { + t.Errorf("listed %d runs with limit %d, want the default %d — a non-positive limit means the default, not none", + len(rows), limit, DefaultHistoryLimit) + } + } +} + +func TestClampLimit(t *testing.T) { + for _, tc := range []struct { + in, want int + }{ + {5, 5}, + {DefaultHistoryLimit, DefaultHistoryLimit}, + {MaxHistoryLimit, MaxHistoryLimit}, + {MaxHistoryLimit + 1, MaxHistoryLimit}, + {0, DefaultHistoryLimit}, + {-7, DefaultHistoryLimit}, + } { + if got := ClampLimit(tc.in); got != tc.want { + t.Errorf("ClampLimit(%d) = %d, want %d", tc.in, got, tc.want) + } + } +} + +// A returned listing is a copy, so a caller cannot edit the history by +// editing what it was handed — the thing a scanned row cannot do to a +// database either. +func TestMemoryStoreHandsOutCopies(t *testing.T) { + s, _ := newTestStore() + insert(t, s, Entry{Text: "the real text"}) + + rows, _ := s.List(context.Background(), 10) + rows[0].Text = "edited by a caller" + rows[0].ID = 999 + + again, _ := s.List(context.Background(), 10) + if again[0].Text != "the real text" || again[0].ID != 1 { + t.Errorf("stored run is %+v after a caller edited its copy", again[0]) + } +} + +// The store is written by the scheduler goroutine while HTTP handlers read +// it, which is the one piece of concurrency this feature actually has. +// Under -race this is what proves the mutex covers both paths. +func TestMemoryStoreIsSafeUnderConcurrentUse(t *testing.T) { + s := NewMemoryStore() + const writers, each = 8, 25 + + var wg sync.WaitGroup + for w := 0; w < writers; w++ { + wg.Add(1) + go func(w int) { + defer wg.Done() + for i := 0; i < each; i++ { + if _, err := s.Insert(context.Background(), Entry{Text: fmt.Sprintf("w%d-%d", w, i)}); err != nil { + t.Errorf("Insert: %v", err) + return + } + } + }(w) + } + wg.Add(1) + go func() { + defer wg.Done() + for i := 0; i < each; i++ { + if _, err := s.List(context.Background(), 5); err != nil { + t.Errorf("List: %v", err) + return + } + } + }() + wg.Wait() + + if got := s.Count(); got != writers*each { + t.Errorf("recorded %d runs, want %d", got, writers*each) + } +} diff --git a/httpapi/digest_handlers.go b/httpapi/digest_handlers.go new file mode 100644 index 0000000..4f68f2b --- /dev/null +++ b/httpapi/digest_handlers.go @@ -0,0 +1,149 @@ +package httpapi + +import ( + "net/http" + "time" + + "github.com/crydensync/cryden/v2" + + "github.com/crydensync/api/digest" +) + +// DigestHandlers answers the admin digest endpoints: the report itself, +// and the history of the ones the schedule produced. +type DigestHandlers struct { + Engine *cryden.Engine + + // Store is the digest history — this repo's own table. Nil unless + // DIGEST_INTERVAL_HOURS asked for a schedule, because the table is + // only ever written by the scheduler: with no schedule there is no + // history to read, and an empty list would be a lie about a + // deployment that has never built one. The handler answers 404, the + // same shape every other unconfigured feature in this API uses. + Store digest.Store +} + +// digestDefaultWindowDays is the reporting window used when the caller +// does not ask for another one. Seven days, matching cryden's own +// admin.DefaultDigestWindow — what makes the report a *weekly* digest — +// restated here rather than imported so this endpoint's bound and the +// engine's default are each readable where they are set, the same way +// hashMigrationDefaultWindowDays states the same week. +const digestDefaultWindowDays = 7 + +// digestRunDTO is one recorded run in the history listing. +type digestRunDTO struct { + ID int64 `json:"id"` + WindowStart time.Time `json:"window_start"` + WindowEnd time.Time `json:"window_end"` + GeneratedAt time.Time `json:"generated_at"` + + // Text is the report exactly as the engine rendered it, stored at + // the time and returned verbatim. It is included in a *listing* + // rather than behind a fetch-by-id, because there is no fetch-by-id: + // reading a past digest is the whole point of the history, and a + // list of windows with no reports in it would be an index to + // nothing. + Text string `json:"text"` +} + +// digestDTO is the on-demand report. +// +// Since and Until are reported alongside the text even though the text +// has the same window in its header, because a client should not have to +// parse English out of a report to learn what it covers — and because +// Until is this repo's clock reading, not something the engine returns. +type digestDTO struct { + Since time.Time `json:"since"` + Until time.Time `json:"until"` + WindowDays int `json:"window_days"` + Text string `json:"text"` +} + +// digestHistoryDTO is the whole history listing. Limit is echoed back so +// a console showing "20 of 143" knows which number it asked for without +// keeping its own copy of the default. +type digestHistoryDTO struct { + Runs []digestRunDTO `json:"runs"` + Count int `json:"count"` + Limit int `json:"limit"` +} + +// Digest — admin required (see router.go). Builds a digest of the audit +// history and returns it. Read-only in the strongest sense available: +// cryden builds it through an interface with no way to write anything +// (admin.AuditReader has no Record), and this handler records nothing +// either — asking twice does not put two entries in the history. See +// CLAUDE.md's hard rule. +// +// An optional window_days query parameter sets the window; it defaults +// to a week and is bounded, so a caller cannot ask for a window so wide +// the report stops meaning anything. +func (h *DigestHandlers) Digest(w http.ResponseWriter, r *http.Request) { + windowDays, err := queryInt(r, "window_days", digestDefaultWindowDays, 1, 365) + if err != nil { + writeBadRequest(w, err.Error()) + return + } + + // The window is computed here and passed to DigestSince, rather than + // calling WeeklyDigest and describing the window afterwards. Both + // produce the same report, but only this way is the Since reported + // above the exact instant the engine was asked to count from — the + // other would report a window a few microseconds out from the one + // the text describes. + since := time.Now().AddDate(0, 0, -windowDays) + + text, err := cryden.DigestSince(r.Context(), h.Engine, since) + if err != nil { + writeErr(w, err) + return + } + + writeData(w, http.StatusOK, digestDTO{ + Since: since.UTC(), + Until: time.Now().UTC(), + WindowDays: windowDays, + Text: text, + }) +} + +// DigestHistory — admin required (see router.go). Lists the digests the +// schedule has recorded, newest first. Read-only: nothing here creates, +// deletes or re-runs a digest, so an operator cannot manufacture history +// through the API. +func (h *DigestHandlers) DigestHistory(w http.ResponseWriter, r *http.Request) { + if h.Store == nil { + writeErr(w, errAdminStoresUnavailable) + return + } + + limit, err := queryInt(r, "limit", digest.DefaultHistoryLimit, 1, digest.MaxHistoryLimit) + if err != nil { + writeBadRequest(w, err.Error()) + return + } + + runs, err := h.Store.List(r.Context(), limit) + if err != nil { + writeErr(w, err) + return + } + + out := make([]digestRunDTO, 0, len(runs)) + for _, run := range runs { + out = append(out, digestRunDTO{ + ID: run.ID, + WindowStart: run.WindowStart, + WindowEnd: run.WindowEnd, + GeneratedAt: run.GeneratedAt, + Text: run.Text, + }) + } + + writeData(w, http.StatusOK, digestHistoryDTO{ + Runs: out, + Count: len(out), + Limit: limit, + }) +} diff --git a/httpapi/digest_handlers_test.go b/httpapi/digest_handlers_test.go new file mode 100644 index 0000000..3d567f9 --- /dev/null +++ b/httpapi/digest_handlers_test.go @@ -0,0 +1,439 @@ +package httpapi + +import ( + "context" + "encoding/json" + "net/http" + "net/http/httptest" + "strings" + "testing" + "time" + + "github.com/crydensync/cryden/v2" + "github.com/crydensync/cryden/v2/admin" + "github.com/crydensync/cryden/v2/store" + "github.com/crydensync/cryden/v2/store/memory" + "github.com/crydensync/cryden/v2/token" + + "github.com/crydensync/api/config" + "github.com/crydensync/api/digest" +) + +// digestResponse and digestHistoryResponse mirror the endpoints' DTOs +// field by field, so a renamed or dropped field fails here rather than +// silently changing the contract an operator's console reads. +type digestResponse struct { + Data struct { + Since time.Time `json:"since"` + Until time.Time `json:"until"` + WindowDays int `json:"window_days"` + Text string `json:"text"` + } `json:"data"` +} + +type digestHistoryResponse struct { + Data struct { + Runs []struct { + ID int64 `json:"id"` + WindowStart time.Time `json:"window_start"` + WindowEnd time.Time `json:"window_end"` + GeneratedAt time.Time `json:"generated_at"` + Text string `json:"text"` + } `json:"runs"` + Count int `json:"count"` + Limit int `json:"limit"` + } `json:"data"` +} + +type digestFixture struct { + engine *cryden.Engine + audit *memory.AuditStore + store *digest.MemoryStore + router http.Handler + + adminToken string + userToken string +} + +// newDigestFixture builds the engine on in-memory stores, with the audit +// store held directly — the digest is a report on exactly that store, so +// the fixture is the only thing that can seed it and check what came back. +func newDigestFixture(t *testing.T) digestFixture { + t.Helper() + ctx := context.Background() + + audit := memory.NewAuditStore() + store := digest.NewMemoryStore() + var adminID string + engine, err := cryden.New(cryden.Config{ + JWTSecret: "test-secret", + Users: memory.NewUserStore(), + Sessions: memory.NewSessionStore(), + Audit: audit, + Verifications: memory.NewVerificationStore(), + EmailSender: stubMailSender{}, + MagicLinkSender: stubMailSender{}, + AccessTokenClaims: token.ClaimsFunc(func(_ context.Context, userID string) (map[string]any, error) { + if userID == adminID { + return map[string]any{"role": "admin"}, nil + } + return nil, nil + }), + }) + if err != nil { + t.Fatalf("cryden.New on the in-memory stores: %v", err) + } + + admin, err := cryden.SignUp(ctx, engine, "operator@example.com", testPassword, "203.0.113.1") + if err != nil { + t.Fatalf("signup (operator): %v", err) + } + adminID = admin.ID + adminTokens, err := cryden.Login(ctx, engine, "operator@example.com", testPassword, "203.0.113.1", chromeOnMacOS) + if err != nil { + t.Fatalf("login (operator): %v", err) + } + + const userEmail = "dana@example.com" + if _, err := cryden.SignUp(ctx, engine, userEmail, testPassword, "203.0.113.2"); err != nil { + t.Fatalf("signup (user): %v", err) + } + userTokens, err := cryden.Login(ctx, engine, userEmail, testPassword, "203.0.113.2", chromeOnMacOS) + if err != nil { + t.Fatalf("login (user): %v", err) + } + + return digestFixture{ + engine: engine, + audit: audit, + store: store, + router: NewRouter(Deps{Engine: engine, Audit: audit, Digests: store, Config: config.Config{}}), + adminToken: adminTokens.AccessToken, + userToken: userTokens.AccessToken, + } +} + +func (f digestFixture) get(t *testing.T, path, token, query string) *httptest.ResponseRecorder { + t.Helper() + req := httptest.NewRequest(http.MethodGet, path+query, nil) + if token != "" { + req.Header.Set("Authorization", "Bearer "+token) + } + rec := httptest.NewRecorder() + f.router.ServeHTTP(rec, req) + return rec +} + +func (f digestFixture) digest(t *testing.T, token, query string) digestResponse { + t.Helper() + rec := f.get(t, "/v1/admin/digest", token, query) + if rec.Code != http.StatusOK { + t.Fatalf("digest: status = %d, want 200 (body %s)", rec.Code, rec.Body.String()) + } + var resp digestResponse + if err := json.Unmarshal(rec.Body.Bytes(), &resp); err != nil { + t.Fatalf("decoding %s: %v", rec.Body.String(), err) + } + return resp +} + +func (f digestFixture) history(t *testing.T, token, query string) digestHistoryResponse { + t.Helper() + rec := f.get(t, "/v1/admin/digest/history", token, query) + if rec.Code != http.StatusOK { + t.Fatalf("history: status = %d, want 200 (body %s)", rec.Code, rec.Body.String()) + } + var resp digestHistoryResponse + if err := json.Unmarshal(rec.Body.Bytes(), &resp); err != nil { + t.Fatalf("decoding %s: %v", rec.Body.String(), err) + } + return resp +} + +// record plants audit events of a type cryden does not define. That is +// what makes the assertions below exact: a host-specific type cannot +// collide with anything the engine records during signup and login, so +// the count in the report is this test's and nothing else's. +func (f digestFixture) record(t *testing.T, eventType store.AuditEventType, n int) { + t.Helper() + for i := 0; i < n; i++ { + if err := f.audit.Record(context.Background(), store.AuditEvent{Type: eventType}); err != nil { + t.Fatalf("recording a %s event: %v", eventType, err) + } + } +} + +// schedule runs the real Scheduler — the same object, on the same +// composition main.go builds — until it has recorded at least n digests, +// then stops it. Going through the scheduler rather than inserting rows by +// hand is the point: what these tests list is what the scheduled job +// produced. +func (f digestFixture) schedule(t *testing.T, n int) { + t.Helper() + ctx, cancel := context.WithCancel(context.Background()) + defer cancel() + + scheduler := &digest.Scheduler{ + Store: f.store, + Interval: 5 * time.Millisecond, + Build: func(ctx context.Context) (digest.Entry, error) { + since := time.Now().Add(-admin.DefaultDigestWindow) + text, err := cryden.DigestSince(ctx, f.engine, since) + if err != nil { + return digest.Entry{}, err + } + return digest.Entry{ + WindowStart: since.UTC(), + WindowEnd: time.Now().UTC(), + Text: text, + }, nil + }, + } + go scheduler.Run(ctx) + + deadline := time.Now().Add(5 * time.Second) + for f.store.Count() < n { + if time.Now().After(deadline) { + t.Fatalf("the scheduler recorded %d digests in five seconds, want %d", f.store.Count(), n) + } + time.Sleep(time.Millisecond) + } +} + +// The window is this repo's arithmetic, not the engine's: cryden's +// WeeklyDigest fixes seven days internally, so the only way to offer a +// configurable window is to compute it here and pass it to DigestSince. +// Both halves of that are asserted — the number echoed back, and the +// instant it was counted from. +func TestDigestCoversTheRequestedWindow(t *testing.T) { + f := newDigestFixture(t) + before := time.Now() + + resp := f.digest(t, f.adminToken, "") + if resp.Data.WindowDays != digestDefaultWindowDays { + t.Errorf("window_days = %d, want the default %d", resp.Data.WindowDays, digestDefaultWindowDays) + } + wantSince := before.AddDate(0, 0, -digestDefaultWindowDays) + if delta := resp.Data.Since.Sub(wantSince); delta > time.Minute || delta < -time.Minute { + t.Errorf("since = %v, want about %v (%v away)", resp.Data.Since, wantSince, delta) + } + // The window ends now, which is a reading only this repo can supply — + // the engine returns the text and nothing else. + if resp.Data.Until.Before(resp.Data.Since) { + t.Errorf("until = %v is before since = %v", resp.Data.Until, resp.Data.Since) + } + if resp.Data.Until.Before(before) { + t.Errorf("until = %v, want an instant at or after the request", resp.Data.Until) + } + if !strings.HasPrefix(resp.Data.Text, "Security digest") { + t.Errorf("text = %q, want the engine's report passed through verbatim", resp.Data.Text) + } + + wider := f.digest(t, f.adminToken, "?window_days=30") + if wider.Data.WindowDays != 30 { + t.Errorf("window_days = %d with window_days=30, want 30", wider.Data.WindowDays) + } + if !wider.Data.Since.Before(resp.Data.Since) { + t.Errorf("a thirty-day window starts at %v, want earlier than the seven-day window's %v", + wider.Data.Since, resp.Data.Since) + } +} + +// The report is built from the audit table the engine writes to, and this +// asserts it against a store the test holds and seeded itself. The type is +// one cryden does not define, which is also the proof that the report is +// the engine's own — a handler that reimplemented the digest would have no +// reason to know about a type it has never heard of. +func TestDigestCountsWhatTheAuditTableHolds(t *testing.T) { + f := newDigestFixture(t) + f.record(t, "host_specific_thing", 3) + + resp := f.digest(t, f.adminToken, "") + if !strings.Contains(resp.Data.Text, "types this engine does not define") { + t.Errorf("text = %q, want the section for types the engine does not define", resp.Data.Text) + } + if !strings.Contains(resp.Data.Text, "3 host_specific_thing") { + t.Errorf("text = %q, want the three events this test recorded", resp.Data.Text) + } + + // A window that reaches back further sees the same three, because they + // happened within the last seven days — the count is the store's answer + // to the window asked for, not a number the handler carries. + if got := f.digest(t, f.adminToken, "?window_days=1").Data.Text; !strings.Contains(got, "3 host_specific_thing") { + t.Errorf("text over one day = %q, want the events recorded moments ago", got) + } +} + +// window_days is bounded and rejected rather than clamped: a caller that +// asked for 5000 days and got 365 back has no way to tell that from a +// window that happens to hold the same history. +func TestDigestWindowIsBounded(t *testing.T) { + f := newDigestFixture(t) + + for _, query := range []string{"?window_days=0", "?window_days=366", "?window_days=week", "?window_days=-7"} { + rec := f.get(t, "/v1/admin/digest", f.adminToken, query) + if rec.Code != http.StatusBadRequest { + t.Errorf("%s: status = %d, want 400 (body %s)", query, rec.Code, rec.Body.String()) + } + } +} + +// The read-only property the whole feature rests on, asserted rather than +// promised: asking for a digest does not record one. Two requests, no +// rows. If this ever fails, an operator reading the history can no longer +// tell what the schedule produced from what somebody happened to open. +func TestDigestRecordsNothing(t *testing.T) { + f := newDigestFixture(t) + f.record(t, "host_specific_thing", 1) + + for i := 0; i < 2; i++ { + if text := f.digest(t, f.adminToken, "").Data.Text; text == "" { + t.Fatal("the digest came back empty") + } + } + if n := f.store.Count(); n != 0 { + t.Errorf("asking for a digest recorded %d runs, want 0", n) + } +} + +// The history shows what the schedule produced. The assertion on the +// planted count is what makes this a test of the whole path — scheduler, +// builder, store, endpoint — rather than of a handler reading rows a test +// inserted for it. +func TestDigestHistoryShowsWhatTheScheduleRecorded(t *testing.T) { + f := newDigestFixture(t) + f.record(t, "host_specific_thing", 3) + + f.schedule(t, 1) + + resp := f.history(t, f.adminToken, "") + if resp.Data.Count == 0 { + t.Fatalf("history is empty after the scheduler ran (count %d)", resp.Data.Count) + } + newest := resp.Data.Runs[0] + if !strings.Contains(newest.Text, "3 host_specific_thing") { + t.Errorf("the recorded digest does not hold the audit history it covered: %q", newest.Text) + } + if newest.ID == 0 || newest.GeneratedAt.IsZero() { + t.Errorf("run %+v was recorded without an id or a timestamp", newest) + } + if newest.WindowEnd.IsZero() || newest.WindowStart.IsZero() { + t.Errorf("run %+v was recorded without its window", newest) + } +} + +// A wired store with no runs answers an empty listing, not a 404: the +// schedule is configured and simply has not fired yet, which is a +// different thing from a deployment that has no history table at all. +// runs is [] rather than null, so a console can iterate it without a +// special case. +func TestDigestHistoryIsEmptyBeforeTheFirstRun(t *testing.T) { + f := newDigestFixture(t) + + rec := f.get(t, "/v1/admin/digest/history", f.adminToken, "") + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200 (body %s)", rec.Code, rec.Body.String()) + } + if !strings.Contains(rec.Body.String(), `"runs":[]`) { + t.Errorf("body = %s, want an empty runs array rather than null", rec.Body.String()) + } + + resp := f.history(t, f.adminToken, "") + if resp.Data.Count != 0 { + t.Errorf("count = %d with nothing scheduled yet, want 0", resp.Data.Count) + } + if resp.Data.Limit != digest.DefaultHistoryLimit { + t.Errorf("limit = %d, want the default %d", resp.Data.Limit, digest.DefaultHistoryLimit) + } +} + +// Ordering, the limit, and its bounds. The rows are inserted directly here +// because this is about the listing rather than about the schedule: the +// clock is controlled so newest-first is a fact and not a hope. +func TestDigestHistoryOrdersAndBoundsTheListing(t *testing.T) { + f := newDigestFixture(t) + base := time.Date(2026, 8, 1, 12, 0, 0, 0, time.UTC) + + for i := 0; i < 5; i++ { + if _, err := f.store.Insert(context.Background(), digest.Entry{ + WindowStart: base.Add(time.Duration(i) * time.Hour), + WindowEnd: base.Add(time.Duration(i+1) * time.Hour), + GeneratedAt: base.Add(time.Duration(i) * time.Hour), + Text: "report " + string(rune('a'+i)), + }); err != nil { + t.Fatalf("inserting run %d: %v", i, err) + } + } + + resp := f.history(t, f.adminToken, "") + if resp.Data.Count != 5 { + t.Fatalf("count = %d, want 5", resp.Data.Count) + } + if resp.Data.Runs[0].Text != "report e" { + t.Errorf("first row = %q, want the newest", resp.Data.Runs[0].Text) + } + if resp.Data.Runs[4].Text != "report a" { + t.Errorf("last row = %q, want the oldest", resp.Data.Runs[4].Text) + } + + limited := f.history(t, f.adminToken, "?limit=2") + if limited.Data.Count != 2 || limited.Data.Limit != 2 { + t.Errorf("count = %d limit = %d with limit=2, want 2 and 2", limited.Data.Count, limited.Data.Limit) + } + if limited.Data.Runs[0].Text != "report e" { + t.Errorf("first row = %q, want the newest of the bounded set", limited.Data.Runs[0].Text) + } + + // Bounded, not clamped — the same rule the logging endpoint's limit + // follows, so one word does not mean two things across the admin + // surface. + for _, query := range []string{"?limit=0", "?limit=201", "?limit=-1", "?limit=lots"} { + rec := f.get(t, "/v1/admin/digest/history", f.adminToken, query) + if rec.Code != http.StatusBadRequest { + t.Errorf("%s: status = %d, want 400 (body %s)", query, rec.Code, rec.Body.String()) + } + } +} + +// A router built without the history store answers 404 rather than 500 — +// a wiring fact, not a server fault, and the same shape every unconfigured +// feature in this API uses. Called through a router with no Digests, which +// is the state a deployment without DIGEST_INTERVAL_HOURS is in. +func TestDigestHistoryWithoutAStoreIsNotFound(t *testing.T) { + f := newDigestFixture(t) + router := NewRouter(Deps{Engine: f.engine, Audit: f.audit, Config: config.Config{}}) + + req := httptest.NewRequest(http.MethodGet, "/v1/admin/digest/history", nil) + req.Header.Set("Authorization", "Bearer "+f.adminToken) + rec := httptest.NewRecorder() + router.ServeHTTP(rec, req) + + if rec.Code != http.StatusNotFound { + t.Fatalf("status = %d, want 404 (body %s)", rec.Code, rec.Body.String()) + } + if !strings.Contains(rec.Body.String(), "not_configured") { + t.Errorf("body = %s, want the not_configured code", rec.Body.String()) + } +} + +// Both routes sit behind the same gate as every other admin report. A +// digest names accounts and failures, and read-only is not a reason to +// widen who can read it. +func TestDigestRoutesAreGatedByRequireAdmin(t *testing.T) { + f := newDigestFixture(t) + + for _, path := range []string{"/v1/admin/digest", "/v1/admin/digest/history"} { + if rec := f.get(t, path, "", ""); rec.Code != http.StatusUnauthorized { + t.Errorf("%s: no token gave %d, want 401", path, rec.Code) + } + if rec := f.get(t, path, "not-a-real-token", ""); rec.Code != http.StatusUnauthorized { + t.Errorf("%s: garbage token gave %d, want 401", path, rec.Code) + } + if rec := f.get(t, path, f.userToken, ""); rec.Code != http.StatusForbidden { + t.Errorf("%s: ordinary user gave %d, want 403 (body %s)", path, rec.Code, rec.Body.String()) + } + if rec := f.get(t, path, f.adminToken, ""); rec.Code != http.StatusOK { + t.Errorf("%s: operator gave %d, want 200 (body %s)", path, rec.Code, rec.Body.String()) + } + } +} diff --git a/httpapi/router.go b/httpapi/router.go index 26bc388..4d15786 100644 --- a/httpapi/router.go +++ b/httpapi/router.go @@ -8,6 +8,7 @@ import ( "github.com/crydensync/cryden/v2/store" "github.com/crydensync/api/config" + "github.com/crydensync/api/digest" "github.com/crydensync/api/shiplog" "github.com/crydensync/api/usermeta" "github.com/crydensync/api/webhook" @@ -59,6 +60,16 @@ type Deps struct { // off, nothing writes rows, and an empty list would be a lie about a // deployment that ships nothing. Shipped shiplog.Store + + // Digests backs GET /v1/admin/digest/history. Nil unless + // DIGEST_INTERVAL_HOURS asked for a schedule: the table is only ever + // written by the scheduler, so with no schedule there is no history to + // read, and the handler answers 404 rather than an empty list an + // operator would read as "nothing has ever happened". + // + // The on-demand GET /v1/admin/digest needs nothing from here — it + // reads the engine's audit history and records nothing. + Digests digest.Store } // NewRouter builds the full route table. Called once from main.go. @@ -81,6 +92,7 @@ func NewRouter(d Deps) http.Handler { metadata := &MetadataHandlers{Users: d.Users, Meta: d.Meta} hooks := &WebhookHandlers{Store: d.Hooks} logging := &LoggingHandlers{Store: d.Shipped} + digests := &DigestHandlers{Engine: engine, Store: d.Digests} mux := http.NewServeMux() @@ -208,5 +220,16 @@ func NewRouter(d Deps) http.Handler { // is the history. mux.HandleFunc("GET /v1/admin/logging/recent", RequireAdmin(engine, logging.Recent)) + // The weekly digest, and the history of the ones the schedule built. + // + // Two endpoints rather than one, because they answer different + // questions and only one of them can write. GET /v1/admin/digest + // reports on the window ending now and records nothing — asking twice + // leaves no trace. GET /v1/admin/digest/history reads what the + // scheduled job recorded, and nothing on this surface can create a + // row there. Both are read-only; see DigestHandlers. + mux.HandleFunc("GET /v1/admin/digest", RequireAdmin(engine, digests.Digest)) + mux.HandleFunc("GET /v1/admin/digest/history", RequireAdmin(engine, digests.DigestHistory)) + return mux } diff --git a/main.go b/main.go index 1089807..bc74ce0 100644 --- a/main.go +++ b/main.go @@ -5,16 +5,19 @@ import ( "database/sql" "log" "net/http" + "time" _ "github.com/lib/pq" "github.com/redis/go-redis/v9" "github.com/crydensync/cryden/v2" + "github.com/crydensync/cryden/v2/admin" "github.com/crydensync/cryden/v2/logger" "github.com/crydensync/cryden/v2/security" "github.com/crydensync/cryden/v2/store/postgres" "github.com/crydensync/api/config" + "github.com/crydensync/api/digest" "github.com/crydensync/api/httpapi" "github.com/crydensync/api/operator" "github.com/crydensync/api/shiplog" @@ -83,6 +86,16 @@ func main() { log.Printf("email templates loaded from %s", cfg.EmailTemplateDir) } + // Digest history: this repo's own table, written only by the scheduler + // below. Declared as the interface rather than as *digest.PostgresStore + // for the same reason webhookStore is — a typed nil in a non-nil + // interface passes every nil check and then panics on use, and the + // router's handler guards on exactly that check. + var digestStore digest.Store + if cfg.DigestInterval > 0 { + digestStore = digest.NewStore(db) + } + engineCfg := cryden.Config{ JWTSecret: cfg.JWTSecret, Users: users, @@ -282,6 +295,42 @@ func main() { log.Printf("webhook deliveries enabled: %d event types, up to %d attempts each", events, cfg.WebhookMaxAttempts) } + // The digest schedule. Started only when DIGEST_INTERVAL_HOURS asked + // for one, so an unconfigured deployment runs no goroutine and writes + // no rows — the same shape as the webhook worker above. + // + // The builder closes over the engine rather than this package taking + // one: the report itself is cryden's (DigestSince), and digest's job is + // only to record what it produced. The window is computed here and + // passed in, so the row states the exact interval the engine was asked + // to count over rather than one reconstructed from the text afterwards. + // + // Run takes context.Background() for the same reason the worker does: + // this repo still has no graceful shutdown, and that is noted in + // PROGRESS.md as owed rather than smuggled in behind a second + // goroutine. + if digestStore != nil { + scheduler := &digest.Scheduler{ + Store: digestStore, + Interval: cfg.DigestInterval, + Log: log.Default(), + Build: func(ctx context.Context) (digest.Entry, error) { + since := time.Now().Add(-admin.DefaultDigestWindow) + text, err := cryden.DigestSince(ctx, engine, since) + if err != nil { + return digest.Entry{}, err + } + return digest.Entry{ + WindowStart: since.UTC(), + WindowEnd: time.Now().UTC(), + Text: text, + }, nil + }, + } + go scheduler.Run(context.Background()) + log.Printf("scheduled digests enabled: one every %s", cfg.DigestInterval) + } + router := httpapi.NewRouter(httpapi.Deps{ Engine: engine, DB: db, @@ -292,6 +341,8 @@ func main() { Hooks: webhookStore, Shipped: shippedLog, + + Digests: digestStore, }) limiter := httpapi.NewEdgeRateLimiter(cfg.EdgeRateLimit, cfg.EdgeRateLimitWindow) handler := httpapi.WithCORS(cfg.CORSOrigins, httpapi.WithEdgeRateLimit(limiter, router)) diff --git a/migrations/012_digest_runs.down.sql b/migrations/012_digest_runs.down.sql new file mode 100644 index 0000000..049aa45 --- /dev/null +++ b/migrations/012_digest_runs.down.sql @@ -0,0 +1,4 @@ +-- 012_digest_runs.down.sql + +DROP INDEX IF EXISTS idx_digest_runs_generated; +DROP TABLE IF EXISTS digest_runs; diff --git a/migrations/012_digest_runs.up.sql b/migrations/012_digest_runs.up.sql new file mode 100644 index 0000000..e60e2d4 --- /dev/null +++ b/migrations/012_digest_runs.up.sql @@ -0,0 +1,50 @@ +-- 012_digest_runs.up.sql +-- +-- The digest history: one row per digest this deployment has produced on +-- a schedule, so the console can list past ones. +-- +-- cryden owns no scheduling concept at all. WeeklyDigest/DigestSince +-- build a report from the audit table and return a string; there is no +-- cron, no run record, and nothing anywhere in the engine that remembers +-- a digest was ever generated. Any "past digests" list is therefore this +-- repo's own infrastructure, and this is its table. +-- +-- The digest TEXT is stored rather than re-derivable counts, which is a +-- deliberate choice about what this table is for. A digest is a report +-- on a window that has ended: the events it counted are still in the +-- audit table, but re-running the same query later would not reproduce +-- it, because "last seven days" is anchored to when the digest was +-- built. Storing the rendered text is what makes reading a past digest +-- the same experience as reading a fresh one, and what makes the row +-- evidence of what was actually reported rather than a recipe for +-- something similar. +-- +-- The two window columns are stored alongside it because the text alone +-- cannot be sorted, filtered or compared, and an operator asking "which +-- windows did we cover" should not have to parse English out of a report +-- to find out. + +CREATE TABLE digest_runs ( + id BIGSERIAL PRIMARY KEY, + + -- The window this digest covers: [window_start, window_end). Stored + -- as the inclusive start and the instant the digest was built, which + -- is what DigestSince was asked for and what its own text prints, so + -- the row and the report never disagree about which week they cover. + window_start TIMESTAMPTZ NOT NULL, + window_end TIMESTAMPTZ NOT NULL, + + -- When this row was written. Distinct from window_end on purpose: + -- a run that is replayed or backfilled after an outage covers a + -- window that closed before the digest was made, and collapsing the + -- two would lose exactly that. + generated_at TIMESTAMPTZ NOT NULL DEFAULT now(), + + -- The rendered report, exactly as DigestSince returned it. + digest_text TEXT NOT NULL +); + +-- The listing's only view: newest first. The id tiebreak is what makes +-- the order total — two runs can share a timestamp, and an unstable sort +-- would show them in a different order on each request. +CREATE INDEX idx_digest_runs_generated ON digest_runs(generated_at DESC, id DESC); diff --git a/openapi/spec.yaml b/openapi/spec.yaml index 027a0d2..046eed3 100644 --- a/openapi/spec.yaml +++ b/openapi/spec.yaml @@ -1,7 +1,7 @@ openapi: 3.0.3 info: title: CrydenSync API - version: "1.3" + version: "1.4" description: > A self-hosted HTTP wrapper around the CrydenSync auth engine. Every response follows one of two envelope shapes: {"data": ...} @@ -27,6 +27,13 @@ info: that PUT and DELETE /admin/users/{userID}/metadata/{key} are this API's first endpoints that change what a subsequent token for another user will contain. See that path's description. + + 1.4 is additive: the weekly security digest and the history of the + ones the schedule built (GET /admin/digest, GET + /admin/digest/history). Both admin only. Neither changes any + existing path, field or status code, and neither can take an + action — the digest endpoint records nothing at all, and only a + background job this API runs writes the history. servers: - url: http://localhost:8080/v1 description: Local dev @@ -330,6 +337,36 @@ components: login emits many records and at second precision they would tie and nothing downstream could order them. + DigestRun: + type: object + description: > + One recorded digest. Text is the report exactly as the engine + rendered it at the time, stored rather than the counts behind it: + a digest covers a window that has ENDED, so re-running its query + later would not reproduce it — "the last seven days" is anchored + to when the digest was built. Storing what was reported is what + makes reading a past digest the same experience as reading a + fresh one. + properties: + id: { type: integer } + window_start: { type: string, format: date-time } + window_end: { type: string, format: date-time } + generated_at: + type: string + format: date-time + description: > + When the row was written. Distinct from window_end on purpose: + a run replayed after an outage covers a window that closed + before the digest was made. + text: + type: string + description: > + The rendered report, verbatim. Included in a listing rather + than behind a fetch-by-id, because there is no fetch-by-id — + reading a past digest is the whole point of the history, and + a list of windows with no reports in it would be an index to + nothing. + responses: BadRequest: description: Malformed request body @@ -1108,3 +1145,111 @@ paths: content: application/json: schema: { $ref: '#/components/schemas/ErrorResponse' } + + /admin/digest: + get: + summary: Build a security digest for a window ending now + description: > + Admin only. Returns cryden's DigestSince report for the window + ending at this request, verbatim — the text is the engine's and + this API does not reformat a report it does not own. since and + until are reported alongside it because a client should not have + to parse English out of a digest to learn what it covers, and + because until is this API's clock reading rather than something + the engine returns. + + Read-only in the strongest sense available: the report is built + through an interface with no way to write anything, and this + handler records nothing either. Asking twice leaves no trace, so + a digest in the history below is only ever something the + schedule produced. + security: [{ bearerAuth: [] }] + parameters: + - name: window_days + in: query + required: false + schema: { type: integer, minimum: 1, maximum: 365, default: 7 } + description: > + How far back to count. Default 7, which is what makes this a + weekly digest. Bounded and rejected rather than clamped, so a + caller that asked for 5000 days is told so instead of being + handed 365 and reading it as a week that happens to hold the + same history. + responses: + '200': + description: The report and the window it covers + content: + application/json: + schema: + type: object + properties: + data: + type: object + properties: + since: { type: string, format: date-time } + until: + type: string + format: date-time + description: The instant the report was built. Always now — a digest is a report on the run-up to the present. + window_days: { type: integer } + text: + type: string + description: > + The engine's rendered report, byte for byte. + Multi-line plain text, meant to be read as it is. + '400': { $ref: '#/components/responses/BadRequest' } + '401': { $ref: '#/components/responses/Unauthorized' } + '403': { $ref: '#/components/responses/Forbidden' } + + /admin/digest/history: + get: + summary: The digests the schedule has recorded, newest first + description: > + Admin only. Lists rows from this API's own digest_runs table, + which only the scheduler writes — nothing on the HTTP surface can + create, re-run or delete one, so an operator cannot manufacture + history through the API. + + cryden has no scheduling concept: WeeklyDigest/DigestSince build a + report on demand and return a string, with no run record. The + table, the job that fills it and this endpoint are all this API's + own. DIGEST_INTERVAL_HOURS turns the schedule on; it is opt-in, + and the first run happens one full interval after startup rather + than at boot, so a process that restarts often does not write one + row per restart. + security: [{ bearerAuth: [] }] + parameters: + - name: limit + in: query + required: false + schema: { type: integer, minimum: 1, maximum: 200, default: 20 } + responses: + '200': + description: The recorded runs, newest first + content: + application/json: + schema: + type: object + properties: + data: + type: object + properties: + runs: + type: array + items: { $ref: '#/components/schemas/DigestRun' } + count: { type: integer } + limit: + type: integer + description: The limit in force, so a console showing "20 of 143" need not keep its own copy of the default. + '400': { $ref: '#/components/responses/BadRequest' } + '401': { $ref: '#/components/responses/Unauthorized' } + '403': { $ref: '#/components/responses/Forbidden' } + '404': + description: > + not_configured — no digest_runs store is wired, which is what + an unset DIGEST_INTERVAL_HOURS produces. Answered rather than + returned as an empty list, which an operator would read as + "no digest has ever been built". + content: + application/json: + schema: { $ref: '#/components/schemas/ErrorResponse' } From 705b820223e4a6d65a47a356ceaeb00b3213a100 Mon Sep 17 00:00:00 2001 From: Raymond Nicholas Date: Tue, 15 Sep 2026 23:35:10 +0100 Subject: [PATCH 03/12] feat: add the support-ticket login diagnosis GET /v1/admin/support/diagnose reports why one account cannot sign in, from its own recorded history: the lock and its expiry, consecutive failed attempts, live sessions and the failures behind them. The report is cryden's, passed through verbatim; an unknown address is the answer rather than a 404, which is the engine's own distinction. Co-Authored-By: Claude Code --- README.md | 11 ++ httpapi/router.go | 8 ++ httpapi/support_handlers.go | 53 +++++++ httpapi/support_handlers_test.go | 240 +++++++++++++++++++++++++++++++ openapi/spec.yaml | 66 ++++++++- 5 files changed, 374 insertions(+), 4 deletions(-) create mode 100644 httpapi/support_handlers.go create mode 100644 httpapi/support_handlers_test.go diff --git a/README.md b/README.md index bd051fa..473a336 100644 --- a/README.md +++ b/README.md @@ -235,6 +235,17 @@ Set `PASSWORD_HASHER=argon2id` and every login whose stored hash is out of date - `estimated_remaining` is therefore *estimated*, floored at zero, and is `total_users - upgraded_events`. - `upgraded_events_in_window` is the field that actually answers "is this draining": the all-time count only ever rises, while a windowed one falls to zero as the last stragglers log in. `window_days` (1–365, default 7) sets that window. +`GET /v1/admin/support/diagnose?email=` answers the support ticket "why can't this person log in", from the account's own recorded history: whether it is locked and until when, its consecutive failed-attempt count, how many sessions it currently holds, and the recent failure-type events behind all of that, newest first. + +```json +{"data": {"email": "dana@example.com", "text": "Login diagnosis for dana@example.com\n\nAccount is LOCKED until 14:32 UTC.\n5 consecutive failed attempts currently recorded…"}} +``` + +- **An unknown address is the answer, not a 404.** cryden's `admin.DiagnoseLogin` returns `Found: false` rather than an error, and the report says "No account exists for this email address." A 404 would be indistinguishable from a broken endpoint, and "you have the wrong address" is exactly what a support agent pasting a typo'd email needs to be told. +- The text is the engine's, passed through verbatim. This repo does not reformat a report it does not own — and `email` is echoed alongside it so an agent working through a queue can see which address was answered. +- **Read-only structurally, not by convention.** The report is built through interfaces carrying no `LockAccount`, `ResetFailedAttempts` or `Revoke`, so the endpoint cannot unlock the very account it is describing, whatever the caller asks for. That is cryden's design and this repo adds nothing on top of it. +- A missing `email` is a `400`, not a diagnosis of the empty string — which would come back as "no account exists", an answer to a question nobody asked. + ## API keys `POST /v1/api-keys` mints a machine-to-machine credential for the calling user and returns the raw key **once** — cryden stores only its SHA-256 hash and can never reproduce it, so a caller that loses it has to mint a new one. The response carries the raw key, the stored record (`id`, `name`, `prefix`, `scopes`, `expires_at`, `expired`, `created_at`, `last_used_at`) and a `notice` saying so; a client that renders the key without that notice is the failure this guards against. diff --git a/httpapi/router.go b/httpapi/router.go index 4d15786..5a93422 100644 --- a/httpapi/router.go +++ b/httpapi/router.go @@ -93,6 +93,7 @@ func NewRouter(d Deps) http.Handler { hooks := &WebhookHandlers{Store: d.Hooks} logging := &LoggingHandlers{Store: d.Shipped} digests := &DigestHandlers{Engine: engine, Store: d.Digests} + support := &SupportHandlers{Engine: engine} mux := http.NewServeMux() @@ -231,5 +232,12 @@ func NewRouter(d Deps) http.Handler { mux.HandleFunc("GET /v1/admin/digest", RequireAdmin(engine, digests.Digest)) mux.HandleFunc("GET /v1/admin/digest/history", RequireAdmin(engine, digests.DigestHistory)) + // The support-ticket assistant: "why can't this person log in", + // answered from the account's own recorded history. Read-only by + // construction — cryden builds it through interfaces carrying no way + // to clear a lockout or reset a counter, so it cannot fix the account + // it is describing. See SupportHandlers. + mux.HandleFunc("GET /v1/admin/support/diagnose", RequireAdmin(engine, support.Diagnose)) + return mux } diff --git a/httpapi/support_handlers.go b/httpapi/support_handlers.go new file mode 100644 index 0000000..d66894f --- /dev/null +++ b/httpapi/support_handlers.go @@ -0,0 +1,53 @@ +package httpapi + +import ( + "net/http" + + "github.com/crydensync/cryden/v2" +) + +// SupportHandlers answers the support-ticket assistant. +type SupportHandlers struct { + Engine *cryden.Engine +} + +// supportDiagnosisDTO is one diagnosis. Text is the report the engine +// built; Email is echoed back because the whole point of this endpoint is +// to answer "why can't THIS person log in", and a support agent pasting +// addresses into a console should be able to see which one was answered +// without scrolling up. +type supportDiagnosisDTO struct { + Email string `json:"email"` + Text string `json:"text"` +} + +// Diagnose — admin required (see router.go). Answers the support ticket +// "why can't this user log in": whether the account is locked and until +// when, its failed-attempt count, how many sessions it holds, and its +// recent failure history. +// +// Read-only structurally, not by convention. cryden builds this through +// admin.DiagnoseLogin, which is handed narrow interfaces carrying no +// Create, LockAccount, ResetFailedAttempts or Revoke — so this endpoint +// cannot unlock the very account it is reporting on, whatever the caller +// asks for. See CLAUDE.md's hard rule. +// +// An account that does not exist is not an error: it is the answer, and +// the report says so. That distinction is the engine's own, and this +// handler passes it through rather than inventing a 404 the engine did +// not mean. +func (h *SupportHandlers) Diagnose(w http.ResponseWriter, r *http.Request) { + email := queryString(r, "email") + if email == "" { + writeBadRequest(w, "email is required") + return + } + + text, err := cryden.DiagnoseLoginIssue(r.Context(), h.Engine, email) + if err != nil { + writeErr(w, err) + return + } + + writeData(w, http.StatusOK, supportDiagnosisDTO{Email: email, Text: text}) +} diff --git a/httpapi/support_handlers_test.go b/httpapi/support_handlers_test.go new file mode 100644 index 0000000..024152c --- /dev/null +++ b/httpapi/support_handlers_test.go @@ -0,0 +1,240 @@ +package httpapi + +import ( + "context" + "encoding/json" + "net/http" + "net/http/httptest" + "strings" + "testing" + "time" + + "github.com/crydensync/cryden/v2" + "github.com/crydensync/cryden/v2/store/memory" + "github.com/crydensync/cryden/v2/token" + + "github.com/crydensync/api/config" +) + +// supportResponse mirrors the endpoint's DTO field by field, so a renamed +// or dropped field fails here rather than silently changing the contract a +// support console reads. +type supportResponse struct { + Data struct { + Email string `json:"email"` + Text string `json:"text"` + } `json:"data"` +} + +type supportFixture struct { + engine *cryden.Engine + router http.Handler + + adminToken string + userToken string + userEmail string +} + +// newSupportFixture builds an engine on the in-memory stores, on the +// lockout settings a real deployment runs — five failures then fifteen +// minutes, which is what config.Config defaults to. The threshold is set +// rather than left zero, because cryden's engine takes these straight from +// its config with no defaulting of its own: a zero threshold locks the +// account on the very first failure, which is not the deployment these +// tests are about. +func newSupportFixture(t *testing.T) supportFixture { + t.Helper() + ctx := context.Background() + + var adminID string + engine, err := cryden.New(cryden.Config{ + JWTSecret: "test-secret", + Users: memory.NewUserStore(), + Sessions: memory.NewSessionStore(), + Audit: memory.NewAuditStore(), + Verifications: memory.NewVerificationStore(), + EmailSender: stubMailSender{}, + MagicLinkSender: stubMailSender{}, + LockoutThreshold: 5, + LockoutDuration: 15 * time.Minute, + AccessTokenClaims: token.ClaimsFunc(func(_ context.Context, userID string) (map[string]any, error) { + if userID == adminID { + return map[string]any{"role": "admin"}, nil + } + return nil, nil + }), + }) + if err != nil { + t.Fatalf("cryden.New on the in-memory stores: %v", err) + } + + admin, err := cryden.SignUp(ctx, engine, "operator@example.com", testPassword, "203.0.113.1") + if err != nil { + t.Fatalf("signup (operator): %v", err) + } + adminID = admin.ID + adminTokens, err := cryden.Login(ctx, engine, "operator@example.com", testPassword, "203.0.113.1", chromeOnMacOS) + if err != nil { + t.Fatalf("login (operator): %v", err) + } + + const userEmail = "dana@example.com" + if _, err := cryden.SignUp(ctx, engine, userEmail, testPassword, "203.0.113.2"); err != nil { + t.Fatalf("signup (user): %v", err) + } + userTokens, err := cryden.Login(ctx, engine, userEmail, testPassword, "203.0.113.2", chromeOnMacOS) + if err != nil { + t.Fatalf("login (user): %v", err) + } + + return supportFixture{ + engine: engine, + router: NewRouter(Deps{Engine: engine, Config: config.Config{}}), + adminToken: adminTokens.AccessToken, + userToken: userTokens.AccessToken, + userEmail: userEmail, + } +} + +// failLogin makes one wrong-password attempt and insists it really failed. +// An assertion that a diagnosis shows failed attempts is worth nothing if +// the attempts before it quietly succeeded. +func (f supportFixture) failLogin(t *testing.T) { + t.Helper() + if _, err := cryden.Login(context.Background(), f.engine, f.userEmail, "not-the-password", "203.0.113.2", chromeOnMacOS); err == nil { + t.Fatal("a login with the wrong password succeeded") + } +} + +func (f supportFixture) diagnose(t *testing.T, token, query string) *httptest.ResponseRecorder { + t.Helper() + req := httptest.NewRequest(http.MethodGet, "/v1/admin/support/diagnose"+query, nil) + if token != "" { + req.Header.Set("Authorization", "Bearer "+token) + } + rec := httptest.NewRecorder() + f.router.ServeHTTP(rec, req) + return rec +} + +func (f supportFixture) diagnosis(t *testing.T, token, query string) supportResponse { + t.Helper() + rec := f.diagnose(t, token, query) + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200 (body %s)", rec.Code, rec.Body.String()) + } + var resp supportResponse + if err := json.Unmarshal(rec.Body.Bytes(), &resp); err != nil { + t.Fatalf("decoding %s: %v", rec.Body.String(), err) + } + return resp +} + +// The healthy case, which is what makes the failure cases mean anything: +// an account nobody has troubled reads as not locked, holding the one +// session its login created. The session count is the assertion that this +// report is reading live state rather than describing a template. +func TestDiagnoseReportsAHealthyAccount(t *testing.T) { + f := newSupportFixture(t) + + resp := f.diagnosis(t, f.adminToken, "?email="+f.userEmail) + if resp.Data.Email != f.userEmail { + t.Errorf("email = %q, want the address asked about echoed back", resp.Data.Email) + } + if !strings.Contains(resp.Data.Text, "Login diagnosis for "+f.userEmail) { + t.Errorf("text = %q, want the engine's report for this address", resp.Data.Text) + } + if !strings.Contains(resp.Data.Text, "Account is not locked.") { + t.Errorf("text = %q, want the account reported as unlocked", resp.Data.Text) + } + if !strings.Contains(resp.Data.Text, "0 consecutive failed attempts") { + t.Errorf("text = %q, want no failed attempts on an account that has none", resp.Data.Text) + } + if !strings.Contains(resp.Data.Text, "1 session is currently active") { + t.Errorf("text = %q, want the session the login created", resp.Data.Text) + } +} + +// The ticket this endpoint exists for: an account that cannot get in. +// Five wrong passwords trip the configured lockout, and both halves of +// that — the lock and the failures behind it — have to show up, because +// "locked until 14:32" and "five bad passwords in a row" lead a support +// agent to different answers. +func TestDiagnoseReportsALockedAccountAndWhy(t *testing.T) { + f := newSupportFixture(t) + for i := 0; i < 5; i++ { + f.failLogin(t) + } + + resp := f.diagnosis(t, f.adminToken, "?email="+f.userEmail) + if !strings.Contains(resp.Data.Text, "Account is LOCKED until") { + t.Errorf("text = %q, want the lockout reported", resp.Data.Text) + } + if !strings.Contains(resp.Data.Text, "5 consecutive failed attempts") { + t.Errorf("text = %q, want the five failures counted", resp.Data.Text) + } + // The history that explains it, not just the state. Both types are the + // engine's own records — nothing here is inferred by this endpoint. + for _, want := range []string{"Recent failure-type events", "login_failed", "account_locked"} { + if !strings.Contains(resp.Data.Text, want) { + t.Errorf("text = %q, want it to contain %q", resp.Data.Text, want) + } + } +} + +// An address nobody has is the answer, not a 404 and not an error. The +// distinction is cryden's own (admin.DiagnoseLogin returns Found=false +// rather than an error) and this handler passes it through instead of +// inventing a status the engine did not mean — a support agent asking +// about a typo'd address needs told there is no account, which a 404 +// would look identical to a broken endpoint. +func TestDiagnoseReportsAnUnknownAccountAsAnAnswer(t *testing.T) { + f := newSupportFixture(t) + + resp := f.diagnosis(t, f.adminToken, "?email=nobody@example.com") + if resp.Data.Email != "nobody@example.com" { + t.Errorf("email = %q, want the address asked about", resp.Data.Email) + } + if !strings.Contains(resp.Data.Text, "No account exists for this email address.") { + t.Errorf("text = %q, want the engine's own no-such-account answer", resp.Data.Text) + } +} + +// The address is the whole request, so a missing one is a 400 rather than +// a diagnosis of the empty string — which would come back as "no account +// exists", an answer to a question nobody asked. +func TestDiagnoseRequiresAnEmail(t *testing.T) { + f := newSupportFixture(t) + + for _, query := range []string{"", "?email=", "?email=%20", "?email=%20%20"} { + rec := f.diagnose(t, f.adminToken, query) + if rec.Code != http.StatusBadRequest { + t.Errorf("%q: status = %d, want 400 (body %s)", query, rec.Code, rec.Body.String()) + continue + } + if !strings.Contains(rec.Body.String(), "email") { + t.Errorf("%q: body = %s, want the missing parameter named", query, rec.Body.String()) + } + } +} + +// Every diagnosis names an account and its failure history, so this sits +// behind the same gate as the rest of the admin surface. Read-only is not +// a reason to widen who can read it. +func TestDiagnoseRouteIsGatedByRequireAdmin(t *testing.T) { + f := newSupportFixture(t) + query := "?email=" + f.userEmail + + if rec := f.diagnose(t, "", query); rec.Code != http.StatusUnauthorized { + t.Errorf("no token: status = %d, want 401", rec.Code) + } + if rec := f.diagnose(t, "not-a-real-token", query); rec.Code != http.StatusUnauthorized { + t.Errorf("garbage token: status = %d, want 401", rec.Code) + } + if rec := f.diagnose(t, f.userToken, query); rec.Code != http.StatusForbidden { + t.Errorf("ordinary user: status = %d, want 403 (body %s)", rec.Code, rec.Body.String()) + } + if rec := f.diagnose(t, f.adminToken, query); rec.Code != http.StatusOK { + t.Errorf("operator: status = %d, want 200 (body %s)", rec.Code, rec.Body.String()) + } +} diff --git a/openapi/spec.yaml b/openapi/spec.yaml index 046eed3..a02f593 100644 --- a/openapi/spec.yaml +++ b/openapi/spec.yaml @@ -30,10 +30,12 @@ info: 1.4 is additive: the weekly security digest and the history of the ones the schedule built (GET /admin/digest, GET - /admin/digest/history). Both admin only. Neither changes any - existing path, field or status code, and neither can take an - action — the digest endpoint records nothing at all, and only a - background job this API runs writes the history. + /admin/digest/history), plus the support-ticket assistant (GET + /admin/support/diagnose). All three admin only. None changes any + existing path, field or status code, and none can take an action — + the digest endpoint records nothing at all, the diagnosis cannot + unlock the account it describes, and only a background job this API + runs writes the digest history. servers: - url: http://localhost:8080/v1 description: Local dev @@ -1146,6 +1148,62 @@ paths: application/json: schema: { $ref: '#/components/schemas/ErrorResponse' } + /admin/support/diagnose: + get: + summary: Why one account cannot log in + description: > + Admin only. Answers the support ticket "why can't this person + log in" from the account's own recorded history: whether it is + locked and until when, its consecutive failed-attempt count, how + many sessions it holds right now, and the recent failure-type + events behind all of that, newest first. + + Read-only structurally, not by convention: the report is built + through interfaces carrying no LockAccount, ResetFailedAttempts + or Revoke, so this cannot unlock the very account it is + describing. + + An address with no account is NOT an error — it is the answer, + and the text says so. A 404 would be indistinguishable from a + broken endpoint, and "you have the wrong address" is exactly what + an agent pasting a typo needs to be told. + security: [{ bearerAuth: [] }] + parameters: + - name: email + in: query + required: true + schema: { type: string, format: email } + description: > + The account to diagnose. Required: absent or blank is a 400 + rather than a diagnosis of the empty string, which would come + back as "no account exists" — an answer to a question nobody + asked. + responses: + '200': + description: The diagnosis, as plain text + content: + application/json: + schema: + type: object + properties: + data: + type: object + properties: + email: + type: string + description: > + Echoed back, so an agent working through a queue + can see which address was answered. + text: + type: string + description: > + The engine's rendered report, byte for byte. + Multi-line plain text, meant to be read as it + is. + '400': { $ref: '#/components/responses/BadRequest' } + '401': { $ref: '#/components/responses/Unauthorized' } + '403': { $ref: '#/components/responses/Forbidden' } + /admin/digest: get: summary: Build a security digest for a window ending now From d74d8a4f6182b5a4a981ab1ebd8cd866f204b6b2 Mon Sep 17 00:00:00 2001 From: Raymond Nicholas Date: Tue, 15 Sep 2026 23:45:59 +0100 Subject: [PATCH 04/12] feat: add the config tuning advisor and put lockout in force MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit GET /v1/admin/config-tuning calls admin.BuildTuningReport directly for the structured suggestions, and accepts GET only — applying one means pre-filling a settings field a human saves, never this endpoint. Passing the lockout knobs to the engine is a real behaviour change: cryden does not default them, so the engine ran with both at zero. Co-Authored-By: Claude Code --- .env.example | 15 ++ README.md | 25 +++ config/config.go | 32 +++ config/config_test.go | 84 ++++++++ httpapi/router.go | 8 + httpapi/tuning_handlers.go | 176 ++++++++++++++++ httpapi/tuning_handlers_test.go | 363 ++++++++++++++++++++++++++++++++ main.go | 12 ++ openapi/spec.yaml | 111 +++++++++- 9 files changed, 822 insertions(+), 4 deletions(-) create mode 100644 httpapi/tuning_handlers.go create mode 100644 httpapi/tuning_handlers_test.go diff --git a/.env.example b/.env.example index b65c5c1..ed3f057 100644 --- a/.env.example +++ b/.env.example @@ -67,6 +67,21 @@ REDIS_URL= RATE_LIMIT_ATTEMPTS= RATE_LIMIT_WINDOW_SECONDS= +# Account lockout: after LOCKOUT_THRESHOLD consecutive failed passwords the +# account is locked for LOCKOUT_DURATION_MINUTES. Defaults are cryden's own +# (5 and 15), restated here because the engine does NOT fill these in — it +# reads whatever it is handed, and both zero values are wrong in the same +# direction. A zero threshold locks every account on its first bad +# password; a zero duration locks it until an instant already past, which +# is to say never. LOCKOUT_THRESHOLD below 1 is refused rather than read as +# "off", because cryden has no way to switch lockout off. +# +# These are also what GET /v1/admin/config-tuning quotes when it says a +# lockout setting is in force, so the report and the engine cannot disagree +# about what this deployment is running. +LOCKOUT_THRESHOLD= +LOCKOUT_DURATION_MINUTES= + # Password hashing. bcrypt is the engine's default; argon2id is the # current recommendation for new deployments (memory-hard, and the knob a # GPU attacker cannot parallelize around). Switching is safe at any time diff --git a/README.md b/README.md index 473a336..4dbedee 100644 --- a/README.md +++ b/README.md @@ -326,6 +326,31 @@ There is no vendor here: this repo ships no SDK, so "shipped" means "recorded in - An unknown `level` is a `400` naming the four valid values, not an empty list — which is indistinguishable from "the engine has been quiet". - The write is **synchronous**, on the goroutine that logged. That is a real cost and is not the shape a busy deployment wants; it is the shape this one can have, because an asynchronous sink needs a flush policy and a shutdown path, and this repo has no graceful shutdown anywhere yet. A buffer that is never flushed on exit is a log that silently drops its last records before a crash, which for a log is the failure that matters most. `LOG_LEVEL` (default `info`) is what keeps the volume sane in the meantime, since the engine's debug records never reach the sink. +## Config tuning advisor + +`GET /v1/admin/config-tuning?window_days=` reads the recent audit history, compares it against the settings **actually in force**, and returns the suggestions as a structured list — one object per knob, ready to render as a card: + +```json +{"data": { + "since": "…", "until": "…", "window_days": 30, + "counts": {"account_locked": 5, "login_failed": 5}, + "suggestions": [{ + "area": "Lockout", + "finding": "5 accounts were locked out of 5 recorded in this window (100%) — LockoutThreshold is currently 5, LockoutDuration 15m0s.", + "suggestion": "If most of these are real users mistyping a password rather than an attack, consider raising LockoutThreshold…" + }] +}} +``` + +- **There is no write path, and there is not going to be one.** No parameter changes a setting, and no counterpart endpoint applies a suggestion. The decision recorded for this surface is **pre-fill, never auto-apply**: a suggestion pre-fills the settings field it concerns, and a human still saves that change through the ordinary settings path. An endpoint that wrote a suggested value straight into live config would be the violation `CLAUDE.md`'s hard rule names, and would let a bad suggestion change production with no confirmation. The route accepts `GET` and nothing else. +- It calls `admin.BuildTuningReport` **directly**, not the flattened `cryden.ConfigTuningReport` text helper. The text is right for a CLI and wrong for a console: a pre-rendered blob cannot become one card per suggestion, and a client would be back to parsing English to find which knob a paragraph was about. +- `counts` is the raw audit evidence the suggestions were computed from, including event types cryden does not define — so a console can show the numbers rather than asking an operator to trust a sentence. +- Every value in the report comes from the config this process built the engine from, never a second reading of the environment. `UsingDefaultRateLimiter` is derived from `RedisURL` being empty — the same condition `main.go` uses to build the Redis limiter — so the report and the wiring cannot drift. +- `window_days` (1–365) defaults to cryden's own **30**-day tuning window, deliberately wider than the digest's week: a config knob should be judged against a month of traffic, not whatever happened this week. +- The password-strength finding always says the breach checker is not set. That is **accurate rather than a stub**: this repo has never wired `Config.BreachedPasswordChecker`, because cryden ships no implementation and every real one calls somebody else's corpus. + +`LOCKOUT_THRESHOLD` and `LOCKOUT_DURATION_MINUTES` (defaults 5 and 15 minutes) are passed through to the engine explicitly, for the two reasons above at once: cryden reads them straight off its config with no defaulting, so a deployment that left them implicit was running a lockout that could never actually trigger; and the tuning report describes the settings in force, so it should not have to guess what the engine was handed. + ## Weekly digest Two endpoints, and only one of them depends on any configuration: diff --git a/config/config.go b/config/config.go index debd264..b495f19 100644 --- a/config/config.go +++ b/config/config.go @@ -116,6 +116,22 @@ type Config struct { RateLimitAttempts int RateLimitWindow time.Duration + // LockoutThreshold and LockoutDuration are the engine's account + // lockout bounds: after LockoutThreshold consecutive failed attempts the + // account is locked for LockoutDuration. Defaults are cryden's own (5 + // and 15 minutes), restated here for the same reason the rate-limit + // bounds are — and with a sharper edge, because the engine does not + // fill these in either way. A zero threshold locks an account on its + // very first failed password; a zero duration locks it until an instant + // already past, which is to say not at all. + // + // They are also what GET /v1/admin/config-tuning describes when it says + // a lockout setting is in force. Passed through to the engine rather + // than left implicit, so the report and the engine cannot disagree + // about what this deployment is actually running. + LockoutThreshold int + LockoutDuration time.Duration + // PasswordHasher selects which algorithm NEW password hashes are // written with — PasswordHasherBcrypt (the engine's default) or // PasswordHasherArgon2id. Switching is safe at any time and needs no @@ -409,6 +425,22 @@ func Load() (Config, error) { return cfg, err } + // Account lockout. Defaulted rather than left at zero — see the field + // comments: a zero threshold would lock every account on its first + // failed password, which is the opposite of a default. A threshold + // below 1 is refused for the same reason rather than read as "off": + // cryden has no way to switch lockout off, so a 0 here can only be a + // typo, and honouring it would lock every account on one bad password. + if cfg.LockoutThreshold, err = envInt("LOCKOUT_THRESHOLD", 5); err != nil { + return cfg, err + } + if cfg.LockoutThreshold < 1 { + return cfg, fmt.Errorf("LOCKOUT_THRESHOLD must be at least 1, got %d — cryden has no way to switch account lockout off", cfg.LockoutThreshold) + } + if cfg.LockoutDuration, err = envMinutes("LOCKOUT_DURATION_MINUTES", 15*time.Minute); err != nil { + return cfg, err + } + // Password hashing. Bcrypt is the engine's default, so the only thing // this repo has to do for it is not pass a hasher — but the argon2id // parameters are assembled either way, because the hash-migration diff --git a/config/config_test.go b/config/config_test.go index b1333b0..3570df8 100644 --- a/config/config_test.go +++ b/config/config_test.go @@ -44,6 +44,9 @@ var tieredEnvVars = []string{ "WEBHOOK_SECRET", "WEBHOOK_EVENTS", "WEBHOOK_MAX_ATTEMPTS", + "LOCKOUT_THRESHOLD", + "LOCKOUT_DURATION_MINUTES", + "DIGEST_INTERVAL_HOURS", } func loadForTest(t *testing.T, env map[string]string) (Config, error) { @@ -400,3 +403,84 @@ func TestTier3WebhookMaxAttemptsIsBounded(t *testing.T) { } } } + +// The Tier 4 defaults: cryden's own lockout numbers restated, because the +// engine takes them straight off its config with no defaulting of its own +// and both zero values are wrong in the same direction — a zero threshold +// locks an account on its first failed password, a zero duration locks it +// until an instant already past. +func TestTier4DefaultsComeFromTheEngine(t *testing.T) { + cfg, err := loadForTest(t, nil) + if err != nil { + t.Fatalf("Load() failed with only the required vars set: %v", err) + } + + if cfg.LockoutThreshold != 5 { + t.Errorf("LockoutThreshold = %d, want cryden's own default 5", cfg.LockoutThreshold) + } + if cfg.LockoutDuration != 15*time.Minute { + t.Errorf("LockoutDuration = %s, want cryden's own default 15m", cfg.LockoutDuration) + } + // Digests are opt-in: no schedule unless one was asked for, so an + // unconfigured deployment runs no goroutine and writes no rows. + if cfg.DigestInterval != 0 { + t.Errorf("DigestInterval = %s, want 0 (no schedule)", cfg.DigestInterval) + } +} + +func TestTier4EnvOverridesLeaveOtherKnobsDefaulted(t *testing.T) { + cfg, err := loadForTest(t, map[string]string{ + "LOCKOUT_THRESHOLD": "9", + "LOCKOUT_DURATION_MINUTES": "45", + "DIGEST_INTERVAL_HOURS": "168", + }) + if err != nil { + t.Fatalf("Load() failed: %v", err) + } + + if cfg.LockoutThreshold != 9 { + t.Errorf("LockoutThreshold = %d, want 9", cfg.LockoutThreshold) + } + if cfg.LockoutDuration != 45*time.Minute { + t.Errorf("LockoutDuration = %s, want 45m", cfg.LockoutDuration) + } + // A week in hours, which is the shape the env var is written in even + // though everything downstream holds a duration. + if cfg.DigestInterval != 168*time.Hour { + t.Errorf("DigestInterval = %s, want 168h", cfg.DigestInterval) + } + // Untouched knobs stay on their defaults. + if cfg.RateLimitAttempts != 10 || cfg.RateLimitWindow != time.Minute { + t.Errorf("rate limit = %d per %s, want the untouched default 10 per minute", cfg.RateLimitAttempts, cfg.RateLimitWindow) + } +} + +// A lockout threshold below 1 and a negative digest interval are both +// refused rather than read as "off": cryden has no way to switch account +// lockout off, and treating a typo as the default is how a setting an +// operator meant to change silently does nothing. +func TestTier4UnusableKnobValuesAreStartupErrors(t *testing.T) { + cases := []struct { + name string + env map[string]string + want string + }{ + {"threshold of zero", map[string]string{"LOCKOUT_THRESHOLD": "0"}, "LOCKOUT_THRESHOLD must be at least 1"}, + {"negative threshold", map[string]string{"LOCKOUT_THRESHOLD": "-1"}, "LOCKOUT_THRESHOLD must be at least 1"}, + {"non-numeric threshold", map[string]string{"LOCKOUT_THRESHOLD": "five"}, "LOCKOUT_THRESHOLD must be a number"}, + {"non-numeric lockout duration", map[string]string{"LOCKOUT_DURATION_MINUTES": "quarter of an hour"}, "LOCKOUT_DURATION_MINUTES must be a number of minutes"}, + {"negative digest interval", map[string]string{"DIGEST_INTERVAL_HOURS": "-1"}, "DIGEST_INTERVAL_HOURS cannot be negative"}, + {"non-numeric digest interval", map[string]string{"DIGEST_INTERVAL_HOURS": "weekly"}, "DIGEST_INTERVAL_HOURS must be a number"}, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + _, err := loadForTest(t, tc.env) + if err == nil { + t.Fatalf("%v was accepted, want an error", tc.env) + } + if !strings.Contains(err.Error(), tc.want) { + t.Errorf("error = %q, want it to contain %q", err, tc.want) + } + }) + } +} diff --git a/httpapi/router.go b/httpapi/router.go index 5a93422..1ef7c6b 100644 --- a/httpapi/router.go +++ b/httpapi/router.go @@ -94,6 +94,7 @@ func NewRouter(d Deps) http.Handler { logging := &LoggingHandlers{Store: d.Shipped} digests := &DigestHandlers{Engine: engine, Store: d.Digests} support := &SupportHandlers{Engine: engine} + tuning := &TuningHandlers{Audit: d.Audit, Config: d.Config} mux := http.NewServeMux() @@ -239,5 +240,12 @@ func NewRouter(d Deps) http.Handler { // it is describing. See SupportHandlers. mux.HandleFunc("GET /v1/admin/support/diagnose", RequireAdmin(engine, support.Diagnose)) + // The config tuning advisor. Suggestions only: there is no endpoint + // that applies one, and no parameter that changes a setting — the + // recorded decision is that a suggestion pre-fills the settings field + // it concerns and a human saves that change through the ordinary + // settings path. See TuningHandlers and CLAUDE.md's hard rule. + mux.HandleFunc("GET /v1/admin/config-tuning", RequireAdmin(engine, tuning.ConfigTuning)) + return mux } diff --git a/httpapi/tuning_handlers.go b/httpapi/tuning_handlers.go new file mode 100644 index 0000000..4140c22 --- /dev/null +++ b/httpapi/tuning_handlers.go @@ -0,0 +1,176 @@ +package httpapi + +import ( + "net/http" + "time" + + "github.com/crydensync/cryden/v2/admin" + "github.com/crydensync/cryden/v2/store" + + "github.com/crydensync/api/config" +) + +// TuningHandlers answers the config tuning advisor. +type TuningHandlers struct { + // Audit is the same store instance main.go handed cryden (see Deps). + // It satisfies admin.AuditReader as it stands — the point of that + // interface is that Record is not in it, so a report built through it + // has no way to write. + Audit store.AuditStore + + // Config is what the suggestions are judged against: the settings + // actually in force, read back off the config this process built. + Config config.Config +} + +// tuningSuggestionDTO is one observation and the change it points at. +// The three fields are cryden's own TuningSuggestion, unchanged — the +// console renders one card per suggestion from them, which is exactly why +// this endpoint returns the structured list rather than the flattened +// prose ConfigTuningReport produces. +type tuningSuggestionDTO struct { + // Area names the knob the suggestion concerns. cryden's names, not + // this repo's: "Lockout", "Rate limiting", "Anomaly detection", + // "Credential stuffing", "Password strength". + Area string `json:"area"` + Finding string `json:"finding"` + Suggestion string `json:"suggestion"` +} + +// tuningDTO is the whole report. +// +// Counts is included alongside the suggestions even though the prose +// above already quotes the numbers: it is the raw audit counts the +// suggestions were computed from, so a console can show the evidence +// behind a recommendation instead of asking an operator to trust a +// sentence. Keys are audit event types, including ones this engine does +// not define. +type tuningDTO struct { + Since time.Time `json:"since"` + Until time.Time `json:"until"` + WindowDays int `json:"window_days"` + Counts map[string]int `json:"counts"` + Suggestions []tuningSuggestionDTO `json:"suggestions"` +} + +// ConfigTuning — admin required (see router.go). Summarises recent audit +// history against this deployment's own tuning settings and suggests +// changes worth considering. +// +// # There is no write path here, and there is not going to be one +// +// This endpoint suggests. It has no counterpart that applies a +// suggestion, and no parameter that changes a setting — the decision +// already recorded for this surface is that a suggestion PRE-FILLS the +// settings field it concerns and a human still saves that change through +// the ordinary settings path. An endpoint that wrote a suggested value +// straight into live config would be exactly the violation CLAUDE.md's +// hard rule names, and it would also mean a bad suggestion silently +// changing production with no confirmation step. +// +// cryden's side of that is structural: admin.BuildTuningReport is handed +// an AuditReader with no Record on it, so the report cannot even write an +// audit event of its own. Running this twice is exactly like running it +// once. +// +// # Why the structured report +// +// This calls admin.BuildTuningReport directly rather than +// cryden.ConfigTuningReport, which wraps it and returns that report's +// Text(). The text is the right thing for a CLI and the wrong thing for a +// console: a pre-rendered blob cannot be rendered as one card per +// suggestion, and a client would be back to parsing English to find out +// which knob a paragraph was about. The Text() rendering remains the +// engine's own, and this repo does not reimplement it. +// +// An optional window_days query parameter sets the window; it defaults to +// cryden's own 30-day tuning window, which is deliberately wider than the +// digest's week — a config knob should be judged against a month of +// traffic, not whatever happened to occur this week. +func (h *TuningHandlers) ConfigTuning(w http.ResponseWriter, r *http.Request) { + if h.Audit == nil { + writeErr(w, errAdminStoresUnavailable) + return + } + + windowDays, err := queryInt(r, "window_days", int(admin.DefaultTuningWindow.Hours()/24), 1, 365) + if err != nil { + writeBadRequest(w, err.Error()) + return + } + + since := time.Now().AddDate(0, 0, -windowDays) + + report, err := admin.BuildTuningReport(r.Context(), h.Audit, h.inputs(), since) + if err != nil { + writeErr(w, err) + return + } + + suggestions := make([]tuningSuggestionDTO, 0, len(report.Suggestions)) + for _, s := range report.Suggestions { + suggestions = append(suggestions, tuningSuggestionDTO{ + Area: s.Area, + Finding: s.Finding, + Suggestion: s.Suggestion, + }) + } + + writeData(w, http.StatusOK, tuningDTO{ + Since: report.Since, + Until: report.Until, + WindowDays: windowDays, + Counts: eventCounts(report.Counts), + Suggestions: suggestions, + }) +} + +// inputs assembles the snapshot of current settings the report judges +// history against. Every value comes from the config this process +// actually built the engine from — never a second reading of the +// environment, which could disagree with what the engine is running. +// +// Two of these are worth reading the reasoning for: +// +// - UsingDefaultRateLimiter is derived from RedisURL being empty, which +// is the same condition main.go uses to decide whether to construct a +// Redis limiter. Derived rather than stored as its own flag, so the +// report and the wiring cannot drift. +// - BreachedPasswordCheckerActive is always false today, and that is +// accurate rather than a stub: this repo has never set +// Config.BreachedPasswordChecker, because cryden ships no checker +// implementation and every real one calls somebody else's breach +// corpus. So the report says the checker is not set, which is true, +// and a deployment that wires one in changes this line and the +// engine wiring together. +func (h *TuningHandlers) inputs() admin.TuningInputs { + return admin.TuningInputs{ + LockoutThreshold: h.Config.LockoutThreshold, + LockoutDuration: h.Config.LockoutDuration, + RateLimitAttempts: h.Config.RateLimitAttempts, + RateLimitWindow: h.Config.RateLimitWindow, + UsingDefaultRateLimiter: h.Config.RedisURL == "", + + AnomaliesEnabled: h.Config.AnomalyDetection, + UserFailureVelocity: h.Config.AnomalyThresholds.UserFailureVelocity, + IPFailureVelocity: h.Config.AnomalyThresholds.IPFailureVelocity, + HistorySize: h.Config.AnomalyThresholds.HistorySize, + StuffingTargetAccounts: h.Config.CredentialStuffingThresholds.TargetAccounts, + StuffingWindow: h.Config.CredentialStuffingThresholds.Window, + StuffingCooldown: h.Config.CredentialStuffingThresholds.Cooldown, + + BreachedPasswordCheckerActive: false, + } +} + +// eventCounts narrows the report's typed count map to plain strings for +// JSON. Done explicitly rather than relying on the encoder: a map key of +// a defined string type marshals the same way, but only by accident of +// encoding/json, and a response shape should not depend on that. +func eventCounts(counts map[store.AuditEventType]int) map[string]int { + out := make(map[string]int, len(counts)) + for eventType, n := range counts { + out[string(eventType)] = n + } + return out +} diff --git a/httpapi/tuning_handlers_test.go b/httpapi/tuning_handlers_test.go new file mode 100644 index 0000000..c38d72c --- /dev/null +++ b/httpapi/tuning_handlers_test.go @@ -0,0 +1,363 @@ +package httpapi + +import ( + "context" + "encoding/json" + "net/http" + "net/http/httptest" + "strings" + "testing" + "time" + + "github.com/crydensync/cryden/v2" + "github.com/crydensync/cryden/v2/admin" + "github.com/crydensync/cryden/v2/store" + "github.com/crydensync/cryden/v2/store/memory" + "github.com/crydensync/cryden/v2/token" + + "github.com/crydensync/api/config" +) + +// tuningResponse mirrors the endpoint's DTO field by field, so a renamed +// or dropped field fails here rather than silently changing the contract a +// console reads — and this one is rendered as a card per suggestion, so a +// renamed field is a broken screen rather than a cosmetic difference. +type tuningResponse struct { + Data struct { + Since time.Time `json:"since"` + Until time.Time `json:"until"` + WindowDays int `json:"window_days"` + Counts map[string]int `json:"counts"` + Suggestions []struct { + Area string `json:"area"` + Finding string `json:"finding"` + Suggestion string `json:"suggestion"` + } `json:"suggestions"` + } `json:"data"` +} + +// findByArea returns the suggestion for one knob, and whether there was +// one. Area is the field a console groups on, so finding by it is finding +// the thing the endpoint promises. +func (r tuningResponse) findByArea(area string) (string, bool) { + for _, s := range r.Data.Suggestions { + if s.Area == area { + return s.Finding + " " + s.Suggestion, true + } + } + return "", false +} + +type tuningFixture struct { + audit *memory.AuditStore + router http.Handler + + adminToken string + userToken string +} + +// newTuningFixture builds an engine on in-memory stores with the audit +// store held directly, because the report is computed from exactly that +// store and the test is the only thing that can seed it. cfg is the +// deployment under test — the report describes the settings in force, so +// the settings are the input. +func newTuningFixture(t *testing.T, cfg config.Config) tuningFixture { + t.Helper() + ctx := context.Background() + + audit := memory.NewAuditStore() + var adminID string + engine, err := cryden.New(cryden.Config{ + JWTSecret: "test-secret", + Users: memory.NewUserStore(), + Sessions: memory.NewSessionStore(), + Audit: audit, + Verifications: memory.NewVerificationStore(), + EmailSender: stubMailSender{}, + MagicLinkSender: stubMailSender{}, + AccessTokenClaims: token.ClaimsFunc(func(_ context.Context, userID string) (map[string]any, error) { + if userID == adminID { + return map[string]any{"role": "admin"}, nil + } + return nil, nil + }), + }) + if err != nil { + t.Fatalf("cryden.New on the in-memory stores: %v", err) + } + + admin, err := cryden.SignUp(ctx, engine, "operator@example.com", testPassword, "203.0.113.1") + if err != nil { + t.Fatalf("signup (operator): %v", err) + } + adminID = admin.ID + adminTokens, err := cryden.Login(ctx, engine, "operator@example.com", testPassword, "203.0.113.1", chromeOnMacOS) + if err != nil { + t.Fatalf("login (operator): %v", err) + } + + const userEmail = "dana@example.com" + if _, err := cryden.SignUp(ctx, engine, userEmail, testPassword, "203.0.113.2"); err != nil { + t.Fatalf("signup (user): %v", err) + } + userTokens, err := cryden.Login(ctx, engine, userEmail, testPassword, "203.0.113.2", chromeOnMacOS) + if err != nil { + t.Fatalf("login (user): %v", err) + } + + return tuningFixture{ + audit: audit, + router: NewRouter(Deps{Engine: engine, Audit: audit, Config: cfg}), + adminToken: adminTokens.AccessToken, + userToken: userTokens.AccessToken, + } +} + +func (f tuningFixture) record(t *testing.T, eventType store.AuditEventType, n int) { + t.Helper() + for i := 0; i < n; i++ { + if err := f.audit.Record(context.Background(), store.AuditEvent{Type: eventType}); err != nil { + t.Fatalf("recording a %s event: %v", eventType, err) + } + } +} + +func (f tuningFixture) report(t *testing.T, method, token, query string) *httptest.ResponseRecorder { + t.Helper() + req := httptest.NewRequest(method, "/v1/admin/config-tuning"+query, nil) + if token != "" { + req.Header.Set("Authorization", "Bearer "+token) + } + rec := httptest.NewRecorder() + f.router.ServeHTTP(rec, req) + return rec +} + +func (f tuningFixture) tuning(t *testing.T, token, query string) tuningResponse { + t.Helper() + rec := f.report(t, http.MethodGet, token, query) + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200 (body %s)", rec.Code, rec.Body.String()) + } + var resp tuningResponse + if err := json.Unmarshal(rec.Body.Bytes(), &resp); err != nil { + t.Fatalf("decoding %s: %v", rec.Body.String(), err) + } + return resp +} + +// The lockout suggestion quotes the threshold and duration in force, and +// this drives the whole path to check it quotes the CONFIGURED ones rather +// than cryden's documented defaults. That distinction is the reason this +// repo passes the lockout knobs to the engine at all: a report describing +// settings the engine is not running would be worse than no report. +func TestConfigTuningJudgesHistoryAgainstTheSettingsInForce(t *testing.T) { + cfg := config.Config{ + LockoutThreshold: 7, + LockoutDuration: 30 * time.Minute, + RateLimitAttempts: 42, + RateLimitWindow: 90 * time.Second, + } + f := newTuningFixture(t, cfg) + + // Five lockouts against five failures: every failed attempt ended in a + // lockout, which is the shape cryden flags as "probably catching real + // users". + f.record(t, store.EventAccountLocked, 5) + f.record(t, store.EventLoginFailed, 5) + f.record(t, "host_specific_thing", 3) + + resp := f.tuning(t, f.adminToken, "") + + lockout, ok := resp.findByArea("Lockout") + if !ok { + t.Fatalf("no Lockout suggestion in %+v, want one for five lockouts", resp.Data.Suggestions) + } + if !strings.Contains(lockout, "LockoutThreshold is currently 7") { + t.Errorf("lockout finding = %q, want the configured threshold", lockout) + } + if !strings.Contains(lockout, "30m0s") { + t.Errorf("lockout finding = %q, want the configured duration", lockout) + } + // And it suggests rather than acts: nothing in the text claims a + // change was made, and no field carries one. + if !strings.Contains(lockout, "consider") { + t.Errorf("lockout text = %q, want a suggestion, not an instruction", lockout) + } + + // The counts are the evidence behind the suggestions, reported raw so a + // console can show them rather than asking an operator to trust a + // sentence. A type cryden does not define is included, which is what + // makes this the engine's own count rather than a list this repo keeps. + if got := resp.Data.Counts[string(store.EventAccountLocked)]; got != 5 { + t.Errorf("counts[account_locked] = %d, want 5", got) + } + if got := resp.Data.Counts["host_specific_thing"]; got != 3 { + t.Errorf("counts[host_specific_thing] = %d, want 3", got) + } + + // The rate limiter suggestion quotes the configured bounds for the same + // reason. + limiter, ok := resp.findByArea("Rate limiting") + if !ok { + t.Fatalf("no Rate limiting suggestion in %+v, want one for the in-process limiter", resp.Data.Suggestions) + } + if !strings.Contains(limiter, "42 attempts per 1m30s") { + t.Errorf("rate limiter finding = %q, want the configured bounds", limiter) + } +} + +// A Redis-backed limiter is the one suggestion that needs no audit data at +// all — it is a fact about the deployment — so it has to disappear when +// the deployment has actually fixed it. Derived from RedisURL, which is +// the same condition main.go uses to build the Redis limiter. +func TestConfigTuningOnlyFlagsTheInProcessRateLimiter(t *testing.T) { + t.Run("in-process", func(t *testing.T) { + f := newTuningFixture(t, config.Config{RateLimitAttempts: 10, RateLimitWindow: time.Minute}) + if _, ok := f.tuning(t, f.adminToken, "").findByArea("Rate limiting"); !ok { + t.Error("no Rate limiting suggestion with the default in-process limiter") + } + }) + + t.Run("shared", func(t *testing.T) { + f := newTuningFixture(t, config.Config{ + RedisURL: "redis://cache.internal:6379", + RateLimitAttempts: 10, + RateLimitWindow: time.Minute, + }) + if _, ok := f.tuning(t, f.adminToken, "").findByArea("Rate limiting"); ok { + t.Error("the in-process limiter was flagged on a deployment running the Redis one") + } + }) +} + +// Anomaly detection off and on are different reports, and the off case is +// the one every deployment starts in. Both are asserted because the +// suggestion is the only place a console learns detection is not running. +func TestConfigTuningReportsWhetherAnomalyDetectionIsOn(t *testing.T) { + t.Run("off", func(t *testing.T) { + f := newTuningFixture(t, config.Config{}) + finding, ok := f.tuning(t, f.adminToken, "").findByArea("Anomaly detection") + if !ok { + t.Fatal("no Anomaly detection suggestion with detection off") + } + if !strings.Contains(finding, "is not set") { + t.Errorf("finding = %q, want it to say detection is not configured", finding) + } + }) + + t.Run("on and quiet", func(t *testing.T) { + cfg := config.Config{AnomalyDetection: true} + cfg.AnomalyThresholds.UserFailureVelocity = 5 + cfg.AnomalyThresholds.IPFailureVelocity = 20 + cfg.AnomalyThresholds.HistorySize = 100 + f := newTuningFixture(t, cfg) + // Enough successful sign-ins that silence is worth remarking on + // rather than being a quiet window. + f.record(t, store.EventLoginSuccess, 60) + + finding, ok := f.tuning(t, f.adminToken, "").findByArea("Anomaly detection") + if !ok { + t.Fatal("no Anomaly detection suggestion with detection on and 60 sign-ins, none flagged") + } + // The thresholds come from this repo's config, so the report quotes + // the numbers the engine is screening against. + for _, want := range []string{"UserFailureVelocity=5", "IPFailureVelocity=20", "HistorySize=100"} { + if !strings.Contains(finding, want) { + t.Errorf("finding = %q, want it to contain %q", finding, want) + } + } + }) +} + +// The password-strength finding is the one that is always present here, +// and it is accurate rather than a stub: this repo has never set a +// breached-password checker, because cryden ships no implementation and +// every real one calls somebody else's corpus. +func TestConfigTuningReportsTheAbsentBreachChecker(t *testing.T) { + f := newTuningFixture(t, config.Config{}) + + finding, ok := f.tuning(t, f.adminToken, "").findByArea("Password strength") + if !ok { + t.Fatal("no Password strength suggestion, want the always-present one") + } + if !strings.Contains(finding, "BreachedPasswordChecker is not set") { + t.Errorf("finding = %q, want it to say the checker is not configured", finding) + } +} + +// window_days defaults to cryden's own 30-day tuning window — deliberately +// wider than the digest's week, because a config knob should be judged +// against a month of traffic — and is bounded rather than clamped. +func TestConfigTuningWindowBounds(t *testing.T) { + f := newTuningFixture(t, config.Config{}) + + if got := f.tuning(t, f.adminToken, "").Data.WindowDays; got != int(admin.DefaultTuningWindow.Hours()/24) { + t.Errorf("window_days = %d, want cryden's default %v", got, admin.DefaultTuningWindow) + } + if got := f.tuning(t, f.adminToken, "?window_days=90").Data.WindowDays; got != 90 { + t.Errorf("window_days = %d with window_days=90, want 90", got) + } + for _, query := range []string{"?window_days=0", "?window_days=366", "?window_days=month"} { + rec := f.report(t, http.MethodGet, f.adminToken, query) + if rec.Code != http.StatusBadRequest { + t.Errorf("%s: status = %d, want 400 (body %s)", query, rec.Code, rec.Body.String()) + } + } +} + +// The route answers GET and nothing else. This is the HTTP-level half of +// "suggests, never applies": there is no POST that takes a suggestion, and +// a suggestion cannot be turned into a config change by a request to this +// path in any method. The recorded decision is that applying one means +// pre-filling a settings field and a human saving it through the ordinary +// settings path — see CLAUDE.md's hard rule. +func TestConfigTuningAcceptsNoWriteMethod(t *testing.T) { + f := newTuningFixture(t, config.Config{}) + + for _, method := range []string{http.MethodPost, http.MethodPut, http.MethodPatch, http.MethodDelete} { + rec := f.report(t, method, f.adminToken, "") + if rec.Code == http.StatusOK { + t.Errorf("%s: status = %d, want this route to refuse anything but GET (body %s)", method, rec.Code, rec.Body.String()) + } + } +} + +// A handler built without the audit store answers 404 rather than 500 — a +// wiring fact, not a server fault, and the same shape every unconfigured +// feature in this API uses. Called directly, because a router built with a +// nil engine cannot serve an authenticated request at all and this is the +// only way to reach the branch. +func TestConfigTuningWithoutAnAuditStoreIsNotFound(t *testing.T) { + req := httptest.NewRequest(http.MethodGet, "/v1/admin/config-tuning", nil) + rec := httptest.NewRecorder() + + h := &TuningHandlers{} + h.ConfigTuning(rec, req) + + if rec.Code != http.StatusNotFound { + t.Fatalf("status = %d, want 404 (body %s)", rec.Code, rec.Body.String()) + } + if !strings.Contains(rec.Body.String(), "not_configured") { + t.Errorf("body = %s, want the not_configured code", rec.Body.String()) + } +} + +// The reports name failed logins and lockouts by account, so this sits +// behind the same gate as every other admin report. +func TestConfigTuningRouteIsGatedByRequireAdmin(t *testing.T) { + f := newTuningFixture(t, config.Config{}) + + if rec := f.report(t, http.MethodGet, "", ""); rec.Code != http.StatusUnauthorized { + t.Errorf("no token: status = %d, want 401", rec.Code) + } + if rec := f.report(t, http.MethodGet, "not-a-real-token", ""); rec.Code != http.StatusUnauthorized { + t.Errorf("garbage token: status = %d, want 401", rec.Code) + } + if rec := f.report(t, http.MethodGet, f.userToken, ""); rec.Code != http.StatusForbidden { + t.Errorf("ordinary user: status = %d, want 403 (body %s)", rec.Code, rec.Body.String()) + } + if rec := f.report(t, http.MethodGet, f.adminToken, ""); rec.Code != http.StatusOK { + t.Errorf("operator: status = %d, want 200 (body %s)", rec.Code, rec.Body.String()) + } +} diff --git a/main.go b/main.go index bc74ce0..9b05914 100644 --- a/main.go +++ b/main.go @@ -196,6 +196,18 @@ func main() { // cryden's own documented trade-off for the shared limiter. engineCfg.RateLimitAttempts = cfg.RateLimitAttempts engineCfg.RateLimitWindow = cfg.RateLimitWindow + + // Account lockout, passed through rather than left implicit. cryden + // reads these straight off its config with no defaulting of its own, + // and the zero values are both wrong in the same direction: a zero + // threshold locks an account on its first failed password, and a zero + // duration locks it until an instant already past — which is to say + // never. config.Load defaults them to cryden's own numbers (5 failures, + // 15 minutes) so the engine runs what this deployment believes it runs, + // and so GET /v1/admin/config-tuning describes settings that are + // actually in force. + engineCfg.LockoutThreshold = cfg.LockoutThreshold + engineCfg.LockoutDuration = cfg.LockoutDuration if cfg.RedisURL != "" { redisOpts, err := redis.ParseURL(cfg.RedisURL) if err != nil { diff --git a/openapi/spec.yaml b/openapi/spec.yaml index a02f593..13db5af 100644 --- a/openapi/spec.yaml +++ b/openapi/spec.yaml @@ -30,12 +30,15 @@ info: 1.4 is additive: the weekly security digest and the history of the ones the schedule built (GET /admin/digest, GET - /admin/digest/history), plus the support-ticket assistant (GET - /admin/support/diagnose). All three admin only. None changes any + /admin/digest/history), the support-ticket assistant (GET + /admin/support/diagnose) and the config tuning advisor (GET + /admin/config-tuning). All four admin only. None changes any existing path, field or status code, and none can take an action — the digest endpoint records nothing at all, the diagnosis cannot - unlock the account it describes, and only a background job this API - runs writes the digest history. + unlock the account it describes, the tuning advisor answers GET and + nothing else and its suggestions pre-fill a settings field for a + human to save rather than being applied, and only a background job + this API runs writes the digest history. servers: - url: http://localhost:8080/v1 description: Local dev @@ -369,6 +372,37 @@ components: a list of windows with no reports in it would be an index to nothing. + TuningSuggestion: + type: object + description: > + One thing worth reconsidering, split into three parts rather than + a sentence. This API returns the structured form + (admin.BuildTuningReport) and not cryden's flattened + ConfigTuningReport text helper, because a console renders one + card per suggestion and a pre-rendered blob cannot be turned + back into cards. The three fields are the three things a console + needs: which knob, what the evidence says, and what might be + worth changing. + properties: + area: + type: string + description: > + Which knob this concerns — the field a console groups or + de-duplicates on. Stable strings, not localized text. + finding: + type: string + description: > + What the recent history shows, in terms of the settings + currently in force. A sentence meant to be read by a human, + not parsed. + suggestion: + type: string + description: > + What might be worth changing. Wording is deliberately + suggestive ("consider ..."), never imperative: nothing here + has been applied, and nothing that reads this response can + apply it either. + responses: BadRequest: description: Malformed request body @@ -1148,6 +1182,75 @@ paths: application/json: schema: { $ref: '#/components/schemas/ErrorResponse' } + /admin/config-tuning: + get: + summary: Settings worth reconsidering, judged against recent history + description: > + Admin only. Summarises the recent audit history against the + settings THIS deployment is actually running and returns the + suggestions as a structured list — one object per knob, ready to + render as a card. Calls admin.BuildTuningReport directly rather + than the flattened cryden.ConfigTuningReport text helper: the + text is right for a CLI and wrong for a console, which cannot + turn a pre-rendered blob into one card per suggestion. + + THERE IS NO WRITE PATH, and there is not going to be one. No + parameter here changes a setting and no counterpart endpoint + applies a suggestion. The recorded decision is pre-fill, never + auto-apply: a suggestion pre-fills the settings field it + concerns and a human still saves that change through the + ordinary settings path. The route accepts GET and nothing else, + and the report is built through an interface with no Record on + it, so producing one cannot even write an audit event. + security: [{ bearerAuth: [] }] + parameters: + - name: window_days + in: query + required: false + schema: { type: integer, minimum: 1, maximum: 365, default: 30 } + description: > + How far back to judge. Default 30 — cryden's own tuning + window, deliberately wider than the digest's week, because a + config knob should be judged against a month of traffic + rather than whatever happened this week. + responses: + '200': + description: The suggestions and the counts they were computed from + content: + application/json: + schema: + type: object + properties: + data: + type: object + properties: + since: { type: string, format: date-time } + until: { type: string, format: date-time } + window_days: { type: integer } + counts: + type: object + additionalProperties: { type: integer } + description: > + The raw audit counts the suggestions were + computed from, so a console can show the + evidence rather than asking an operator to + trust a sentence. Includes event types cryden + does not define, if this deployment records + any. + suggestions: + type: array + items: { $ref: '#/components/schemas/TuningSuggestion' } + '400': { $ref: '#/components/responses/BadRequest' } + '401': { $ref: '#/components/responses/Unauthorized' } + '403': { $ref: '#/components/responses/Forbidden' } + '404': + description: > + not_configured — the router was built without the audit + store, so there is no history to judge anything against. + content: + application/json: + schema: { $ref: '#/components/schemas/ErrorResponse' } + /admin/support/diagnose: get: summary: Why one account cannot log in From 68e0a57f67f4f2bfe144cb47fac534caa14d3d81 Mon Sep 17 00:00:00 2001 From: Raymond Nicholas Date: Tue, 15 Sep 2026 23:47:27 +0100 Subject: [PATCH 05/12] docs: document Tier 4 Stage 1 and the lockout passthrough MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Records the digest, support diagnosis and tuning endpoints in CURRENT-STATE.md and NEXT.md, and logs the session in PROGRESS.md — including that 012 has never been applied anywhere and that passing the lockout knobs to the engine changes what every deployment does. Co-Authored-By: Claude Code --- docs/development/CURRENT-STATE.md | 91 ++++++++++++++++++-- docs/development/NEXT.md | 39 +++++++++ docs/development/PROGRESS.md | 134 ++++++++++++++++++++++++++++++ 3 files changed, 257 insertions(+), 7 deletions(-) diff --git a/docs/development/CURRENT-STATE.md b/docs/development/CURRENT-STATE.md index fb84a9f..6ce57b5 100644 --- a/docs/development/CURRENT-STATE.md +++ b/docs/development/CURRENT-STATE.md @@ -15,9 +15,12 @@ has its own `httpapi/apple.go` — see `NEXT.md` Tier 1). Tier 2 added one admin endpoint on top of those, the first in this repo — see below. Tier 3 added three more admin endpoints and this repo's first three tables of its own, plus the config that lights up Argon2id, -cloud logging and email templates — see below. Every admin endpoint in -this repo is either read-only or an explicit operator action on a named -key; nothing on that surface applies a suggestion by itself. +cloud logging and email templates — see below. Tier 4's Stage 1 added +three more admin endpoints and this repo's fourth table — the weekly +digest and its recorded history, the support-ticket login diagnosis and +the config tuning advisor. Every admin endpoint in this repo is either +read-only or an explicit operator action on a named key; nothing on that +surface applies a suggestion by itself. Tier 1 also added the second-factor surface: TOTP enroll/confirm/ disable, passkey registration/list/delete, magic-link request/complete, @@ -305,10 +308,84 @@ in-memory double, not against Postgres `FOR UPDATE SKIP LOCKED`, and that double cannot reproduce two workers racing. `PROGRESS.md` says all of this plainly. -## Tier 4 and 5 +## Tier 4 — AI-assisted admin endpoints (Stage 1): DONE + +Built in two stages on `feat/tier4-ai-admin-endpoints`, for the same +reason Tier 3 was: the three read-only reports below had their decisions +already made in `NEXT.md`, while Stage 2 needs two decisions that are +not a build session's to make (see the end of this section). `go build`, +`go vet`, `gofmt -l` and `go test ./...` are clean, `httpapi` is also +green under `-race`, and `PROGRESS.md` records what that does and does +not cover. + +Everything here is `RequireAdmin`, read-only, and buildable on the +engine alone — no LLM, no second database connection, no outbound call: + +- **`GET /v1/admin/digest`** and **`GET /v1/admin/digest/history`**. + The engine's `cryden.DigestSince` renders a report over a window on + demand; the history is this repo's own (`digest/`, + `migrations/012_digest_runs`), written only by the scheduled job. The + on-demand endpoint **records nothing**, so an operator hitting it + twenty times does not fill the history with twenty near-identical + reports. The engine has no scheduling concept at all, so the job, the + table and the history endpoint are all entirely this repo's. The first + run lands one full interval after startup rather than at boot, since a + process that restarts more often than the interval elapses would + otherwise write a row per restart. Unset `DIGEST_INTERVAL_HOURS` (or + `0`) means no schedule, no goroutine and a `404 not_configured` + history, while the on-demand endpoint keeps working. +- **`GET /v1/admin/support/diagnose?email=...`** → + `cryden.DiagnoseLoginIssue`. An unknown address is an **answer** + (`found: false`), not a `404`: "we have never seen this address" is + what a support ticket needs to be told. It describes a locked account; + it cannot unlock one. +- **`GET /v1/admin/config-tuning`** → `admin.BuildTuningReport` called + **directly**, not `cryden.ConfigTuningReport`. The structured + `TuningSuggestion{Area, Finding, Suggestion}` list is the point — a + console renders one card per suggestion, and a pre-rendered text blob + cannot be turned back into cards. The raw audit `counts` are returned + alongside, so the evidence is visible rather than a sentence asking to + be trusted. `window_days` defaults to cryden's own 30 days (wider than + the digest's week on purpose — a config knob should be judged against + a month of traffic) and is bounded rather than clamped. The route + accepts **GET and nothing else**, which is the HTTP-level half of the + pre-fill-never-auto-apply decision: there is no POST that takes a + suggestion, and applying one means pre-filling a settings field a + human saves through the ordinary settings path. + +**One behaviour change came out of this tier, and it was not the point +of it.** The tuning report is asked to judge the audit history against +"the settings in force", and building it surfaced that this repo had +never passed `LockoutThreshold`/`LockoutDuration` to the engine — so +every deployment so far ran with both at Go's zero value, and cryden +defaults neither. A zero threshold locks an account on its first failed +password; a zero duration locks it until an instant already past, which +is to say not at all. `config` now owns both knobs (cryden's own 5 and +15 minutes by default), `main.go` passes them, and a threshold below 1 +is a startup error rather than being read as "off", because cryden has +no way to switch lockout off. It is a real change to what every existing +deployment does on its next restart, so it is called out in `README.md`, +`.env.example` and its commit message rather than buried. + +The digest's new table (`012`) has **never been applied to a real +database**, the same as `009`–`011` — there is no Postgres in this +sandbox — and the digest schedule is a goroutine on +`context.Background()`, because this repo still has no graceful +shutdown. `PROGRESS.md` says both plainly. + +## Tier 4 Stage 2, and Tier 5 Not started. See `NEXT.md` for the full, ordered, specced-in-detail -queue. Tier 4 is all behind `RequireAdmin` and stays read-only by -construction, with the decision already made that an AI suggestion -**pre-fills** a settings form and never auto-applies. +queue. Stage 2 is the LLM provider config, the read-only database +provider config and the ask-AI widget config; +`ai.LLMProvider`/`ai.QueryableStore` have no implementation in this repo +yet, so nothing in it has an endpoint. Two decisions are open and were +left for the user rather than guessed at: whether this repo ships a live +LLM client against a real vendor (an outbound integration, which this +repo has so far shipped none of), and where the at-rest encryption key +for the stored provider credential comes from (reusing the existing +`ENCRYPTION_KEY` is the obvious candidate and still a decision with a +blast radius). Tier 4 stays read-only by construction, with the +decision already made that an AI suggestion **pre-fills** a settings +form and never auto-applies. diff --git a/docs/development/NEXT.md b/docs/development/NEXT.md index 396d0b2..41b42e8 100644 --- a/docs/development/NEXT.md +++ b/docs/development/NEXT.md @@ -16,6 +16,10 @@ Tier 2 is done — see the status note under Tier 2 and `PROGRESS.md`'s migrations to copy. Tier 3 is done, in two stages on `feat/tier3-config-and-endpoints` — see the status note under Tier 3 and `PROGRESS.md`'s 2026-09-15 entries. +Tier 4 is **in progress** on `feat/tier4-ai-admin-endpoints`: Stage 1 +(digest + scheduling + history, support diagnosis, config tuning +advisor) is built — see the status note under Tier 4. Stage 2 (the LLM +and database providers, and the ask-AI widget config) is not started. --- @@ -284,6 +288,41 @@ Two details were decided rather than assumed, and are recorded in ## Tier 4 — AI-assisted admin endpoints (all behind `RequireAdmin`) +> **Status: Stage 1 is built on `feat/tier4-ai-admin-endpoints`.** +> The weekly digest and its schedule and history, the support-ticket +> assistant and the config tuning advisor all exist, are wired in +> `main.go`, and are tested end to end on the in-memory stores — +> `go build`/`go vet`/`go test ./...` clean, `httpapi` also green under +> `-race`. Stage 2 — the LLM provider config, the read-only database +> provider config and the ask-AI widget config — is **not started**: +> `ai.LLMProvider` and `ai.QueryableStore` still have no implementation +> in this repo, so nothing in Stage 2 has an endpoint yet. +> +> What is still owed from Stage 1, said plainly: **migration +> `012_digest_runs` has never been applied to a database** (no Postgres +> in this sandbox — the same is true of `009`–`011`), so the digest +> history's real `PostgresStore` has only been reasoned about, not run; +> and the digest schedule is a goroutine on `context.Background()`, +> because this repo still has no graceful shutdown. +> +> Two things Stage 1 changed that were not in the spec below, both +> recorded because they are behaviour rather than plumbing: +> +> - **`LOCKOUT_THRESHOLD`/`LOCKOUT_DURATION_MINUTES` are now passed to +> the engine.** cryden does not default these — it reads whatever it +> is handed, and `0`/`0` means an account is locked on its first +> failed password until an instant already past, which is to say +> never. Until this tier the engine ran with both at zero. So this is +> a real behaviour change, not a tidy-up, and it is why the tuning +> report can quote the lockout settings in force rather than cryden's +> documented defaults. +> - **`GET /v1/admin/digest` records nothing.** The spec puts scheduling +> and history in this repo, and that is still exactly where the +> writing happens — but the on-demand endpoint deliberately does not +> write a row, so an operator hitting it twenty times does not fill +> the history with twenty near-identical reports. Only the scheduled +> job writes. + Every endpoint in this tier stays read-only/surface-only, no exceptions — see `CLAUDE.md`'s hard rule at the top. diff --git a/docs/development/PROGRESS.md b/docs/development/PROGRESS.md index edaf397..42cae43 100644 --- a/docs/development/PROGRESS.md +++ b/docs/development/PROGRESS.md @@ -678,3 +678,137 @@ smuggled in behind the other. Tier 3 is complete. Next is Tier 4, which stays read-only by construction with the pre-fill-never-auto-apply decision already made. + +## 2026-09-15 — Tier 4, Stage 1 (digest, support diagnosis, config tuning) + +Tier 4 is split for the same reason Tier 3 was: the first half is three +read-only reports with their decisions already made in `NEXT.md`, and +the second half needs two decisions that are not this session's to +make (see "Stage 2" below). Branch `feat/tier4-ai-admin-endpoints`. + +Three commits, one logical step each: the digest and its history +(`21ac94c`), the support-ticket login diagnosis (`705b820`), and the +config tuning advisor (`d74d8a4`). + +What each one is, and the one thing about it worth knowing: + +- **`GET /v1/admin/digest`** and **`GET /v1/admin/digest/history`**. + `digest/` is a new repo-owned package (interface + `PostgresStore` + + in-memory double in one file, the convention every store here + follows) over `migrations/012_digest_runs`. The on-demand endpoint + **records nothing**: an operator hitting it twenty times should not + fill a history with twenty near-identical reports, so only the + scheduled job writes. The schedule is this repo's own — cryden has no + concept of one — and the first run lands one full interval after + startup, not at boot, because a process that restarts more often than + the interval elapses would otherwise write a row per restart. +- **`GET /v1/admin/support/diagnose?email=`** → `cryden.DiagnoseLoginIssue`. + An unknown account is an **answer** (`Found:false`), not a 404 or a + 500: "we have never seen this address" is exactly what a support + ticket needs to be told, and dressing it up as a server error would + hide it. +- **`GET /v1/admin/config-tuning`** → `admin.BuildTuningReport` called + **directly**, not `cryden.ConfigTuningReport`. The structured + `TuningSuggestion{Area, Finding, Suggestion}` list is the point: a + console renders one card per suggestion, and a pre-rendered text blob + cannot be turned back into cards. The counts are returned raw + alongside, so the evidence is visible rather than a sentence asking + to be trusted. + +### The lockout passthrough is a behaviour change, not a tidy-up + +Building the tuning advisor surfaced something: **this repo was never +passing `LockoutThreshold`/`LockoutDuration` to the engine.** `config` +had no such fields, so `main.go` left them at Go's zero values, so the +engine ran with a threshold of 0 and a duration of 0 — and cryden does +no defaulting of either. A zero threshold locks an account on its very +first failed password; a zero duration locks it until an instant +already past, which is to say not at all. Every deployment of this API +so far has been in that second state. + +The report is what made it visible: `BuildTuningReport` is asked to +judge the audit history against "the settings in force", and the +settings in force were not what anyone thought they were. So `config` +gained `LockoutThreshold`/`LockoutDuration` (defaulting to cryden's own +5 and 15 minutes, with the values written down here for the same reason +the rate-limit bounds are), `main.go` passes them, and a threshold +below 1 is a **startup error** rather than being read as "off" — +cryden has no way to switch lockout off, so accepting 0 would be +accepting a setting that means something else. + +This is called out in `README.md`, `.env.example` and the commit +message because it changes what every existing deployment does the next +time it restarts. It is the right direction — an account that can be +guessed at forever was not a design decision anyone made — but it is a +change nobody asked for, and burying it in a commit about a reporting +endpoint would have been the wrong way to ship it. + +### Verification: what this does NOT cover + +- **`migrations/012_digest_runs` has never been applied to a + database**, the same as `009`–`011`. Everything above is tested + through the in-memory doubles. +- **`digest.PostgresStore`'s `List` has not been run.** Its limit + clamp is asserted through the in-memory double and through + `ClampLimit` directly, which is the shared rule — but the SQL that + applies it is a copy of a design, not a verified query. The + TIMESTAMPTZ round trip in particular cannot be reproduced by a double + that stores `time.Time` as `time.Time`. +- **`memory.AuditStore` stamps `time.Now()` with no injectable clock**, + so window-*boundary* exclusion cannot be driven through the endpoint. + The digest and tuning tests assert the positive direction (events + recorded moments ago do appear inside a one-day window) and the text + the engine actually renders, rather than backdating an event. +- **The digest schedule is a goroutine on `context.Background()`.** The + scheduler takes a `context.Context` and is tested with a real + cancellable one, but `main.go` has nothing to cancel it with, because + this repo still has no graceful shutdown — the debt Stage 1 of Tier 3 + flagged, now with one more holder. +- **No live LLM call and no live database provider exist to test**, + because Stage 2 is not built. Nothing in Stage 1 touches + `ai.LLMProvider` or `ai.QueryableStore`. +- **`internal/smoketest` still has never been run** against a database, + unchanged from every previous tier's note. + +### Stage 2, and the two decisions it needs + +Stage 2 is the LLM provider config, the database provider config and +the ask-AI widget config. `NEXT.md` settles the shape of all three +(settings endpoints, this repo's own config table, pre-fill never +auto-apply, validate the read-only role by attempting a write). Two +things it does not settle, both of which change what gets built: + +1. **Whether this repo ships a live LLM client at all.** `ai.LLMProvider` + is an interface; implementing it against a real vendor means an + outbound HTTP client, a vendor choice, and a credential that leaves + the building. A console that configures a provider it cannot call is + not useful, so this is likely yes — but it is an integration + decision, not a wrapper decision, and this repo has so far shipped + no outbound integration of its own (the webhooks are cryden calling + a URL this repo hands it). +2. **Where the at-rest encryption key comes from.** `NEXT.md` requires + the stored provider credential be encrypted at rest and treated with + the same care as `JWT_SECRET`. `ENCRYPTION_KEY` already exists and + already encrypts TOTP secrets, so reusing it is the obvious + candidate — but reusing one key across two purposes is a decision + with a blast radius, and the alternative (a second key, or a KMS) + is a deployment change. + +Both were left for the user rather than guessed at. + +### Noticed while working, not fixed + +- **`openapi/spec.yaml` is now at 1.4 and covers Tiers 1–4 Stage 1**, + which closes the gap every previous entry flagged — `NEXT.md`'s Tier 1 + note that the spec "still predates Tier 1" is no longer true. The + document is large and hand-maintained, so it can drift again. +- **The unconfigured-store answer is still `404 not_configured`**, now + used by the digest history and the tuning endpoint too. Consistent + with every other unconfigured feature here, and still + indistinguishable from "this resource genuinely does not exist". +- **`config.Load` now refuses three knobs at startup** that it used to + accept silently (`LOG_LEVEL`, `LOCKOUT_THRESHOLD`, `DIGEST_INTERVAL_HOURS`). + That is the intended direction — a setting that silently does nothing + is worse than one that refuses to start — but it means an existing + deployment with a typo in one of them will fail to boot rather than + run with a default. From bf1abaae070692f0e5ae91fd2dc3e42fdaca1a75 Mon Sep 17 00:00:00 2001 From: Raymond Nicholas Date: Tue, 15 Sep 2026 23:54:16 +0100 Subject: [PATCH 06/12] feat: add the settings store and its at-rest encryption MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Credentials for the AI features are this repo's to store — cryden's ai.LLMProvider and ai.QueryableStore are interfaces it never implements — so migrations/013_settings holds them as AES-GCM ciphertext, sealed by cryden's own encryptor. The store deals in opaque bytes and has no helper that could write a plaintext credential. SETTINGS_ENCRYPTION_KEY is its own key, not a reuse of ENCRYPTION_KEY. Co-Authored-By: Claude Code --- .env.example | 23 +++++ config/config.go | 25 +++++ migrations/013_settings.down.sql | 3 + migrations/013_settings.up.sql | 36 +++++++ settings/memory.go | 85 +++++++++++++++ settings/postgres.go | 87 ++++++++++++++++ settings/secrets.go | 113 ++++++++++++++++++++ settings/secrets_test.go | 172 +++++++++++++++++++++++++++++++ settings/store.go | 75 ++++++++++++++ settings/store_test.go | 149 ++++++++++++++++++++++++++ 10 files changed, 768 insertions(+) create mode 100644 migrations/013_settings.down.sql create mode 100644 migrations/013_settings.up.sql create mode 100644 settings/memory.go create mode 100644 settings/postgres.go create mode 100644 settings/secrets.go create mode 100644 settings/secrets_test.go create mode 100644 settings/store.go create mode 100644 settings/store_test.go diff --git a/.env.example b/.env.example index ed3f057..dfd4c01 100644 --- a/.env.example +++ b/.env.example @@ -35,6 +35,29 @@ WEBAUTHN_RP_ID= WEBAUTHN_RP_DISPLAY_NAME= WEBAUTHN_RP_ORIGINS= +# Seals the credentials this api stores ITSELF, in the settings table: +# the LLM provider's API key and the read-only database's password behind +# the AI-assisted admin features. AES-256-GCM, the same encryptor cryden +# uses for TOTP secrets, with the key derived from this value — so treat +# it with the same care as JWT_SECRET. +# +# Deliberately a separate value from ENCRYPTION_KEY rather than a reuse +# of it. That one is cryden's, the engine derives from it whatever it +# needs to read TOTP secrets, and the two have different lifetimes: a +# rotation of either must not silently make the other's rows unreadable. +# The same reasoning is why CLOUD_LOG_HASH_KEY is its own value. +# +# Leave it unset to run without the AI settings screens; GET/PUT +# /v1/admin/settings/llm-provider and its database counterpart then answer +# 404 not_configured rather than the server refusing to start, the same +# shape ENCRYPTION_KEY itself uses for the second factors. +# +# If it ever changes, rows written under the old value become unreadable +# and say so (a decryption failure, distinct from "never configured") — +# re-enter those credentials. Clearing a setting still works without the +# key, so there is a way out that is not direct database access. +SETTINGS_ENCRYPTION_KEY= + # Login anomaly detection and credential-stuffing detection. Both are # report-only — a flagged attempt writes an audit event and nothing else, # no login is ever blocked — and both are off until ANOMALY_DETECTION is diff --git a/config/config.go b/config/config.go index b495f19..8b824a2 100644 --- a/config/config.go +++ b/config/config.go @@ -62,6 +62,25 @@ type Config struct { // JWT_SECRET. EncryptionKey string + // SettingsEncryptionKey seals the credentials this repo stores for + // the AI-assisted admin features — an LLM provider's API key, a + // read-only database's password — at rest, in the settings table. + // + // Deliberately NOT EncryptionKey, and this is the one place in this + // repo where a second key is spent rather than reused. EncryptionKey + // is cryden's: the engine derives from it whatever it needs to read + // TOTP secrets, and this repo never sees those bytes. This key is + // this repo's own, and the two have different lifetimes and different + // blast radii — rotating one must not silently make the other's rows + // unreadable. Every other purpose-keyed secret here follows the same + // rule (see CLOUD_LOG_HASH_KEY). + // + // Empty means the AI settings endpoints answer 404 not_configured + // rather than the server refusing to start, the same shape + // ENCRYPTION_KEY itself uses for the second factors: a deployment + // that has never opened that screen should still run. + SettingsEncryptionKey string + // TOTPIssuerName is what the user's authenticator app shows next to // the account. Cosmetic. Empty means cryden's own default ("Cryden"). TOTPIssuerName string @@ -370,6 +389,12 @@ func Load() (Config, error) { } } + // This repo's own key, for the credentials it stores itself in the + // settings table. Read here rather than defaulted from EncryptionKey + // above, on purpose — see the field comment for why the two are + // separate secrets with separate lifetimes. + cfg.SettingsEncryptionKey = os.Getenv("SETTINGS_ENCRYPTION_KEY") + // Anomaly detection and credential-stuffing detection — one switch, // because they are one store. Both threshold sets begin as the // engine's defaults and every knob below only replaces the one it diff --git a/migrations/013_settings.down.sql b/migrations/013_settings.down.sql new file mode 100644 index 0000000..67a7a14 --- /dev/null +++ b/migrations/013_settings.down.sql @@ -0,0 +1,3 @@ +-- 013_settings.down.sql + +DROP TABLE IF EXISTS settings; diff --git a/migrations/013_settings.up.sql b/migrations/013_settings.up.sql new file mode 100644 index 0000000..767aa52 --- /dev/null +++ b/migrations/013_settings.up.sql @@ -0,0 +1,36 @@ +-- 013_settings.up.sql +-- +-- Runtime configuration this api layer owns, for the values an operator +-- changes through a settings screen rather than a redeploy. +-- +-- Why a table and not env vars: cryden's ai.LLMProvider and +-- ai.QueryableStore are Go interfaces a host implements in code, and +-- everything else in this repo is configured by environment variable at +-- startup. That works for a threshold. It does not work for an LLM API +-- key an operator rotates, or for pointing the AI query surface at a +-- second, read-only database — both are things a deployment wants to +-- change without a restart and without shell access to the box. +-- +-- value is BYTEA, not TEXT or JSONB, and that is the point of the table +-- rather than a storage detail. Every credential written here — an API +-- key, a database password — is sealed with AES-GCM before it reaches +-- this store (see settings/secrets.go), so the bytes the database holds +-- are ciphertext and nothing else. A TEXT column would invite a caller to +-- write a readable value "just this once"; a BYTEA column holding +-- authenticated ciphertext means a plaintext credential in this table is +-- not a mistake someone can make by choosing the wrong helper. +-- +-- There is no user_id and no created_by. This is deployment-wide +-- configuration, not per-user data, and the audit trail for changing it +-- is the engine's own audit table — the handlers behind these rows are +-- admin-only and every write is attributable to a token this api issued. + +CREATE TABLE settings ( + -- A short, stable name ("llm_provider", "database_provider"). Not a + -- surrogate id: the name IS the identity, and a lookup by anything + -- else would be a lookup by something a caller had to be told. + key TEXT PRIMARY KEY, + -- Authenticated ciphertext. Never plaintext — see above. + value BYTEA NOT NULL, + updated_at TIMESTAMPTZ NOT NULL DEFAULT now() +); diff --git a/settings/memory.go b/settings/memory.go new file mode 100644 index 0000000..74754a7 --- /dev/null +++ b/settings/memory.go @@ -0,0 +1,85 @@ +package settings + +import ( + "context" + "fmt" + "sort" + "sync" +) + +// MemoryStore is the in-process Store, for tests and for any embedding +// host that wants the settings wiring without a database behind it. +// +// It is a faithful double rather than a convenient one in the one place +// that matters here: the bytes are copied on the way in and on the way +// out, exactly as a driver round trip through BYTEA would leave them +// unrelated to the caller's slice. A double that stored the caller's +// slice directly would let a test mutate a "stored" value through the +// variable it wrote with, and then assert something the real store cannot +// reproduce. +type MemoryStore struct { + mu sync.RWMutex + rows map[string][]byte +} + +func NewMemoryStore() *MemoryStore { + return &MemoryStore{rows: make(map[string][]byte)} +} + +var _ Store = (*MemoryStore)(nil) + +func (s *MemoryStore) Get(_ context.Context, key string) ([]byte, error) { + s.mu.RLock() + defer s.mu.RUnlock() + + value, ok := s.rows[key] + if !ok { + return nil, fmt.Errorf("%w: %q", ErrNotFound, key) + } + return append([]byte(nil), value...), nil +} + +func (s *MemoryStore) Put(_ context.Context, key string, value []byte) error { + s.mu.Lock() + defer s.mu.Unlock() + + s.rows[key] = append([]byte(nil), value...) + return nil +} + +func (s *MemoryStore) Delete(_ context.Context, key string) error { + s.mu.Lock() + defer s.mu.Unlock() + + if _, ok := s.rows[key]; !ok { + return fmt.Errorf("%w: %q", ErrNotFound, key) + } + delete(s.rows, key) + return nil +} + +func (s *MemoryStore) Keys(_ context.Context) ([]string, error) { + s.mu.RLock() + defer s.mu.RUnlock() + + keys := make([]string, 0, len(s.rows)) + for k := range s.rows { + keys = append(keys, k) + } + sort.Strings(keys) + return keys, nil +} + +// Raw is a test-only accessor the PostgresStore has no equivalent for: +// it returns what is actually in the store, so a test can assert that a +// credential went in as ciphertext rather than taking Secrets' word for +// it. Deliberately not part of Store, which is the interface production +// code gets — the point of that interface is that nothing above it can +// read the stored bytes without decrypting them first. +func (s *MemoryStore) Raw(key string) ([]byte, bool) { + s.mu.RLock() + defer s.mu.RUnlock() + + value, ok := s.rows[key] + return append([]byte(nil), value...), ok +} diff --git a/settings/postgres.go b/settings/postgres.go new file mode 100644 index 0000000..b94cb22 --- /dev/null +++ b/settings/postgres.go @@ -0,0 +1,87 @@ +package settings + +import ( + "context" + "database/sql" + "errors" + "fmt" +) + +// PostgresStore is the real store, over migrations/013_settings.up.sql. +// Constructed once in main.go with the same *sql.DB every other store in +// this repo gets. +type PostgresStore struct { + db *sql.DB +} + +func NewStore(db *sql.DB) *PostgresStore { + return &PostgresStore{db: db} +} + +var _ Store = (*PostgresStore)(nil) + +func (s *PostgresStore) Get(ctx context.Context, key string) ([]byte, error) { + var value []byte + err := s.db.QueryRowContext(ctx, + `SELECT value FROM settings WHERE key = $1`, key).Scan(&value) + if errors.Is(err, sql.ErrNoRows) { + return nil, fmt.Errorf("%w: %q", ErrNotFound, key) + } + if err != nil { + return nil, err + } + return value, nil +} + +func (s *PostgresStore) Put(ctx context.Context, key string, value []byte) error { + // []byte against a BYTEA column is the right pairing, unlike the + // JSONB columns elsewhere in this repo where a []byte param would be + // sent as bytea hex and rejected — there, params go as string(raw). + // Here the column IS bytea, so the driver's default encoding is the + // one the column wants. + // + // The upsert is one statement rather than a SELECT-then-INSERT, so + // two operators saving the same setting at the same moment cannot + // lose one of the writes. + _, err := s.db.ExecContext(ctx, ` + INSERT INTO settings (key, value) VALUES ($1, $2) + ON CONFLICT (key) DO UPDATE SET value = EXCLUDED.value, updated_at = now() + `, key, value) + return err +} + +func (s *PostgresStore) Delete(ctx context.Context, key string) error { + res, err := s.db.ExecContext(ctx, `DELETE FROM settings WHERE key = $1`, key) + if err != nil { + return err + } + // Reported from the statement's own row count rather than a prior + // SELECT, so "was there one" and "delete it" cannot disagree because + // another operator deleted it in between. + n, err := res.RowsAffected() + if err != nil { + return err + } + if n == 0 { + return fmt.Errorf("%w: %q", ErrNotFound, key) + } + return nil +} + +func (s *PostgresStore) Keys(ctx context.Context) ([]string, error) { + rows, err := s.db.QueryContext(ctx, `SELECT key FROM settings ORDER BY key`) + if err != nil { + return nil, err + } + defer rows.Close() + + keys := make([]string, 0, 8) + for rows.Next() { + var key string + if err := rows.Scan(&key); err != nil { + return nil, err + } + keys = append(keys, key) + } + return keys, rows.Err() +} diff --git a/settings/secrets.go b/settings/secrets.go new file mode 100644 index 0000000..566d6ca --- /dev/null +++ b/settings/secrets.go @@ -0,0 +1,113 @@ +package settings + +import ( + "context" + "errors" + "fmt" + + "github.com/crydensync/cryden/v2/security" +) + +// ErrNoEncryptionKey means Secrets was built without a key. Every method +// on Secrets returns it rather than storing anything, because the +// alternative — writing a credential through unencrypted — is the one +// outcome this package exists to prevent, and a "just this once" path +// would be found by exactly the caller least likely to think about it. +var ErrNoEncryptionKey = errors.New("settings: an encryption key is required to store credentials") + +// ErrUndecryptable means a value is stored under that key but the +// configured encryption key cannot open it. +var ErrUndecryptable = errors.New("settings: stored value could not be decrypted") + +// Secrets is the only thing a handler is given. It wraps a Store and +// seals every value on the way in and opens every value on the way out, +// so no caller has to remember to encrypt and no caller can forget to. +// +// The encryption itself is cryden's, not this repo's: +// security.NewAESGCMEncryptor is the same AES-256-GCM encryptor the +// engine already uses for TOTP secrets. Reimplementing it here would mean +// two cryptographic implementations to keep correct and one more place +// for a nonce to be reused, and there is nothing about a provider API key +// that needs different treatment from a TOTP secret. What this repo adds +// is only the storage and the key it derives from. +type Secrets struct { + store Store + // encryptor is nil when no key is configured. Held as cryden's + // interface rather than its concrete type so the nil case is one + // check here and not a type assertion at each call. + encryptor security.Encryptor +} + +// NewSecrets returns a Secrets over store, sealing with an AES-256-GCM +// key derived from key. +// +// An empty key is not an error at construction — it produces a Secrets +// that refuses every read and write with ErrNoEncryptionKey. That is +// deliberate: this repo's convention is that an optional feature which +// isn't configured answers 404 rather than stopping the server from +// starting (see the ENCRYPTION_KEY gate in main.go), and a deployment +// that has never opened the AI settings screen should not fail to boot +// because it has not set a key for a feature it isn't using. +func NewSecrets(store Store, key string) (*Secrets, error) { + if key == "" { + return &Secrets{store: store}, nil + } + encryptor, err := security.NewAESGCMEncryptor(key) + if err != nil { + return nil, fmt.Errorf("settings: building the encryptor: %w", err) + } + return &Secrets{store: store, encryptor: encryptor}, nil +} + +// Configured reports whether a key is set. Handlers use it to answer 404 +// not_configured, the same shape every other unconfigured feature in this +// api uses, rather than surfacing ErrNoEncryptionKey as a server fault. +func (s *Secrets) Configured() bool { + return s != nil && s.encryptor != nil +} + +// Put seals plaintext and stores it under key. +func (s *Secrets) Put(ctx context.Context, key string, plaintext []byte) error { + if !s.Configured() { + return ErrNoEncryptionKey + } + sealed, err := s.encryptor.Encrypt(string(plaintext)) + if err != nil { + return fmt.Errorf("settings: sealing %q: %w", key, err) + } + return s.store.Put(ctx, key, []byte(sealed)) +} + +// Get returns the decrypted value stored under key, or ErrNotFound. +// +// A decryption failure is reported as its own error rather than as +// ErrNotFound, and that distinction is load-bearing: "nothing is stored" +// and "something is stored that this key cannot open" call for different +// answers from an operator. The second almost always means the +// encryption key changed, and telling them the setting is missing would +// send them to re-enter a credential that is still there. +func (s *Secrets) Get(ctx context.Context, key string) ([]byte, error) { + if !s.Configured() { + return nil, ErrNoEncryptionKey + } + sealed, err := s.store.Get(ctx, key) + if err != nil { + return nil, err + } + plaintext, err := s.encryptor.Decrypt(string(sealed)) + if err != nil { + return nil, fmt.Errorf("%w: %q (was the encryption key changed?)", ErrUndecryptable, key) + } + return []byte(plaintext), nil +} + +// Delete removes key through the store. It needs no key of its own: a +// deployment that has lost its encryption key must still be able to clear +// the row it can no longer read, or the only way out is direct database +// access. +func (s *Secrets) Delete(ctx context.Context, key string) error { + if s == nil || s.store == nil { + return ErrNoEncryptionKey + } + return s.store.Delete(ctx, key) +} diff --git a/settings/secrets_test.go b/settings/secrets_test.go new file mode 100644 index 0000000..8ab636c --- /dev/null +++ b/settings/secrets_test.go @@ -0,0 +1,172 @@ +package settings + +import ( + "context" + "errors" + "strings" + "testing" +) + +const testKey = "a-test-encryption-key-long-enough-to-be-plausible" + +func newTestSecrets(t *testing.T, key string) (*Secrets, *MemoryStore) { + t.Helper() + store := NewMemoryStore() + secrets, err := NewSecrets(store, key) + if err != nil { + t.Fatalf("NewSecrets: %v", err) + } + return secrets, store +} + +func TestSecretsRoundTripsACredential(t *testing.T) { + ctx := context.Background() + secrets, _ := newTestSecrets(t, testKey) + + const credential = "sk-ant-api03-this-is-the-api-key" + if err := secrets.Put(ctx, KeyLLMProvider, []byte(credential)); err != nil { + t.Fatalf("Put: %v", err) + } + got, err := secrets.Get(ctx, KeyLLMProvider) + if err != nil { + t.Fatalf("Get: %v", err) + } + if string(got) != credential { + t.Errorf("Get = %q, want the credential back", got) + } +} + +// The whole reason this type exists, asserted against what is actually in +// the store rather than against Secrets' own answer: the bytes that land +// in the row must not contain the credential. +func TestSecretsNeverStoresThePlaintext(t *testing.T) { + ctx := context.Background() + secrets, store := newTestSecrets(t, testKey) + + const credential = "sk-ant-api03-this-is-the-api-key" + if err := secrets.Put(ctx, KeyLLMProvider, []byte(credential)); err != nil { + t.Fatalf("Put: %v", err) + } + + raw, ok := store.Raw(KeyLLMProvider) + if !ok { + t.Fatal("nothing was stored") + } + if strings.Contains(string(raw), "sk-ant") { + t.Errorf("the stored bytes contain the credential: %q", raw) + } + if strings.Contains(string(raw), credential) { + t.Errorf("the stored bytes are the credential verbatim: %q", raw) + } +} + +// Two writes of the same credential must not produce the same bytes. A +// fresh nonce per seal is what makes that true, and it is the property a +// store-and-compare attacker needs to be denied. +func TestSecretsSealsWithAFreshNonce(t *testing.T) { + ctx := context.Background() + secrets, store := newTestSecrets(t, testKey) + + const credential = "same-credential-both-times" + if err := secrets.Put(ctx, KeyLLMProvider, []byte(credential)); err != nil { + t.Fatalf("first Put: %v", err) + } + first, _ := store.Raw(KeyLLMProvider) + + if err := secrets.Put(ctx, KeyLLMProvider, []byte(credential)); err != nil { + t.Fatalf("second Put: %v", err) + } + second, _ := store.Raw(KeyLLMProvider) + + if string(first) == string(second) { + t.Error("sealing the same credential twice produced identical bytes, so the nonce is not fresh") + } +} + +// A changed key is a different failure from a missing setting, and an +// operator sent to re-enter a credential that is still on disk would be +// sent the wrong way. This asserts the two are distinguishable. +func TestSecretsSeparatesAChangedKeyFromAMissingSetting(t *testing.T) { + ctx := context.Background() + secrets, store := newTestSecrets(t, testKey) + + if _, err := secrets.Get(ctx, KeyLLMProvider); !errors.Is(err, ErrNotFound) { + t.Errorf("Get with nothing stored = %v, want ErrNotFound", err) + } + + if err := secrets.Put(ctx, KeyLLMProvider, []byte("a-credential")); err != nil { + t.Fatalf("Put: %v", err) + } + + // The same rows, a different key — a redeploy with + // SETTINGS_ENCRYPTION_KEY rotated and the table left alone. + rotated, err := NewSecrets(store, "a-completely-different-encryption-key") + if err != nil { + t.Fatalf("NewSecrets with a rotated key: %v", err) + } + if _, err := rotated.Get(ctx, KeyLLMProvider); !errors.Is(err, ErrUndecryptable) { + t.Errorf("Get with a rotated key = %v, want ErrUndecryptable", err) + } +} + +// Without a key every path that touches a credential refuses. This is the +// state a deployment that has never opened the AI settings screen is in, +// and the failure has to be a refusal rather than a write. +func TestSecretsWithoutAKeyRefusesEverything(t *testing.T) { + ctx := context.Background() + secrets, store := newTestSecrets(t, "") + + if secrets.Configured() { + t.Error("Configured = true with no key, so a handler would offer a screen that cannot save") + } + if err := secrets.Put(ctx, KeyLLMProvider, []byte("a-credential")); !errors.Is(err, ErrNoEncryptionKey) { + t.Errorf("Put with no key = %v, want ErrNoEncryptionKey", err) + } + if _, err := secrets.Get(ctx, KeyLLMProvider); !errors.Is(err, ErrNoEncryptionKey) { + t.Errorf("Get with no key = %v, want ErrNoEncryptionKey", err) + } + if _, ok := store.Raw(KeyLLMProvider); ok { + t.Error("Put with no key stored something anyway, which is the one outcome this must prevent") + } +} + +// Clearing a setting must not need the key that sealed it. A deployment +// that has lost SETTINGS_ENCRYPTION_KEY would otherwise have no way to +// remove the row it can no longer read short of direct database access. +func TestSecretsDeletesWithoutAKey(t *testing.T) { + ctx := context.Background() + secrets, store := newTestSecrets(t, testKey) + + if err := secrets.Put(ctx, KeyLLMProvider, []byte("a-credential")); err != nil { + t.Fatalf("Put: %v", err) + } + + noKey, err := NewSecrets(store, "") + if err != nil { + t.Fatalf("NewSecrets with no key: %v", err) + } + if err := noKey.Delete(ctx, KeyLLMProvider); err != nil { + t.Fatalf("Delete with no key: %v", err) + } + if _, ok := store.Raw(KeyLLMProvider); ok { + t.Error("the setting is still stored after Delete") + } +} + +// A nil Secrets is the shape a router built without the store has. Every +// method must answer rather than panic, because "not configured" is a +// wiring fact a handler reports, not a crash. +func TestNilSecretsIsUsable(t *testing.T) { + ctx := context.Background() + var secrets *Secrets + + if secrets.Configured() { + t.Error("a nil Secrets reports itself configured") + } + if _, err := secrets.Get(ctx, KeyLLMProvider); !errors.Is(err, ErrNoEncryptionKey) { + t.Errorf("Get on a nil Secrets = %v, want ErrNoEncryptionKey", err) + } + if err := secrets.Delete(ctx, KeyLLMProvider); !errors.Is(err, ErrNoEncryptionKey) { + t.Errorf("Delete on a nil Secrets = %v, want ErrNoEncryptionKey", err) + } +} diff --git a/settings/store.go b/settings/store.go new file mode 100644 index 0000000..3c4b736 --- /dev/null +++ b/settings/store.go @@ -0,0 +1,75 @@ +// Package settings stores the runtime configuration an operator changes +// through a settings screen, as opposed to the environment variables +// everything else in this repo is configured by at startup. +// +// What lives here is exactly the configuration cryden cannot own, because +// cryden has no settings concept at all: ai.LLMProvider and +// ai.QueryableStore are Go interfaces a host implements in code, so the +// credential behind each one is the host's to store. The table is +// migrations/013_settings.up.sql. +// +// The one rule this package exists to make unavoidable: **a credential +// never reaches the store as plaintext.** The store below deals in opaque +// bytes and knows nothing about what they mean; encryption happens in +// Secrets, which is the only thing a handler is given. A caller who +// reaches for the Store directly is writing ciphertext-shaped bytes or +// nothing, because the Store has no other interface to offer — see the +// migration's own comment on why the column is BYTEA. +package settings + +import ( + "context" + "errors" +) + +// The names this package stores under. Constants rather than string +// literals at each call site, so a typo is a compile error rather than a +// second, silently-empty setting nobody reads. +const ( + // KeyLLMProvider is the provider backing ai.LLMProvider: which model, + // and the API key to reach it with. + KeyLLMProvider = "llm_provider" + // KeyDatabaseProvider is the connection ai.QueryableStore runs + // against. It MUST be a read-only role — see + // httpapi/database_provider_handlers.go for the check that enforces + // it rather than trusting the form. + KeyDatabaseProvider = "database_provider" + // KeyAskAIWidget is the embed and scope configuration for the + // end-user ask-ai widget. Contains no credential: it is snippets and + // limits a console renders, so it is the one key here that is not + // secret. + KeyAskAIWidget = "ask_ai_widget" +) + +// ErrNotFound means nothing is stored under that key yet. Reported +// rather than returning empty bytes, because "never configured" and +// "configured as empty" are different answers and an endpoint that +// conflates them cannot tell an operator which one they are looking at. +var ErrNotFound = errors.New("settings: no such setting") + +// Store is the persistence. It deals in opaque bytes on purpose: nothing +// above it can store a readable credential through this interface without +// having encrypted it first, because this interface has no way to know +// the difference and no helper that would hide the step. +// +// Two implementations, the same reason every other store in this repo has +// two: the handlers have to be testable without a database. +type Store interface { + // Get returns the stored bytes for key, or ErrNotFound. The returned + // slice is a copy, so a caller mutating it does not reach into the + // store's own memory. + Get(ctx context.Context, key string) ([]byte, error) + + // Put creates or replaces key. One statement, so two operators + // saving at the same moment cannot lose one of the writes. + Put(ctx context.Context, key string, value []byte) error + + // Delete removes key, or returns ErrNotFound. Deleting an absent + // setting is reported rather than ignored, so a console that cleared + // the wrong row says so instead of showing success. + Delete(ctx context.Context, key string) error + + // Keys lists every stored key, sorted. Names only — never values, + // which is what makes it safe to answer a diagnostic question with. + Keys(ctx context.Context) ([]string, error) +} diff --git a/settings/store_test.go b/settings/store_test.go new file mode 100644 index 0000000..8cdf7b0 --- /dev/null +++ b/settings/store_test.go @@ -0,0 +1,149 @@ +package settings + +import ( + "context" + "errors" + "strings" + "testing" +) + +func newTestStore() *MemoryStore { return NewMemoryStore() } + +// The store's contract is small, and both implementations have to agree +// on it — the PostgresStore is what production runs and the MemoryStore is +// what every handler test asserts against, so a divergence between them +// would make the tests pass for a store production does not use. +// +// What is asserted here is the rule both share. The PostgresStore's own +// SQL is not exercised: there is no Postgres in this sandbox, so its +// statements are a copy of a design rather than a verified query (see +// PROGRESS.md). +func TestMemoryStoreRoundTripsBytes(t *testing.T) { + ctx := context.Background() + store := newTestStore() + + // Ciphertext, not text: arbitrary bytes including NULs, because that + // is what an AES-GCM seal produces and a store that mangled them + // would still pass a test written with ASCII. + value := []byte{0x00, 0xff, 0x10, 0x00, 'h', 'i', 0x7f} + + if err := store.Put(ctx, KeyLLMProvider, value); err != nil { + t.Fatalf("Put: %v", err) + } + got, err := store.Get(ctx, KeyLLMProvider) + if err != nil { + t.Fatalf("Get: %v", err) + } + if string(got) != string(value) { + t.Errorf("Get = %v, want %v", got, value) + } +} + +func TestMemoryStoreReportsAMissingKey(t *testing.T) { + ctx := context.Background() + store := newTestStore() + + if _, err := store.Get(ctx, KeyLLMProvider); !errors.Is(err, ErrNotFound) { + t.Errorf("Get on an empty store = %v, want ErrNotFound", err) + } + if err := store.Delete(ctx, KeyLLMProvider); !errors.Is(err, ErrNotFound) { + t.Errorf("Delete on an empty store = %v, want ErrNotFound", err) + } +} + +func TestMemoryStoreReplacesRatherThanAppends(t *testing.T) { + ctx := context.Background() + store := newTestStore() + + if err := store.Put(ctx, KeyLLMProvider, []byte("first")); err != nil { + t.Fatalf("first Put: %v", err) + } + if err := store.Put(ctx, KeyLLMProvider, []byte("second")); err != nil { + t.Fatalf("second Put: %v", err) + } + + got, err := store.Get(ctx, KeyLLMProvider) + if err != nil { + t.Fatalf("Get: %v", err) + } + if string(got) != "second" { + t.Errorf("Get = %q, want %q — the second write replaced the first", got, "second") + } + if keys, _ := store.Keys(ctx); len(keys) != 1 { + t.Errorf("Keys = %v, want one entry after two writes to the same key", keys) + } +} + +// The store must not hand out its own memory. A caller that kept the +// slice it wrote with, and then mutated it, would otherwise be editing +// what the store holds — which for a double means a handler test could +// assert a value production cannot produce. +func TestMemoryStoreCopiesOnTheWayInAndOut(t *testing.T) { + ctx := context.Background() + store := newTestStore() + + written := []byte("original") + if err := store.Put(ctx, KeyLLMProvider, written); err != nil { + t.Fatalf("Put: %v", err) + } + written[0] = 'X' // the caller's slice, after the write + + first, err := store.Get(ctx, KeyLLMProvider) + if err != nil { + t.Fatalf("Get: %v", err) + } + if string(first) != "original" { + t.Errorf("Get = %q after the caller mutated its own slice, want %q", first, "original") + } + + first[0] = 'Y' // the returned slice, after the read + second, err := store.Get(ctx, KeyLLMProvider) + if err != nil { + t.Fatalf("second Get: %v", err) + } + if string(second) != "original" { + t.Errorf("Get = %q after the caller mutated the returned slice, want %q", second, "original") + } +} + +func TestMemoryStoreListsKeysSortedAndWithoutValues(t *testing.T) { + ctx := context.Background() + store := newTestStore() + + for _, key := range []string{KeyLLMProvider, KeyAskAIWidget, KeyDatabaseProvider} { + if err := store.Put(ctx, key, []byte("secret-value")); err != nil { + t.Fatalf("Put(%s): %v", key, err) + } + } + + keys, err := store.Keys(ctx) + if err != nil { + t.Fatalf("Keys: %v", err) + } + want := []string{KeyAskAIWidget, KeyDatabaseProvider, KeyLLMProvider} + if strings.Join(keys, ",") != strings.Join(want, ",") { + t.Errorf("Keys = %v, want %v (sorted)", keys, want) + } + // Keys is names only. It is the one read an endpoint can make safely, + // and that only holds while it cannot carry a value. + for _, key := range keys { + if strings.Contains(key, "secret") { + t.Errorf("Keys returned %q, which looks like a value rather than a name", key) + } + } +} + +func TestMemoryStoreDeleteRemovesIt(t *testing.T) { + ctx := context.Background() + store := newTestStore() + + if err := store.Put(ctx, KeyLLMProvider, []byte("x")); err != nil { + t.Fatalf("Put: %v", err) + } + if err := store.Delete(ctx, KeyLLMProvider); err != nil { + t.Fatalf("Delete: %v", err) + } + if _, err := store.Get(ctx, KeyLLMProvider); !errors.Is(err, ErrNotFound) { + t.Errorf("Get after Delete = %v, want ErrNotFound", err) + } +} From bc7de0eae00e43154f42de15025906e18fcadfee Mon Sep 17 00:00:00 2001 From: Raymond Nicholas Date: Wed, 16 Sep 2026 00:00:23 +0100 Subject: [PATCH 07/12] feat: build a live ai.LLMProvider over the Anthropic Messages API MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit cryden ships no provider on purpose, so this repo brings one: the answer is constrained to a JSON schema whose enums are built from cryden's own allowlists rather than restated, so a model that asked for password_hash could not express it. cryden still validates every intent — this is the second lock, not the first. Tested against a local fake in the Messages API's wire shape, not the live service. Co-Authored-By: Claude Code --- aiprovider/anthropic.go | 287 ++++++++++++++++++++++++++++++ aiprovider/anthropic_test.go | 331 +++++++++++++++++++++++++++++++++++ aiprovider/schema.go | 27 +++ go.mod | 12 ++ go.sum | 30 ++++ settings/llm.go | 136 ++++++++++++++ 6 files changed, 823 insertions(+) create mode 100644 aiprovider/anthropic.go create mode 100644 aiprovider/anthropic_test.go create mode 100644 aiprovider/schema.go create mode 100644 settings/llm.go diff --git a/aiprovider/anthropic.go b/aiprovider/anthropic.go new file mode 100644 index 0000000..b66edfa --- /dev/null +++ b/aiprovider/anthropic.go @@ -0,0 +1,287 @@ +// Package aiprovider holds this repo's implementations of the interfaces +// cryden's ai and widget packages define. +// +// cryden ships none on purpose: ai.LLMProvider and ai.QueryableStore are +// shaped so a host brings its own vendor and its own database connection +// (see ai/types.go — "Ships zero implementations here — the consumer +// brings its own provider and API key, the same pattern as +// notify.EmailSender and logger.Logger"). This is that consumer. +// +// Nothing in this package is the safety boundary. cryden validates every +// QueryIntent against its allowlist before any query runs (ai.validateIntent) +// and widget.Ask force-scopes every intent to the calling user before +// that. What happens here is narrower and easier to state: turn a +// question into a candidate intent, and run an already-validated intent +// against a connection that has been checked to be read-only. +package aiprovider + +import ( + "context" + "encoding/json" + "errors" + "fmt" + "strings" + + "github.com/anthropics/anthropic-sdk-go" + "github.com/anthropics/anthropic-sdk-go/option" + + crydenai "github.com/crydensync/cryden/v2/ai" +) + +// ErrNoAPIKey is returned when the provider was built without a +// credential. Constructing one anyway and failing at the first call would +// move the failure from startup to the middle of an admin's question. +var ErrNoAPIKey = errors.New("aiprovider: an API key is required") + +// ErrUnexpectedAnswer means the model returned something that is not a +// QueryIntent. It is a real possibility even with the output schema +// enforced — a refusal, a truncated response — and it is reported as its +// own error rather than as a validation failure, because "the model did +// not answer the question" and "the model answered with something unsafe" +// call for different words in front of an operator. +var ErrUnexpectedAnswer = errors.New("aiprovider: the model did not return a query intent") + +// Anthropic implements cryden's ai.LLMProvider (and widget.Composer) +// against the Anthropic Messages API. +// +// It is deliberately thin. The interesting work — deciding whether a +// model's answer is safe to run — belongs to cryden, and duplicating any +// of it here would create a second place for the rules to drift. +type Anthropic struct { + client anthropic.Client + model string + // maxTokens bounds one response. The answer is a small JSON object, + // so this is a ceiling on cost rather than on usefulness. + maxTokens int +} + +// AnthropicConfig is what the settings table stores, unpacked into what +// this constructor needs. +type AnthropicConfig struct { + APIKey string + Model string + MaxTokens int +} + +// NewAnthropic builds a provider from a stored configuration. +func NewAnthropic(cfg AnthropicConfig, opts ...option.RequestOption) (*Anthropic, error) { + if strings.TrimSpace(cfg.APIKey) == "" { + return nil, ErrNoAPIKey + } + opts = append(opts, option.WithAPIKey(cfg.APIKey)) + + return &Anthropic{ + client: anthropic.NewClient(opts...), + model: cfg.Model, + maxTokens: cfg.MaxTokens, + }, nil +} + +var ( + _ crydenai.LLMProvider = (*Anthropic)(nil) +) + +// intentSchema is the JSON schema the model's answer is constrained to. +// +// This is the second lock on a door cryden already bolts. The enums below +// are built from cryden's own allowlists rather than restated, so a model +// that has been talked into asking for the password hash cannot even +// express it: the field is not in the schema, and the API enforces the +// schema, not the prompt. cryden would reject the intent anyway — this +// just means the rejection almost never has to happen. +// +// Built from cryden's maps rather than hardcoded on purpose. If the +// engine allowlists a new field tomorrow, this schema follows it with no +// edit here; if the engine ever *removes* one, a hardcoded copy would +// keep offering it. +func intentSchema() map[string]any { + entities := sortedKeys(crydenai.AllowedEntities) + + // group_by is a single string field, and the schema is flat — it has + // no way to say "this enum depends on the entity you chose". So the + // enum it offers is the union across every entity's allowed fields, + // which is the honest superset: cryden checks group_by against the + // chosen entity's own list and rejects a mismatch. Narrowing here + // instead would mean duplicating cryden's per-entity rule in a form + // JSON Schema cannot express, and a stale copy of it would silently + // refuse a field the engine would have accepted. + groupable := map[string]bool{} + for _, fields := range crydenai.AllowedFields { + for field := range fields { + groupable[field] = true + } + } + + return map[string]any{ + "type": "object", + "properties": map[string]any{ + "entity": enumOf(entities), + "filters": map[string]any{ + "type": "array", + "items": map[string]any{ + "type": "object", + "properties": map[string]any{ + "field": map[string]any{"type": "string"}, + "operator": enumOf(sortedKeys(crydenai.AllowedOperators)), + "value": map[string]any{"type": "string"}, + }, + "required": []string{"field", "operator", "value"}, + "additionalProperties": false, + }, + }, + "aggregate": enumOf([]string{"", "count", "group_by"}), + "group_by": enumOf(sortedKeys(groupable)), + "limit": map[string]any{"type": "integer"}, + }, + "required": []string{"entity", "filters", "aggregate", "limit"}, + "additionalProperties": false, + } +} + +// systemPrompt is the whole of this repo's prompt. It is short because +// the schema does the constraining: the prompt's job is to say what the +// fields mean, not to enumerate what is allowed, and a prompt that listed +// the allowlist would be one more copy of it to keep in sync. +const systemPrompt = `You translate an administrator's question about their user database into a structured query intent. + +The intent names one entity, any filters narrowing it, and how to present the result. Use "count" when the question asks how many, "group_by" when it asks for a breakdown, and the empty aggregate when it asks for the rows themselves. + +Only the fields the schema offers exist. If the question cannot be expressed with them, choose the closest entity and omit the filters you cannot express rather than inventing a field name.` + +// ParseQueryIntent asks the model for a QueryIntent. The returned intent +// is unvalidated: cryden's ai.ExecuteIntent checks it against the +// allowlist before anything runs, and this function must not be assumed +// to have done so. +func (p *Anthropic) ParseQueryIntent(ctx context.Context, naturalLanguage string) (crydenai.QueryIntent, error) { + response, err := p.client.Messages.New(ctx, anthropic.MessageNewParams{ + Model: anthropic.Model(p.model), + MaxTokens: int64(p.maxTokens), + System: []anthropic.TextBlockParam{{ + Text: systemPrompt, + // The prompt and the schema are fixed for the life of the + // process, so every request after the first reads this from + // cache instead of paying for it again. + CacheControl: anthropic.NewCacheControlEphemeralParam(), + }}, + OutputConfig: anthropic.OutputConfigParam{ + Format: anthropic.JSONOutputFormatParam{Schema: intentSchema()}, + }, + Messages: []anthropic.MessageParam{ + anthropic.NewUserMessage(anthropic.NewTextBlock(naturalLanguage)), + }, + }) + if err != nil { + return crydenai.QueryIntent{}, fmt.Errorf("aiprovider: asking the model: %w", err) + } + + // Checked before the content is read, because a refusal carries no + // usable text and reading it first would report a refusal as a parse + // failure — sending an operator to look at their schema when the + // model simply declined the question. + if response.StopReason == anthropic.StopReasonRefusal { + return crydenai.QueryIntent{}, fmt.Errorf("%w: the model declined to answer (%s)", + ErrUnexpectedAnswer, response.StopDetails.Category) + } + + text := firstText(response.Content) + if text == "" { + return crydenai.QueryIntent{}, fmt.Errorf("%w: the response carried no text (stop reason %q)", + ErrUnexpectedAnswer, response.StopReason) + } + + var payload queryIntentPayload + if err := json.Unmarshal([]byte(text), &payload); err != nil { + return crydenai.QueryIntent{}, fmt.Errorf("%w: %v", ErrUnexpectedAnswer, err) + } + + intent := crydenai.QueryIntent{ + Entity: payload.Entity, + Aggregate: payload.Aggregate, + GroupBy: payload.GroupBy, + Limit: payload.Limit, + } + for _, f := range payload.Filters { + intent.Filters = append(intent.Filters, crydenai.QueryFilter{ + Field: f.Field, + Operator: f.Operator, + Value: f.Value, + }) + } + return intent, nil +} + +// queryIntentPayload mirrors cryden's ai.QueryIntent as JSON. A separate +// type rather than unmarshalling into ai.QueryIntent directly, because +// the wire shape and the engine's own struct are allowed to differ — +// QueryFilter is a struct here and an element of a slice there — and +// because a named type is where the JSON tags can be documented. +type queryIntentPayload struct { + Entity string `json:"entity"` + Filters []queryFilterPayload `json:"filters"` + Aggregate string `json:"aggregate"` + GroupBy string `json:"group_by"` + Limit int `json:"limit"` +} + +type queryFilterPayload struct { + Field string `json:"field"` + Operator string `json:"operator"` + Value string `json:"value"` +} + +// ComposeAnswer implements widget.Composer: it turns an already-validated, +// already-owner-scoped result into a sentence for an end user. +// +// The result it is given has been scoped to one identity by cryden's +// widget.Ask before it arrives — this function never sees another user's +// rows, and does not need to know that scoping exists. That is why it can +// be a plain presentation call with nothing to check. +func (p *Anthropic) ComposeAnswer(ctx context.Context, question string, result crydenai.QueryResult) (string, error) { + response, err := p.client.Messages.New(ctx, anthropic.MessageNewParams{ + Model: anthropic.Model(p.model), + MaxTokens: int64(p.maxTokens), + System: []anthropic.TextBlockParam{{ + Text: "You answer a user's question using only the rows provided. If the rows do not answer it, say so plainly. Never mention SQL, tables or column names.", + CacheControl: anthropic.NewCacheControlEphemeralParam(), + }}, + Messages: []anthropic.MessageParam{ + anthropic.NewUserMessage( + anthropic.NewTextBlock("Question: "+question), + anthropic.NewTextBlock("Rows:\n"+renderRows(result)), + ), + }, + }) + if err != nil { + return "", fmt.Errorf("aiprovider: composing an answer: %w", err) + } + if response.StopReason == anthropic.StopReasonRefusal { + return "", fmt.Errorf("%w: the model declined to answer (%s)", + ErrUnexpectedAnswer, response.StopDetails.Category) + } + return firstText(response.Content), nil +} + +// renderRows is a compact, positional rendering of a QueryResult for the +// composer. Headers once, then one line per row — the model is reading +// this, so a repeated header would be noise it might mistake for data. +func renderRows(result crydenai.QueryResult) string { + var b strings.Builder + b.WriteString(strings.Join(result.Columns, ", ")) + for _, row := range result.Rows { + b.WriteString("\n") + b.WriteString(strings.Join(row, ", ")) + } + return b.String() +} + +// firstText returns the first text block's content, or "". A response can +// carry thinking blocks before the text one, so this walks rather than +// indexing. +func firstText(blocks []anthropic.ContentBlockUnion) string { + for _, block := range blocks { + if text, ok := block.AsAny().(anthropic.TextBlock); ok { + return text.Text + } + } + return "" +} diff --git a/aiprovider/anthropic_test.go b/aiprovider/anthropic_test.go new file mode 100644 index 0000000..9c320d9 --- /dev/null +++ b/aiprovider/anthropic_test.go @@ -0,0 +1,331 @@ +package aiprovider + +import ( + "context" + "encoding/json" + "errors" + "fmt" + "io" + "net/http" + "net/http/httptest" + "sort" + "strings" + "testing" + + "github.com/anthropics/anthropic-sdk-go/option" + + crydenai "github.com/crydensync/cryden/v2/ai" +) + +// The provider is tested against a local HTTP server that answers in the +// Messages API's wire shape, rather than against the live API. That is +// not a shortcut around testing the interesting part: everything this +// package does — building the request, reading the response, turning it +// into a QueryIntent — happens on this side of the socket, and a fake +// server is the only way to drive the failure paths (a refusal, a +// truncated answer, an HTTP error) on purpose rather than by luck. +// +// What is NOT covered: the real API's behaviour. Whether the model +// answers well, and whether the output schema is accepted as written, is +// only knowable against the live service. PROGRESS.md says so. +type fakeAPI struct { + *httptest.Server + // lastBody is the decoded request body of the most recent call, so a + // test can assert what was actually sent rather than only what came + // back. + lastBody map[string]any + // calls counts requests, so a test can prove a refusal was not + // retried into a second charge. + calls int +} + +// newFakeAPI answers every request with the given assistant text. +func newFakeAPI(t *testing.T, answer string) *fakeAPI { + t.Helper() + f := &fakeAPI{} + f.Server = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + raw, _ := io.ReadAll(r.Body) + var body map[string]any + _ = json.Unmarshal(raw, &body) + f.lastBody = body + f.calls++ + + w.Header().Set("Content-Type", "application/json") + fmt.Fprintf(w, `{ + "id": "msg_test", "type": "message", "role": "assistant", + "model": "claude-opus-5", "stop_reason": "end_turn", + "content": [{"type": "text", "text": %s}], + "usage": {"input_tokens": 1, "output_tokens": 1} + }`, mustJSON(t, answer)) + })) + t.Cleanup(f.Close) + return f +} + +// newFakeAPIResponse answers with a complete response body, for the cases +// where the shape matters more than the text. +func newFakeAPIResponse(t *testing.T, body string) *fakeAPI { + t.Helper() + f := &fakeAPI{} + f.Server = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + raw, _ := io.ReadAll(r.Body) + var decoded map[string]any + _ = json.Unmarshal(raw, &decoded) + f.lastBody = decoded + f.calls++ + + w.Header().Set("Content-Type", "application/json") + _, _ = io.WriteString(w, body) + })) + t.Cleanup(f.Close) + return f +} + +func mustJSON(t *testing.T, v any) string { + t.Helper() + raw, err := json.Marshal(v) + if err != nil { + t.Fatalf("marshalling %v: %v", v, err) + } + return string(raw) +} + +func (f *fakeAPI) provider(t *testing.T) *Anthropic { + t.Helper() + p, err := NewAnthropic( + AnthropicConfig{APIKey: "test-key", Model: "claude-opus-5", MaxTokens: 1024}, + option.WithBaseURL(f.URL), + ) + if err != nil { + t.Fatalf("NewAnthropic: %v", err) + } + return p +} + +func TestNewAnthropicRefusesAnEmptyKey(t *testing.T) { + for _, key := range []string{"", " "} { + if _, err := NewAnthropic(AnthropicConfig{APIKey: key}); !errors.Is(err, ErrNoAPIKey) { + t.Errorf("NewAnthropic(APIKey %q) = %v, want ErrNoAPIKey", key, err) + } + } +} + +func TestParseQueryIntentReadsTheModelsAnswer(t *testing.T) { + api := newFakeAPI(t, `{"entity":"users","filters":[{"field":"email","operator":"=","value":"dana@example.com"}],"aggregate":"","group_by":"","limit":10}`) + + intent, err := api.provider(t).ParseQueryIntent(context.Background(), "show me dana") + if err != nil { + t.Fatalf("ParseQueryIntent: %v", err) + } + if intent.Entity != "users" { + t.Errorf("Entity = %q, want %q", intent.Entity, "users") + } + if len(intent.Filters) != 1 { + t.Fatalf("Filters = %+v, want exactly one", intent.Filters) + } + got := intent.Filters[0] + if got.Field != "email" || got.Operator != "=" || got.Value != "dana@example.com" { + t.Errorf("Filters[0] = %+v, want the email filter the model returned", got) + } + if intent.Limit != 10 { + t.Errorf("Limit = %d, want 10", intent.Limit) + } +} + +// The schema is the second lock on the door: a model that has been argued +// into asking for a password hash must not even be able to express it. +func TestIntentSchemaOffersOnlyWhatCrydenWouldAccept(t *testing.T) { + schema := intentSchema() + props, ok := schema["properties"].(map[string]any) + if !ok { + t.Fatalf("schema has no properties object: %+v", schema) + } + + entityEnum := enumValues(t, props["entity"]) + for entity := range crydenai.AllowedEntities { + if !contains(entityEnum, entity) { + t.Errorf("schema omits the allowlisted entity %q", entity) + } + } + if len(entityEnum) != len(crydenai.AllowedEntities) { + t.Errorf("entity enum = %v, want exactly cryden's allowlist %v", entityEnum, crydenai.AllowedEntities) + } + + // The columns that must never be reachable through this path. cryden + // leaves them out of AllowedFields; this asserts the schema does too, + // because the schema is what the model is physically able to emit. + groupBy := enumValues(t, props["group_by"]) + for _, forbidden := range []string{"password_hash", "token_hash", "PasswordHash", "TokenHash"} { + if contains(groupBy, forbidden) { + t.Errorf("group_by enum offers %q, which cryden deliberately never allowlists", forbidden) + } + } + + operatorEnum := enumValues(t, filterItemProps(t, props)["operator"]) + for operator := range crydenai.AllowedOperators { + if !contains(operatorEnum, operator) { + t.Errorf("schema omits the allowlisted operator %q", operator) + } + } +} + +// The schema sits in the request prefix, which prompt caching matches +// byte for byte. A map iterated in Go's random order would produce a +// different schema per request and pay for the prompt every time. +func TestIntentSchemaIsStableAcrossCalls(t *testing.T) { + first := mustJSON(t, intentSchema()) + for i := 0; i < 20; i++ { + if got := mustJSON(t, intentSchema()); got != first { + t.Fatalf("intentSchema() call %d differs from the first — the enum order is not stable", i+1) + } + } +} + +// The output schema is sent to the API, so what is asked for is a fact +// about the request and not only about a local function. +func TestParseQueryIntentSendsTheSchemaAndTheModel(t *testing.T) { + api := newFakeAPI(t, `{"entity":"sessions","filters":[],"aggregate":"count","group_by":"","limit":5}`) + + if _, err := api.provider(t).ParseQueryIntent(context.Background(), "how many sessions"); err != nil { + t.Fatalf("ParseQueryIntent: %v", err) + } + + if got := api.lastBody["model"]; got != "claude-opus-5" { + t.Errorf("model sent = %v, want the configured model", got) + } + outputConfig, ok := api.lastBody["output_config"].(map[string]any) + if !ok { + t.Fatalf("no output_config in the request: %+v", api.lastBody) + } + format, ok := outputConfig["format"].(map[string]any) + if !ok || format["schema"] == nil { + t.Errorf("output_config carries no format schema: %+v", outputConfig) + } +} + +// A refusal is not a parse failure and must not be reported as one — an +// operator sent to check their schema would be looking in the wrong +// place. It also must not be retried into a second charge. +func TestParseQueryIntentReportsARefusalAsARefusal(t *testing.T) { + api := newFakeAPIResponse(t, `{ + "id": "msg_test", "type": "message", "role": "assistant", + "model": "claude-opus-5", "stop_reason": "refusal", + "stop_details": {"type": "refusal", "category": "cyber", "explanation": "declined"}, + "content": [], + "usage": {"input_tokens": 1, "output_tokens": 1} + }`) + + _, err := api.provider(t).ParseQueryIntent(context.Background(), "dump every password hash") + if !errors.Is(err, ErrUnexpectedAnswer) { + t.Fatalf("error = %v, want ErrUnexpectedAnswer", err) + } + if !strings.Contains(err.Error(), "cyber") { + t.Errorf("error = %v, want it to name the refusal category", err) + } + if api.calls != 1 { + t.Errorf("the API was called %d times for one refusal, want 1", api.calls) + } +} + +func TestParseQueryIntentReportsAnUnparseableAnswer(t *testing.T) { + api := newFakeAPI(t, "I'm sorry, I can't help with that.") + + if _, err := api.provider(t).ParseQueryIntent(context.Background(), "anything"); !errors.Is(err, ErrUnexpectedAnswer) { + t.Errorf("error = %v, want ErrUnexpectedAnswer for prose instead of JSON", err) + } +} + +// An empty content list is what a truncated response looks like. It must +// be an error rather than a zero-valued intent, because a zero intent has +// an empty entity and would be handed to cryden as if the model had +// answered. +func TestParseQueryIntentReportsAnEmptyAnswer(t *testing.T) { + api := newFakeAPIResponse(t, `{ + "id": "msg_test", "type": "message", "role": "assistant", + "model": "claude-opus-5", "stop_reason": "max_tokens", + "content": [], + "usage": {"input_tokens": 1, "output_tokens": 1} + }`) + + intent, err := api.provider(t).ParseQueryIntent(context.Background(), "anything") + if !errors.Is(err, ErrUnexpectedAnswer) { + t.Fatalf("error = %v, want ErrUnexpectedAnswer", err) + } + if intent.Entity != "" { + t.Errorf("Entity = %q on a failed parse, want the zero value", intent.Entity) + } +} + +// A server error is reported as itself rather than as a model problem: it +// says nothing about the question and everything about the deployment. +func TestParseQueryIntentReportsATransportFailure(t *testing.T) { + api := &fakeAPI{} + api.Server = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { + http.Error(w, `{"type":"error","error":{"type":"api_error","message":"boom"}}`, http.StatusInternalServerError) + })) + t.Cleanup(api.Close) + + _, err := api.provider(t).ParseQueryIntent(context.Background(), "anything") + if err == nil { + t.Fatal("ParseQueryIntent against a 500 returned no error") + } + if errors.Is(err, ErrUnexpectedAnswer) { + t.Errorf("error = %v, want a transport failure rather than ErrUnexpectedAnswer", err) + } +} + +func TestComposeAnswerReturnsTheModelsProse(t *testing.T) { + api := newFakeAPI(t, "You signed in from three devices this week.") + + text, err := api.provider(t).ComposeAnswer(context.Background(), "where did I sign in from?", + crydenai.QueryResult{Columns: []string{"ip"}, Rows: [][]string{{"203.0.113.1"}}}) + if err != nil { + t.Fatalf("ComposeAnswer: %v", err) + } + if text != "You signed in from three devices this week." { + t.Errorf("ComposeAnswer = %q, want the model's text", text) + } +} + +// enumValues pulls the values out of a JSON Schema enum node. +func enumValues(t *testing.T, node any) []string { + t.Helper() + object, ok := node.(map[string]any) + if !ok { + t.Fatalf("expected an enum object, got %T (%v)", node, node) + } + raw, ok := object["enum"].([]string) + if !ok { + t.Fatalf("expected a string enum, got %T (%v)", object["enum"], object["enum"]) + } + out := append([]string(nil), raw...) + sort.Strings(out) + return out +} + +// filterItemProps reaches into the filters array's item schema. +func filterItemProps(t *testing.T, props map[string]any) map[string]any { + t.Helper() + filters, ok := props["filters"].(map[string]any) + if !ok { + t.Fatalf("schema has no filters object: %+v", props) + } + items, ok := filters["items"].(map[string]any) + if !ok { + t.Fatalf("filters has no items schema: %+v", filters) + } + itemProps, ok := items["properties"].(map[string]any) + if !ok { + t.Fatalf("filter items have no properties: %+v", items) + } + return itemProps +} + +func contains(haystack []string, needle string) bool { + for _, h := range haystack { + if h == needle { + return true + } + } + return false +} diff --git a/aiprovider/schema.go b/aiprovider/schema.go new file mode 100644 index 0000000..9b31bdb --- /dev/null +++ b/aiprovider/schema.go @@ -0,0 +1,27 @@ +package aiprovider + +import "sort" + +// sortedKeys returns the keys of a string-keyed set in a stable order. +// +// The order matters and is not cosmetic. These feed JSON Schema enums +// that sit inside the request's system-prompt prefix, and prompt caching +// is a prefix match — a map iterated in Go's random order would produce a +// different schema on every request, invalidating the cache every time +// and paying full price for a prompt that never changed. +func sortedKeys(set map[string]bool) []string { + keys := make([]string, 0, len(set)) + for key := range set { + keys = append(keys, key) + } + sort.Strings(keys) + return keys +} + +// enumOf wraps values as a JSON Schema enum of strings. +func enumOf(values []string) map[string]any { + return map[string]any{ + "type": "string", + "enum": values, + } +} diff --git a/go.mod b/go.mod index 8978bc9..6679b7c 100644 --- a/go.mod +++ b/go.mod @@ -3,6 +3,7 @@ module github.com/crydensync/api go 1.25.0 require ( + github.com/anthropics/anthropic-sdk-go v1.73.0 github.com/crydensync/cryden/v2 v2.5.0 github.com/golang-jwt/jwt/v5 v5.3.1 github.com/lib/pq v1.12.3 @@ -10,7 +11,9 @@ require ( ) require ( + github.com/bahlo/generic-list-go v0.2.0 // indirect github.com/boombuler/barcode v1.0.1-0.20190219062509-6c824513bacc // indirect + github.com/buger/jsonparser v1.1.2 // indirect github.com/cespare/xxhash/v2 v2.3.0 // indirect github.com/fxamacker/cbor/v2 v2.9.3 // indirect github.com/go-viper/mapstructure/v2 v2.5.0 // indirect @@ -18,11 +21,20 @@ require ( github.com/go-webauthn/x v0.3.0 // indirect github.com/google/go-tpm v0.9.8 // indirect github.com/google/uuid v1.6.0 // indirect + github.com/invopop/jsonschema v0.14.0 // indirect + github.com/pb33f/ordered-map/v2 v2.3.1 // indirect github.com/philhofer/fwd v1.2.0 // indirect github.com/pquerna/otp v1.5.0 // indirect + github.com/standard-webhooks/standard-webhooks/libraries v0.0.1 // indirect + github.com/tidwall/gjson v1.18.0 // indirect + github.com/tidwall/match v1.1.1 // indirect + github.com/tidwall/pretty v1.2.1 // indirect + github.com/tidwall/sjson v1.2.5 // indirect github.com/tinylib/msgp v1.6.4 // indirect github.com/x448/float16 v0.8.4 // indirect go.uber.org/atomic v1.11.0 // indirect + go.yaml.in/yaml/v4 v4.0.0-rc.2 // indirect golang.org/x/crypto v0.55.0 // indirect + golang.org/x/sync v0.16.0 // indirect golang.org/x/sys v0.47.0 // indirect ) diff --git a/go.sum b/go.sum index 14a2ddb..dc9b28e 100644 --- a/go.sum +++ b/go.sum @@ -1,9 +1,15 @@ +github.com/anthropics/anthropic-sdk-go v1.73.0 h1:brLt0oVV6LH0YcG6O6G3pBnIVDJQ/QPoXul+OJHXEFc= +github.com/anthropics/anthropic-sdk-go v1.73.0/go.mod h1:x+lPk/cCl48uRegeP0hlYYBN1b7bEBTveInIMgLicnY= +github.com/bahlo/generic-list-go v0.2.0 h1:5sz/EEAK+ls5wF+NeqDpk5+iNdMDXrh3z3nPnH1Wvgk= +github.com/bahlo/generic-list-go v0.2.0/go.mod h1:2KvAjgMlE5NNynlg/5iLrrCCZ2+5xWbdbCW3pNTGyYg= github.com/boombuler/barcode v1.0.1-0.20190219062509-6c824513bacc h1:biVzkmvwrH8WK8raXaxBx6fRVTlJILwEwQGL1I/ByEI= github.com/boombuler/barcode v1.0.1-0.20190219062509-6c824513bacc/go.mod h1:paBWMcWSl3LHKBqUq+rly7CNSldXjb2rDl3JlRe0mD8= github.com/bsm/ginkgo/v2 v2.12.0 h1:Ny8MWAHyOepLGlLKYmXG4IEkioBysk6GpaRTLC8zwWs= github.com/bsm/ginkgo/v2 v2.12.0/go.mod h1:SwYbGRRDovPVboqFv0tPTcG1sN61LM1Z4ARdbAV9g4c= github.com/bsm/gomega v1.27.10 h1:yeMWxP2pV2fG3FgAODIY8EiRE3dy0aeFYt4l7wh6yKA= github.com/bsm/gomega v1.27.10/go.mod h1:JyEr/xRbxbtgWNi8tIEVPUYZ5Dzef52k01W3YH0H+O0= +github.com/buger/jsonparser v1.1.2 h1:frqHqw7otoVbk5M8LlE/L7HTnIq2v9RX6EJ48i9AxJk= +github.com/buger/jsonparser v1.1.2/go.mod h1:6RYKKt7H4d4+iWqouImQ9R2FZql3VbhNgx27UK13J/0= github.com/cespare/xxhash/v2 v2.3.0 h1:UL815xU9SqsFlibzuggzjXhog7bL6oX9BbNZnL2UFvs= github.com/cespare/xxhash/v2 v2.3.0/go.mod h1:VGX0DQ3Q6kWi7AoAeZDth3/j3BFtOZR5XLFGgcrjCOs= github.com/crydensync/cryden/v2 v2.5.0 h1:hIeLxF3SGnZgN7ZyuvTS6WsjmdwLoiRY8LqSunE09pY= @@ -11,6 +17,8 @@ github.com/crydensync/cryden/v2 v2.5.0/go.mod h1:SNrasjDCREkJoSjE++/+jXG1QpdCB6v github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38= github.com/descope/virtualwebauthn v1.0.5 h1:fMXji5UMepJC51Ge6d4v5IAjiJQRKmXE9hlo/B9SczQ= github.com/descope/virtualwebauthn v1.0.5/go.mod h1:lLCfN+DpCM3iisM4bCILZlFEWkC1Zo7ZgsxC45CUapI= +github.com/dnaeon/go-vcr v1.2.0 h1:zHCHvJYTMh1N7xnV7zf1m1GPBF9Ad0Jk/whtQ1663qI= +github.com/dnaeon/go-vcr v1.2.0/go.mod h1:R4UdLID7HZT3taECzJs4YgbbH6PIGXB6W/sc5OLb6RQ= github.com/fxamacker/cbor/v2 v2.9.3 h1:oQBnFATpNdY8gJHTndDDv5Xl4QqNaz51G5LLEPhng3Q= github.com/fxamacker/cbor/v2 v2.9.3/go.mod h1:vM4b+DJCtHn+zz7h3FFp/hDAI9WNWCsZj23V5ytsSxQ= github.com/go-viper/mapstructure/v2 v2.5.0 h1:vM5IJoUAy3d7zRSVtIwQgBj7BiWtMPfmPEgAXnvj1Ro= @@ -27,10 +35,14 @@ github.com/google/go-tpm-tools v0.3.13-0.20230620182252-4639ecce2aba h1:qJEJcuLz github.com/google/go-tpm-tools v0.3.13-0.20230620182252-4639ecce2aba/go.mod h1:EFYHy8/1y2KfgTAsx7Luu7NGhoxtuVHnNo8jE7FikKc= github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0= github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo= +github.com/invopop/jsonschema v0.14.0 h1:MHQqLhvpNUZfw+hM3AZDYK7jxO8FZoQeQM77g8iyZjg= +github.com/invopop/jsonschema v0.14.0/go.mod h1:ygm6C2EaVNMBDPpaPlnOA2pFAxBnxGjFlMZABxm9n2I= github.com/klauspost/cpuid/v2 v2.2.10 h1:tBs3QSyvjDyFTq3uoc/9xFpCuOsJQFNPiAhYdw2skhE= github.com/klauspost/cpuid/v2 v2.2.10/go.mod h1:hqwkgyIinND0mEev00jJYCxPNVRVXFQeu1XKlok6oO0= github.com/lib/pq v1.12.3 h1:tTWxr2YLKwIvK90ZXEw8GP7UFHtcbTtty8zsI+YjrfQ= github.com/lib/pq v1.12.3/go.mod h1:/p+8NSbOcwzAEI7wiMXFlgydTwcgTr3OSKMsD2BitpA= +github.com/pb33f/ordered-map/v2 v2.3.1 h1:5319HDO0aw4DA4gzi+zv4FXU9UlSs3xGZ40wcP1nBjY= +github.com/pb33f/ordered-map/v2 v2.3.1/go.mod h1:qxFQgd0PkVUtOMCkTapqotNgzRhMPL7VvaHKbd1HnmQ= github.com/philhofer/fwd v1.2.0 h1:e6DnBTl7vGY+Gz322/ASL4Gyp1FspeMvx1RNDoToZuM= github.com/philhofer/fwd v1.2.0/go.mod h1:RqIHx9QI14HlwKwm98g9Re5prTQ6LdeRQn+gXJFxsJM= github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4= @@ -38,10 +50,22 @@ github.com/pquerna/otp v1.5.0 h1:NMMR+WrmaqXU4EzdGJEE1aUUI0AMRzsp96fFFWNPwxs= github.com/pquerna/otp v1.5.0/go.mod h1:dkJfzwRKNiegxyNb54X/3fLwhCynbMspSyWKnvi1AEg= github.com/redis/go-redis/v9 v9.22.0 h1:laDvpYXTJtZLloinw1fA5Kqd6HAEH2XKxOkG/PDq2F0= github.com/redis/go-redis/v9 v9.22.0/go.mod h1:y2g0Wj8rQvuK0ELM+oxSudcLtC09JScs98I/X9gRWY4= +github.com/standard-webhooks/standard-webhooks/libraries v0.0.1 h1:uOfcYT+3QungH6tIGSVCR/Y3KJmgJiHcojJbMTPDZAI= +github.com/standard-webhooks/standard-webhooks/libraries v0.0.1/go.mod h1:L1MQhA6x4dn9r007T033lsaZMv9EmBAdXyU/+EF40fo= github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME= github.com/stretchr/testify v1.3.0/go.mod h1:M5WIy9Dh21IEIfnGCwXGc5bZfKNJtfHm1UVUgZn+9EI= github.com/stretchr/testify v1.12.1 h1:EuwCh5fleGS7H32xRwO3wRGT7DxrDhLAT6FF8MpWDWE= github.com/stretchr/testify v1.12.1/go.mod h1:MDEgiDPPsNp5cuIrHPPCyornHKgEVbtFUmoNlxoYthg= +github.com/tidwall/gjson v1.14.2/go.mod h1:/wbyibRr2FHMks5tjHJ5F8dMZh3AcwJEMf5vlfC0lxk= +github.com/tidwall/gjson v1.18.0 h1:FIDeeyB800efLX89e5a8Y0BNH+LOngJyGrIWxG2FKQY= +github.com/tidwall/gjson v1.18.0/go.mod h1:/wbyibRr2FHMks5tjHJ5F8dMZh3AcwJEMf5vlfC0lxk= +github.com/tidwall/match v1.1.1 h1:+Ho715JplO36QYgwN9PGYNhgZvoUSc9X2c80KVTi+GA= +github.com/tidwall/match v1.1.1/go.mod h1:eRSPERbgtNPcGhD8UCthc6PmLEQXEWd3PRB5JTxsfmM= +github.com/tidwall/pretty v1.2.0/go.mod h1:ITEVvHYasfjBbM0u2Pg8T2nJnzm8xPwvNhhsoaGGjNU= +github.com/tidwall/pretty v1.2.1 h1:qjsOFOWWQl+N3RsoF5/ssm1pHmJJwhjlSbZ51I6wMl4= +github.com/tidwall/pretty v1.2.1/go.mod h1:ITEVvHYasfjBbM0u2Pg8T2nJnzm8xPwvNhhsoaGGjNU= +github.com/tidwall/sjson v1.2.5 h1:kLy8mja+1c9jlljvWTlSazM7cKDRfJuR/bOJhcY5NcY= +github.com/tidwall/sjson v1.2.5/go.mod h1:Fvgq9kS/6ociJEDnK0Fk1cpYF4FIW6ZF7LAe+6jwd28= github.com/tinylib/msgp v1.6.4 h1:mOwYbyYDLPj35mkA2BjjYejgJk9BuHxDdvRnb6v2ZcQ= github.com/tinylib/msgp v1.6.4/go.mod h1:RSp0LW9oSxFut3KzESt5Voq4GVWyS+PSulT77roAqEA= github.com/x448/float16 v0.8.4 h1:qLwI1I70+NjRFUR3zs1JPUCgaCXSh3SW62uAKT1mSBM= @@ -54,7 +78,13 @@ go.uber.org/mock v0.6.0 h1:hyF9dfmbgIX5EfOdasqLsWD6xqpNZlXblLB/Dbnwv3Y= go.uber.org/mock v0.6.0/go.mod h1:KiVJ4BqZJaMj4svdfmHM0AUx4NJYO8ZNpPnZn1Z+BBU= go.yaml.in/yaml/v3 v3.0.5 h1:N6y/pJk8buWs9NY5ERU2HSMfm+IuD/OtfdAnq6kESPw= go.yaml.in/yaml/v3 v3.0.5/go.mod h1:HVTZu1O7/Vkt2N+BFy8Zza+lnLsABggaTM2ZpNIGuKg= +go.yaml.in/yaml/v4 v4.0.0-rc.2 h1:/FrI8D64VSr4HtGIlUtlFMGsm7H7pWTbj6vOLVZcA6s= +go.yaml.in/yaml/v4 v4.0.0-rc.2/go.mod h1:aZqd9kCMsGL7AuUv/m/PvWLdg5sjJsZ4oHDEnfPPfY0= golang.org/x/crypto v0.55.0 h1:+KWHjbgOaAQ66dh/YlkZKHlz9ZUlq61AFirAR9ntP8M= golang.org/x/crypto v0.55.0/go.mod h1:uq0V9dE/fzQuJtbnL+2EhWOE63vo164FY8xqEnV9xis= +golang.org/x/sync v0.16.0 h1:ycBJEhp9p4vXvUZNszeOq0kGTPghopOL8q0fq3vstxw= +golang.org/x/sync v0.16.0/go.mod h1:1dzgHSNfp02xaA81J2MS99Qcpr2w7fw1gpm99rleRqA= golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs= golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw= +gopkg.in/yaml.v2 v2.2.8 h1:obN1ZagJSUGI0Ek/LBmuj4SNLPfIny3KsKFopxRdj10= +gopkg.in/yaml.v2 v2.2.8/go.mod h1:hI93XBmqTisBFMUTm0b8Fm+jr3Dg1NNxqwp+5A1VGuI= diff --git a/settings/llm.go b/settings/llm.go new file mode 100644 index 0000000..ee0708d --- /dev/null +++ b/settings/llm.go @@ -0,0 +1,136 @@ +package settings + +import ( + "encoding/json" + "errors" + "fmt" + "strings" +) + +// LLMProviderKindAnthropic is the one provider this repo can construct a +// live ai.LLMProvider for. A kind rather than a hardcoded shape, because +// the stored config has to say which client to build and a console needs +// something to put in a dropdown — but the list is deliberately short: +// this repo ships an implementation for what it names and nothing else, +// so "openai" here would be a setting that saves and then does nothing. +const LLMProviderKindAnthropic = "anthropic" + +// DefaultLLMModel is what an operator gets if they configure a provider +// without naming a model. Claude Opus 5 is the current default for this +// kind of work. +const DefaultLLMModel = "claude-opus-5" + +// The bounds on MaxTokens. The floor is not zero: a provider asked for +// zero output tokens would answer with an empty response that this repo +// would then fail to parse, which reads as a broken provider rather than +// as a bad setting. The ceiling keeps one AI query from consuming a +// month's budget in a single call. +const ( + MinLLMMaxTokens = 256 + MaxLLMMaxTokens = 8192 + // DefaultLLMMaxTokens is room for a QueryIntent plus the model's + // reasoning about it. The output here is a small JSON object, not + // prose, so this is generous rather than tight. + DefaultLLMMaxTokens = 2048 +) + +// ErrInvalidLLMProvider means the supplied configuration cannot back a +// working provider. Returned instead of storing it, because a saved +// setting that fails at the first real use is worse than a refused save: +// the operator finds out from a user, not from the form. +var ErrInvalidLLMProvider = errors.New("settings: invalid LLM provider configuration") + +// LLMProviderConfig is what a console reads and writes. It is the +// plaintext shape — it is marshalled to JSON, sealed, and only then +// handed to Secrets, so it exists in memory and inside the ciphertext and +// nowhere else. +type LLMProviderConfig struct { + // Kind names which client to build. See LLMProviderKindAnthropic. + Kind string `json:"kind"` + // Model is the vendor's model id. + Model string `json:"model"` + // APIKey is the credential. Never returned by a read — see Redacted. + APIKey string `json:"api_key"` + // MaxTokens bounds one response. + MaxTokens int `json:"max_tokens"` +} + +// Validate checks a configuration that is about to be stored. The stored +// value is the one every later read has to work with, so this is the only +// place the rule can be enforced — a handler that trusted a form would be +// enforcing it in the wrong layer. +func (c LLMProviderConfig) Validate() error { + if c.Kind != LLMProviderKindAnthropic { + return fmt.Errorf("%w: unknown provider kind %q, this api can only build %q", + ErrInvalidLLMProvider, c.Kind, LLMProviderKindAnthropic) + } + if strings.TrimSpace(c.Model) == "" { + return fmt.Errorf("%w: model is required", ErrInvalidLLMProvider) + } + if strings.TrimSpace(c.APIKey) == "" { + return fmt.Errorf("%w: api_key is required", ErrInvalidLLMProvider) + } + if c.MaxTokens < MinLLMMaxTokens || c.MaxTokens > MaxLLMMaxTokens { + return fmt.Errorf("%w: max_tokens must be between %d and %d, got %d", + ErrInvalidLLMProvider, MinLLMMaxTokens, MaxLLMMaxTokens, c.MaxTokens) + } + return nil +} + +// Redacted returns the copy of this configuration that is safe to send to +// a console: everything except the credential, with a flag saying whether +// one is set. +// +// This is a method on the config rather than something each handler +// remembers to do, and that is the point. A response that carried the API +// key would put it in a browser's memory, in a devtools panel, in a +// screenshot of the settings screen — and the whole reason the key is +// encrypted at rest is that it is worth stealing. +func (c LLMProviderConfig) Redacted() RedactedLLMProvider { + return RedactedLLMProvider{ + Kind: c.Kind, + Model: c.Model, + MaxTokens: c.MaxTokens, + APIKeySet: c.APIKey != "", + Configured: true, + } +} + +// RedactedLLMProvider is what GET returns. Separate from +// LLMProviderConfig rather than the same struct with an empty APIKey +// field, so a handler cannot accidentally serialise the credential by +// using the wrong type — the credential is not a field here at all. +type RedactedLLMProvider struct { + Kind string `json:"kind"` + Model string `json:"model"` + MaxTokens int `json:"max_tokens"` + // APIKeySet says whether a credential is stored. Enough for a console + // to render "•••••••• saved" versus "not set", and not enough to + // reconstruct anything. + APIKeySet bool `json:"api_key_set"` + // Configured distinguishes a stored provider from the zero value this + // struct has when nothing is stored. Both have empty strings, and a + // console that cannot tell them apart shows a blank form where it + // should show "not configured yet". + Configured bool `json:"configured"` +} + +// MarshalLLMProvider seals and encodes a configuration for storage. +// +// The seal-and-encode pair lives here rather than in the handler so that +// "JSON, then encrypted" has exactly one spelling: a second caller that +// encoded differently would produce rows the reader below cannot open. +func MarshalLLMProvider(c LLMProviderConfig) ([]byte, error) { + return json.Marshal(c) +} + +// UnmarshalLLMProvider decodes what MarshalLLMProvider produced. It +// receives plaintext, so it must only ever be handed the output of +// Secrets.Get. +func UnmarshalLLMProvider(plaintext []byte) (LLMProviderConfig, error) { + var c LLMProviderConfig + if err := json.Unmarshal(plaintext, &c); err != nil { + return LLMProviderConfig{}, fmt.Errorf("settings: decoding the stored LLM provider: %w", err) + } + return c, nil +} From b31ef899f4f0f1d38b750320e227e76869d5f825 Mon Sep 17 00:00:00 2001 From: Raymond Nicholas Date: Wed, 16 Sep 2026 00:02:55 +0100 Subject: [PATCH 08/12] feat: add the read-only QueryableStore and its role check MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit cryden's ai.QueryableStore asks for a read-only Postgres role by name, and that is the guarantee that holds when the allowlist has a bug — so the role is verified by attempting a write and confirming the server refuses it, not by reading a checkbox or a role attribute. A refusal is a pass, a success is a refusal, and anything else is "could not verify" rather than a pass. Co-Authored-By: Claude Code --- aiprovider/query.go | 358 +++++++++++++++++++++++++++++++++++++++ aiprovider/query_test.go | 198 ++++++++++++++++++++++ settings/database.go | 191 +++++++++++++++++++++ 3 files changed, 747 insertions(+) create mode 100644 aiprovider/query.go create mode 100644 aiprovider/query_test.go create mode 100644 settings/database.go diff --git a/aiprovider/query.go b/aiprovider/query.go new file mode 100644 index 0000000..9de2097 --- /dev/null +++ b/aiprovider/query.go @@ -0,0 +1,358 @@ +package aiprovider + +import ( + "context" + "database/sql" + "errors" + "fmt" + "strings" + "time" + + "github.com/lib/pq" + + crydenai "github.com/crydensync/cryden/v2/ai" +) + +// ErrNotReadOnly means the supplied role CAN write, so it must not back +// the AI query surface. +var ErrNotReadOnly = errors.New("aiprovider: the database role is not read-only") + +// ErrCannotVerifyReadOnly means the check could not reach a conclusion. +// Separate from ErrNotReadOnly on purpose: "this role can write" and "we +// could not find out" call for different words, and the second one is +// almost always a connection problem the operator can fix. +var ErrCannotVerifyReadOnly = errors.New("aiprovider: could not verify the database role is read-only") + +// readOnlyProbeTimeout bounds the whole check. It is short because this +// runs inside an HTTP request from a settings form — a check that hangs +// for a minute would look like a broken console, and an operator +// configuring a second database expects a few seconds at most. +const readOnlyProbeTimeout = 10 * time.Second + +// probeTable is the scratch table the write attempt targets. It lives in +// pg_temp, the session's own temporary schema, which is what makes this +// check safe to run: the table exists only for the life of this +// connection, is invisible to every other session, and is dropped by +// Postgres when the connection closes — so a failed CREATE leaves nothing +// behind for the operator to clean up. +// +// It is a single-column table with a single row. The point is not the +// data, it is whether the server says yes. +const ( + probeCreate = `CREATE TEMP TABLE cryden_readonly_probe (id int)` + probeInsert = `INSERT INTO cryden_readonly_probe (id) VALUES (1)` +) + +// CheckReadOnly verifies that dsn names a role which cannot write. +// +// This is the check cryden's ai.QueryableStore interface asks for by name: +// "MUST use a read-only Postgres role for this connection — that's a real +// credential-level guarantee, not just a promise made in code, so a bug +// in validation still can't cause a write." The allowlist in ai.validate.go +// is the first line of defence; this is the line that holds when the +// first one has a bug, because a role that cannot INSERT cannot INSERT +// whatever a query builder does with its input. +// +// Which is also why this is checked by *attempting a write and confirming +// it is rejected*, rather than by reading the role's attributes. Trusting +// a checkbox, or pg_roles.rolsuper, or the presence of "readonly" in a +// connection parameter would all be trusting a claim. Only the server's +// refusal is evidence — and it is evidence about the actual role, on the +// actual database, through the actual credentials, which no amount of +// reading metadata can substitute for. +// +// The three outcomes are deliberately distinct: +// +// - the write is refused -> nil, the role is read-only +// - the write succeeds -> ErrNotReadOnly +// - anything else -> ErrCannotVerifyReadOnly +// +// The third is not a pass. A connection that never opened, a timeout, a +// missing table privilege that fails the CREATE for a reason other than +// read-onlyness — none of those prove anything, and treating them as +// success would make this check pass exactly when it is least able to +// tell. +func CheckReadOnly(ctx context.Context, dsn string) error { + ctx, cancel := context.WithTimeout(ctx, readOnlyProbeTimeout) + defer cancel() + + db, err := sql.Open("postgres", dsn) + if err != nil { + return fmt.Errorf("%w: opening the connection: %v", ErrCannotVerifyReadOnly, err) + } + defer db.Close() + + // One connection, used for both statements. pg_temp is per-session, + // so a pool that handed the CREATE and the INSERT to different + // connections would have the INSERT fail on a missing table — a + // refusal that looks like proof of read-onlyness and is not. + db.SetMaxOpenConns(1) + db.SetMaxIdleConns(1) + + if err := db.PingContext(ctx); err != nil { + return fmt.Errorf("%w: connecting: %v", ErrCannotVerifyReadOnly, err) + } + + // A read that must work, before any write is attempted. Without it, a + // role with no rights at all would fail the CREATE below and be + // reported as read-only — which is the right answer by accident, on a + // connection that cannot serve the feature either. + if _, err := db.ExecContext(ctx, `SELECT 1`); err != nil { + return fmt.Errorf("%w: the connection cannot run a query at all: %v", ErrCannotVerifyReadOnly, err) + } + + if _, err := db.ExecContext(ctx, probeCreate); err != nil { + // The CREATE failing is the expected outcome for a role that + // cannot write, and it is also what a dozen unrelated problems + // look like. Postgres distinguishes them: 42501 is + // insufficient_privilege, which is the server saying "this role + // may not do that". Anything else is reported as unverifiable + // rather than assumed to be a refusal. + if isInsufficientPrivilege(err) { + return nil + } + return fmt.Errorf("%w: creating the probe table: %v", ErrCannotVerifyReadOnly, err) + } + + // The role could create a table. It is not read-only, and the INSERT + // is not needed to know that — but running it keeps the failure + // message specific about what succeeded, which is what an operator + // needs to go and fix the grant. + if _, err := db.ExecContext(ctx, probeInsert); err != nil { + return fmt.Errorf("%w: the role may create tables (INSERT failed separately: %v)", ErrNotReadOnly, err) + } + return fmt.Errorf("%w: the role both created a table and inserted a row", ErrNotReadOnly) +} + +// isInsufficientPrivilege reports whether err is Postgres' own "this role +// may not do that" — SQLSTATE 42501. +// +// Checked by code rather than by matching the message, because the +// message is localized and reworded between major versions while the code +// is not, and this is the branch that decides whether a connection is +// accepted. +func isInsufficientPrivilege(err error) bool { + var pqErr *pq.Error + if errors.As(err, &pqErr) { + return pqErr.Code == "42501" + } + // lib/pq returns the parsed error for anything the server answers, + // so a non-pq error here means the statement never reached Postgres. + // Reported as unverifiable rather than as a refusal, which is what + // the caller does with a false return. + return false +} + +// PostgresSnapshot implements cryden's ai.QueryableStore over a +// connection that CheckReadOnly has already accepted. +// +// It runs an already-validated QueryIntent. "Already-validated" is not a +// hope: cryden's ai.ExecuteIntent runs validateIntent before it calls +// RunSafeQuery at all, and widget.Ask runs it too — there is no path to +// this method that skips that. What this type adds is that the statement +// it builds is assembled from cryden's own EntityColumns list rather than +// from anything on the intent, so an entity that somehow got past +// validation still cannot put text of its own into the SQL. +type PostgresSnapshot struct { + db *sql.DB + maxRows int +} + +// NewPostgresSnapshot opens the read-only connection. It does NOT itself +// check that the role is read-only — that is CheckReadOnly's job and it +// belongs to the settings write, where the operator is present to be told +// about a failure. Opening here is deliberately cheap so that a +// deployment whose stored connection has since gone bad reports a query +// error rather than failing to start. +func NewPostgresSnapshot(dsn string, maxRows int) (*PostgresSnapshot, error) { + db, err := sql.Open("postgres", dsn) + if err != nil { + return nil, fmt.Errorf("aiprovider: opening the read-only connection: %w", err) + } + if maxRows < 1 { + maxRows = crydenai.DefaultLimit + } + if maxRows > crydenai.MaxLimit { + maxRows = crydenai.MaxLimit + } + return &PostgresSnapshot{db: db, maxRows: maxRows}, nil +} + +// Close releases the pool. +func (s *PostgresSnapshot) Close() error { return s.db.Close() } + +var _ crydenai.QueryableStore = (*PostgresSnapshot)(nil) + +// RunSafeQuery executes an intent. +// +// Every value that reaches the SQL text comes from cryden's own +// EntityColumns / AllowedFields maps. Filter *values* — the only part +// that originates with a user — are bind parameters, never interpolated. +// Column and operator names cannot be bound in SQL, which is exactly why +// they are taken from a fixed map on this side rather than trusted from +// the intent: a name that is not in the map is refused, not quoted. +func (s *PostgresSnapshot) RunSafeQuery(ctx context.Context, intent crydenai.QueryIntent) (crydenai.QueryResult, error) { + columns, ok := crydenai.EntityColumns[intent.Entity] + if !ok { + return crydenai.QueryResult{}, fmt.Errorf("aiprovider: unknown entity %q", intent.Entity) + } + + // The intent's own Limit was already defaulted and clamped by + // ai.ExecuteIntent; this is the deployment's own ceiling on top of + // the engine's, so an operator can make the AI surface cheaper than + // cryden's maximum without changing the engine. + limit := intent.Limit + if limit <= 0 || limit > s.maxRows { + limit = s.maxRows + } + + var ( + where []string + args []any + ) + for _, filter := range intent.Filters { + if err := checkFilter(intent.Entity, filter); err != nil { + return crydenai.QueryResult{}, err + } + args = append(args, filterArgument(filter.Operator, filter.Value)) + where = append(where, fmt.Sprintf("%s %s $%d", filter.Field, sqlOperator(filter.Operator), len(args))) + } + + statement := buildStatement(intent, columns, where, limit) + rows, err := s.db.QueryContext(ctx, statement, args...) + if err != nil { + return crydenai.QueryResult{}, fmt.Errorf("aiprovider: running the query: %w", err) + } + defer rows.Close() + + return scanResult(rows) +} + +// checkFilter refuses a filter whose field or operator is not in cryden's +// allowlist for that entity. +// +// Both names are checked here as well as in the engine. That is not +// redundancy for its own sake: SQL cannot bind a column or an operator +// name, so those two are the only parts of this query that are spliced +// into text rather than passed as parameters, and a second check at the +// point of assembly is what makes "the statement can only contain names +// from a fixed map" a property of this function rather than a claim about +// its callers. +func checkFilter(entity string, filter crydenai.QueryFilter) error { + if !crydenai.AllowedFields[entity][filter.Field] { + return fmt.Errorf("aiprovider: field %q is not allowed on %q", filter.Field, entity) + } + if !crydenai.AllowedOperators[filter.Operator] { + return fmt.Errorf("aiprovider: operator %q is not allowed", filter.Operator) + } + return nil +} + +// buildStatement assembles the SELECT. Every piece spliced into the text +// is a name drawn from cryden's maps — see RunSafeQuery — so the only +// thing a caller influences is the shape, never the vocabulary. +func buildStatement(intent crydenai.QueryIntent, columns, where []string, limit int) string { + var b strings.Builder + + switch intent.Aggregate { + case "count": + b.WriteString("SELECT count(*) FROM ") + b.WriteString(intent.Entity) + case "group_by": + b.WriteString("SELECT ") + b.WriteString(intent.GroupBy) + b.WriteString(", count(*) FROM ") + b.WriteString(intent.Entity) + default: + b.WriteString("SELECT ") + b.WriteString(strings.Join(columns, ", ")) + b.WriteString(" FROM ") + b.WriteString(intent.Entity) + } + + if len(where) > 0 { + b.WriteString(" WHERE ") + b.WriteString(strings.Join(where, " AND ")) + } + if intent.Aggregate == "group_by" { + b.WriteString(" GROUP BY ") + b.WriteString(intent.GroupBy) + } + fmt.Fprintf(&b, " LIMIT %d", limit) + return b.String() +} + +// sqlOperator maps cryden's operator vocabulary onto SQL. "contains" is +// the one that is not a plain symbol: it becomes a LIKE. +func sqlOperator(operator string) string { + if operator == "contains" { + return "LIKE" + } + return operator +} + +// filterArgument prepares a filter's value for its place in the query. +// +// "contains" is the only operator that needs anything done to its value: +// LIKE's wildcards are part of the pattern, so a caller writing +// "contains: a%" would otherwise get substring semantics they did not ask +// for and probably did not intend. The wildcards are added here, around +// the whole value, so the value is always a literal. +// +// This is not a safety measure — the value is a bind parameter either +// way, so neither form can reach the SQL text. It is about the operator +// meaning what it says. +func filterArgument(operator, value string) string { + if operator == "contains" { + // Escaped so a literal % or _ in the question stays literal: + // the default LIKE escape character is a backslash. + escaped := strings.NewReplacer(`\`, `\\`, `%`, `\%`, `_`, `\_`).Replace(value) + return "%" + escaped + "%" + } + return value +} + +func scanResult(rows *sql.Rows) (crydenai.QueryResult, error) { + // The header comes from the query rather than from the entity + // definition, because an aggregate's columns are not the entity's + // own — a count returns one column that no entity lists. + names, err := rows.Columns() + if err != nil { + return crydenai.QueryResult{}, err + } + + result := crydenai.QueryResult{Columns: names, Rows: [][]string{}} + for rows.Next() { + cells := make([]any, len(names)) + pointers := make([]any, len(names)) + for i := range cells { + pointers[i] = &cells[i] + } + if err := rows.Scan(pointers...); err != nil { + return crydenai.QueryResult{}, err + } + row := make([]string, len(names)) + for i, cell := range cells { + row[i] = renderCell(cell) + } + result.Rows = append(result.Rows, row) + } + return result, rows.Err() +} + +// renderCell turns a scanned value into the string form ai.QueryResult +// promises. Bytes become a string, a time keeps RFC 3339, and a NULL +// becomes empty rather than the word "NULL" — an empty cell is what a +// table shows and what a model reads as "nothing here". +func renderCell(cell any) string { + switch value := cell.(type) { + case nil: + return "" + case []byte: + return string(value) + case time.Time: + return value.Format(time.RFC3339) + default: + return fmt.Sprintf("%v", value) + } +} diff --git a/aiprovider/query_test.go b/aiprovider/query_test.go new file mode 100644 index 0000000..701d034 --- /dev/null +++ b/aiprovider/query_test.go @@ -0,0 +1,198 @@ +package aiprovider + +import ( + "errors" + "strings" + "testing" + + crydenai "github.com/crydensync/cryden/v2/ai" +) + +// The statement builder is where the two spliced names — a column and an +// operator — become SQL text. Everything else in a query is a bind +// parameter. So this is the file that has to be sure about them. +func TestBuildStatementUsesOnlyAllowlistedNames(t *testing.T) { + columns, ok := crydenai.EntityColumns["users"] + if !ok { + t.Fatal("cryden no longer defines EntityColumns for users") + } + + statement := buildStatement( + crydenai.QueryIntent{Entity: "users", Aggregate: "", Limit: 25}, + columns, + nil, + 25, + ) + + want := "SELECT " + strings.Join(columns, ", ") + " FROM users LIMIT 25" + if statement != want { + t.Errorf("statement =\n %s\nwant\n %s", statement, want) + } +} + +func TestBuildStatementShapesAnAggregate(t *testing.T) { + t.Run("count", func(t *testing.T) { + got := buildStatement(crydenai.QueryIntent{Entity: "audit_events", Aggregate: "count", Limit: 5}, nil, nil, 5) + if got != "SELECT count(*) FROM audit_events LIMIT 5" { + t.Errorf("statement = %q", got) + } + }) + + t.Run("group_by", func(t *testing.T) { + got := buildStatement( + crydenai.QueryIntent{Entity: "audit_events", Aggregate: "group_by", GroupBy: "type", Limit: 5}, + nil, nil, 5, + ) + want := "SELECT type, count(*) FROM audit_events GROUP BY type LIMIT 5" + if got != want { + t.Errorf("statement = %q, want %q", got, want) + } + }) +} + +func TestBuildStatementBindsFilterValuesRatherThanInlining(t *testing.T) { + got := buildStatement( + crydenai.QueryIntent{Entity: "users", Limit: 10}, + nil, + []string{"email = $1", "created_at > $2"}, + 10, + ) + if !strings.Contains(got, "WHERE email = $1 AND created_at > $2") { + t.Errorf("statement = %q, want the filters joined with AND and left as placeholders", got) + } +} + +// A column or operator that is not in cryden's allowlist must be refused +// before it can be spliced into a statement. Asserted with values that +// would be an injection if they ever reached the text. +func TestCheckFilterRefusesAnythingNotAllowlisted(t *testing.T) { + cases := []struct { + name string + entity string + filter crydenai.QueryFilter + }{ + {"unknown field", "users", crydenai.QueryFilter{Field: "password_hash", Operator: "=", Value: "x"}}, + {"injected field", "users", crydenai.QueryFilter{Field: "email; DROP TABLE users", Operator: "=", Value: "x"}}, + {"unknown operator", "users", crydenai.QueryFilter{Field: "email", Operator: "OR 1=1 --", Value: "x"}}, + {"field from another entity", "users", crydenai.QueryFilter{Field: "user_agent", Operator: "=", Value: "x"}}, + {"unknown entity", "secrets", crydenai.QueryFilter{Field: "id", Operator: "=", Value: "x"}}, + } + + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + if err := checkFilter(tc.entity, tc.filter); err == nil { + t.Errorf("checkFilter(%q, %+v) was allowed, want a refusal", tc.entity, tc.filter) + } + }) + } +} + +func TestCheckFilterAllowsWhatCrydenAllows(t *testing.T) { + // Every field cryden allowlists on every entity has to pass here too, + // or this repo would refuse queries the engine considers safe. + for entity, fields := range crydenai.AllowedFields { + for field := range fields { + for operator := range crydenai.AllowedOperators { + if err := checkFilter(entity, crydenai.QueryFilter{Field: field, Operator: operator}); err != nil { + t.Errorf("checkFilter(%q, %s %s) = %v, want nil", entity, field, operator, err) + } + } + } + } +} + +// "contains" means substring, and the wildcards that make it one are +// added here rather than taken from the caller's value. +func TestFilterArgumentWrapsContainsAndEscapesItsWildcards(t *testing.T) { + cases := []struct { + value string + want string + }{ + {"dana", "%dana%"}, + {"", "%%"}, + // A caller's own % must stay a literal percent rather than + // becoming a wildcard they did not ask for. + {"100%", `%100\%%`}, + {"a_b", `%a\_b%`}, + {`back\slash`, `%back\\slash%`}, + } + for _, tc := range cases { + if got := filterArgument("contains", tc.value); got != tc.want { + t.Errorf("filterArgument(contains, %q) = %q, want %q", tc.value, got, tc.want) + } + } + + // Every other operator passes its value through untouched — an "=" + // that mangled its value would silently match nothing. + for _, operator := range []string{"=", ">", "<"} { + if got := filterArgument(operator, "100%"); got != "100%" { + t.Errorf("filterArgument(%q, ...) = %q, want the value unchanged", operator, got) + } + } +} + +func TestSqlOperatorMapsOnlyContains(t *testing.T) { + if got := sqlOperator("contains"); got != "LIKE" { + t.Errorf("sqlOperator(contains) = %q, want LIKE", got) + } + for _, operator := range []string{"=", ">", "<"} { + if got := sqlOperator(operator); got != operator { + t.Errorf("sqlOperator(%q) = %q, want it unchanged", operator, got) + } + } +} + +func TestRenderCell(t *testing.T) { + cases := []struct { + name string + cell any + want string + }{ + // NULL is an empty cell. The word "NULL" would be read by a model + // as the four-character string it looks like. + {"null", nil, ""}, + {"bytes", []byte("dana@example.com"), "dana@example.com"}, + {"string", "already a string", "already a string"}, + {"int", 42, "42"}, + {"bool", true, "true"}, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + if got := renderCell(tc.cell); got != tc.want { + t.Errorf("renderCell(%v) = %q, want %q", tc.cell, got, tc.want) + } + }) + } +} + +// Only Postgres' own insufficient_privilege is read as "the role is +// read-only". Everything else has to fall through to unverifiable, or a +// connection failure would be accepted as proof. +func TestIsInsufficientPrivilegeAcceptsOnly42501(t *testing.T) { + if isInsufficientPrivilege(nil) { + t.Error("a nil error was read as insufficient privilege") + } + if isInsufficientPrivilege(errors.New("dial tcp: connection refused")) { + t.Error("a transport error was read as insufficient privilege") + } + if isInsufficientPrivilege(errors.New(`pq: permission denied for table users`)) { + t.Error("a plain error was read as insufficient privilege — only a parsed SQLSTATE may count") + } +} + +// The probe statements are the evidence this whole check rests on, so +// they must be a write and must not touch anything an operator would have +// to clean up afterwards. +func TestProbeStatementsWriteOnlyToATemporaryTable(t *testing.T) { + if !strings.HasPrefix(strings.ToUpper(probeCreate), "CREATE TEMP TABLE") { + t.Errorf("probeCreate = %q, want a TEMP table so a successful probe leaves nothing behind", probeCreate) + } + if !strings.HasPrefix(strings.ToUpper(probeInsert), "INSERT") { + t.Errorf("probeInsert = %q, want a genuine write", probeInsert) + } + for _, statement := range []string{probeCreate, probeInsert} { + if strings.Contains(strings.ToUpper(statement), "PG_CATALOG") || strings.Contains(strings.ToUpper(statement), "PUBLIC.") { + t.Errorf("probe statement %q touches a real schema", statement) + } + } +} diff --git a/settings/database.go b/settings/database.go new file mode 100644 index 0000000..f7716b0 --- /dev/null +++ b/settings/database.go @@ -0,0 +1,191 @@ +package settings + +import ( + "encoding/json" + "errors" + "fmt" + "net/url" + "strings" +) + +// DefaultDatabaseMaxRows mirrors cryden's ai.MaxLimit. The store clamps to +// the engine's own bounds regardless, so this is what an operator gets +// offered rather than a second limit the engine would honour. +const DefaultDatabaseMaxRows = 50 + +// ErrInvalidDatabaseProvider means the supplied connection cannot be used +// for the AI query surface. +var ErrInvalidDatabaseProvider = errors.New("settings: invalid database provider configuration") + +// DatabaseProviderConfig is the connection the AI query surface runs +// against. It exists as a setting at all because cryden's +// ai.QueryableStore is an interface the host implements, and the +// connection behind it is the host's. +// +// The one requirement that is not negotiable: **this must be a read-only +// role.** cryden says so in the interface's own comment — "The only +// production implementation (store/postgres) MUST use a read-only +// Postgres role for this connection — that's a real credential-level +// guarantee, not just a promise made in code, so a bug in validation +// still can't cause a write." The allowlist is the first line; this is +// the line that holds when the allowlist has a bug. +// +// Validate below can only check the shape. Whether the role is ACTUALLY +// read-only is a fact about a database, not about a string, so it is +// checked by connecting and attempting a write — see +// aiprovider.CheckReadOnly, which the handler runs before storing this. +type DatabaseProviderConfig struct { + // Label is what a console shows instead of the whole connection + // string ("the reporting replica"), so an operator can tell two + // configurations apart without reading a password. + Label string `json:"label"` + // DSN is the full connection string, including the credential. + DSN string `json:"dsn"` + // MaxRows caps a single AI-driven query. Bounded by Validate. + MaxRows int `json:"max_rows"` +} + +// Validate checks the shape of a configuration about to be stored. +// +// It deliberately does NOT claim to check that the role is read-only — +// nothing here can. See the type's comment: that check needs a live +// connection and happens in the handler. +func (c DatabaseProviderConfig) Validate() error { + if strings.TrimSpace(c.DSN) == "" { + return fmt.Errorf("%w: dsn is required", ErrInvalidDatabaseProvider) + } + if err := ValidateDSNShape(c.DSN); err != nil { + return err + } + if c.MaxRows < 1 || c.MaxRows > maxDatabaseRows { + return fmt.Errorf("%w: max_rows must be between 1 and %d, got %d", + ErrInvalidDatabaseProvider, maxDatabaseRows, c.MaxRows) + } + if strings.TrimSpace(c.Label) == "" { + return fmt.Errorf("%w: label is required — a console listing two database providers needs to tell them apart without showing a password", ErrInvalidDatabaseProvider) + } + return nil +} + +// maxDatabaseRows bounds MaxRows. cryden clamps any query to its own +// ai.MaxLimit of 500 whatever this says, so a larger value here would be +// a setting that does not do what it appears to; this is the ceiling that +// keeps the two from disagreeing. +const maxDatabaseRows = 500 + +// ValidateDSNShape checks that dsn is a Postgres connection string this +// repo can hand to lib/pq. It is a shape check and nothing more — it does +// not connect, and it knowingly says nothing about whether the credentials +// work, whether the host exists, or what the role may do. +func ValidateDSNShape(dsn string) error { + trimmed := strings.TrimSpace(dsn) + + // The keyword/value form ("host=… user=… ") is valid for lib/pq and + // has no scheme to parse, so it is accepted on the strength of + // carrying a host and a user rather than being forced through URL + // parsing it was never written for. + if !strings.Contains(trimmed, "://") { + if !strings.Contains(trimmed, "host=") { + return fmt.Errorf("%w: dsn is neither a postgres:// URL nor a keyword/value string with a host", ErrInvalidDatabaseProvider) + } + return nil + } + + parsed, err := url.Parse(trimmed) + if err != nil { + return fmt.Errorf("%w: dsn is not a parseable URL: %v", ErrInvalidDatabaseProvider, err) + } + if scheme := parsed.Scheme; scheme != "postgres" && scheme != "postgresql" { + return fmt.Errorf("%w: dsn scheme is %q, want postgres or postgresql", ErrInvalidDatabaseProvider, scheme) + } + if parsed.Host == "" { + return fmt.Errorf("%w: dsn names no host", ErrInvalidDatabaseProvider) + } + return nil +} + +// Redacted returns the copy that is safe to send to a console. The DSN is +// replaced by a description of it rather than by a mask: a masked +// connection string still leaks the host, the database and the role, and +// an operator does not need any of that back to recognise which one is +// stored — the label and the host are enough. +func (c DatabaseProviderConfig) Redacted() RedactedDatabaseProvider { + return RedactedDatabaseProvider{ + Label: c.Label, + MaxRows: c.MaxRows, + DSNSet: c.DSN != "", + Host: dsnHost(c.DSN), + Database: dsnDatabase(c.DSN), + Configured: true, + } +} + +// RedactedDatabaseProvider is what GET returns. The DSN is not a field +// here at all, so a handler cannot serialise it by reaching for the wrong +// type. +type RedactedDatabaseProvider struct { + Label string `json:"label"` + MaxRows int `json:"max_rows"` + DSNSet bool `json:"dsn_set"` + Host string `json:"host"` + Database string `json:"database"` + Configured bool `json:"configured"` +} + +// dsnHost and dsnDatabase pull the two harmless halves out of a +// connection string for display. Best-effort: an unparseable DSN returns +// "" rather than an error, because this is only ever used to decorate a +// response and a display helper that could fail a request would be a +// worse trade than a blank field. +func dsnHost(dsn string) string { + if dsn == "" { + return "" + } + if !strings.Contains(dsn, "://") { + for _, field := range strings.Fields(dsn) { + if after, ok := strings.CutPrefix(field, "host="); ok { + return after + } + } + return "" + } + parsed, err := url.Parse(dsn) + if err != nil { + return "" + } + return parsed.Host +} + +func dsnDatabase(dsn string) string { + if dsn == "" { + return "" + } + if !strings.Contains(dsn, "://") { + for _, field := range strings.Fields(dsn) { + if after, ok := strings.CutPrefix(field, "dbname="); ok { + return after + } + } + return "" + } + parsed, err := url.Parse(dsn) + if err != nil { + return "" + } + return strings.TrimPrefix(parsed.Path, "/") +} + +// MarshalDatabaseProvider and UnmarshalDatabaseProvider are the +// encode/decode pair for storage, the same shape as the LLM provider's, so +// "JSON, then encrypted" has one spelling per setting. +func MarshalDatabaseProvider(c DatabaseProviderConfig) ([]byte, error) { + return json.Marshal(c) +} + +func UnmarshalDatabaseProvider(plaintext []byte) (DatabaseProviderConfig, error) { + var c DatabaseProviderConfig + if err := json.Unmarshal(plaintext, &c); err != nil { + return DatabaseProviderConfig{}, fmt.Errorf("settings: decoding the stored database provider: %w", err) + } + return c, nil +} From 2815e900188b46f567a872f233e33a4839661c46 Mon Sep 17 00:00:00 2001 From: Raymond Nicholas Date: Wed, 16 Sep 2026 00:12:14 +0100 Subject: [PATCH 09/12] feat: expose the AI provider settings endpoints GET/PUT/DELETE /v1/admin/settings/llm-provider and .../database-provider, storing both through settings.Secrets so neither credential reaches the table unsealed. PUT database-provider connects with the supplied credentials and refuses to store anything until the server has rejected a write on that connection, which is the read-only check NEXT.md asks for by name. Co-Authored-By: Claude Code --- httpapi/errors.go | 32 ++ httpapi/router.go | 41 +- httpapi/settings_handlers.go | 375 ++++++++++++++++++ httpapi/settings_handlers_test.go | 608 ++++++++++++++++++++++++++++++ main.go | 18 + 5 files changed, 1072 insertions(+), 2 deletions(-) create mode 100644 httpapi/settings_handlers.go create mode 100644 httpapi/settings_handlers_test.go diff --git a/httpapi/errors.go b/httpapi/errors.go index 63e6451..6b17597 100644 --- a/httpapi/errors.go +++ b/httpapi/errors.go @@ -10,6 +10,8 @@ import ( "github.com/crydensync/cryden/v2/store" "github.com/crydensync/cryden/v2/token" + "github.com/crydensync/api/aiprovider" + "github.com/crydensync/api/settings" "github.com/crydensync/api/usermeta" ) @@ -151,6 +153,36 @@ func mapError(err error) apiError { return apiError{http.StatusBadRequest, "invalid_metadata_key", "a metadata key must start with a letter or underscore and contain only letters, digits, underscores, dots and dashes, up to 64 characters"} case errors.Is(err, usermeta.ErrNotFound): return apiError{http.StatusNotFound, "metadata_key_not_found", "no such metadata key on this user"} + // The settings behind the AI-assisted admin features. The two invalid + // cases are a form the operator can fix, so they say which field and + // which bound rather than a generic "bad request" — the caller is + // already an operator, and none of it is secret. + case errors.Is(err, errSettingsNotConfigured): + return apiError{http.StatusNotFound, "not_configured", "the AI settings endpoints are not configured on this deployment"} + case errors.Is(err, settings.ErrInvalidLLMProvider): + return apiError{http.StatusBadRequest, "invalid_llm_provider", "that LLM provider configuration is not usable"} + case errors.Is(err, settings.ErrInvalidDatabaseProvider): + return apiError{http.StatusBadRequest, "invalid_database_provider", "that database provider configuration is not usable"} + case errors.Is(err, settings.ErrInvalidAskAIWidget): + return apiError{http.StatusBadRequest, "invalid_ask_ai_widget", "that ask-ai widget configuration is not usable"} + // The read-only check. A role that can write is refused outright + // rather than stored with a warning: cryden's design decision is that + // this credential boundary, not the allowlist, is what makes the AI + // query surface safe, so a writable role is a broken guarantee and + // not a preference. "Could not verify" is a different answer for the + // same reason — it is not a pass. + case errors.Is(err, aiprovider.ErrNotReadOnly): + return apiError{http.StatusBadRequest, "database_role_not_read_only", "that database role can write, so it cannot back the AI query surface — create a role with SELECT only and use that"} + case errors.Is(err, aiprovider.ErrCannotVerifyReadOnly): + return apiError{http.StatusBadRequest, "database_role_unverified", "could not verify that the database role is read-only, so it was not stored — check the connection string and that the role can connect"} + // A stored credential this deployment's key cannot open. Distinct + // from "not configured" on purpose: the row is still there, and + // telling an operator it is missing would send them to re-enter a + // credential that nothing is wrong with. + case errors.Is(err, settings.ErrUndecryptable): + return apiError{http.StatusConflict, "setting_undecryptable", "this setting is stored but cannot be decrypted — SETTINGS_ENCRYPTION_KEY has probably changed; re-enter the credential or clear the setting"} + case errors.Is(err, settings.ErrNoEncryptionKey): + return apiError{http.StatusNotFound, "not_configured", "the AI settings endpoints are not configured on this deployment"} // The five "not configured" sentinels below mean this deployment has // not enabled that feature, not that the caller did anything wrong. // 404 rather than 500 so a client can hide the option instead of diff --git a/httpapi/router.go b/httpapi/router.go index 1ef7c6b..3f16659 100644 --- a/httpapi/router.go +++ b/httpapi/router.go @@ -9,6 +9,7 @@ import ( "github.com/crydensync/api/config" "github.com/crydensync/api/digest" + "github.com/crydensync/api/settings" "github.com/crydensync/api/shiplog" "github.com/crydensync/api/usermeta" "github.com/crydensync/api/webhook" @@ -70,6 +71,14 @@ type Deps struct { // The on-demand GET /v1/admin/digest needs nothing from here — it // reads the engine's audit history and records nothing. Digests digest.Store + + // Settings backs the /v1/admin/settings/* endpoints: the LLM provider + // and the read-only database behind the AI-assisted admin features. + // Nil unless SETTINGS_ENCRYPTION_KEY is set — without a key there is + // nowhere safe to put a credential, so those endpoints answer 404 + // not_configured rather than accepting one they would have to store in + // the clear. See settings.Secrets. + Settings *settings.Secrets } // NewRouter builds the full route table. Called once from main.go. @@ -95,6 +104,7 @@ func NewRouter(d Deps) http.Handler { digests := &DigestHandlers{Engine: engine, Store: d.Digests} support := &SupportHandlers{Engine: engine} tuning := &TuningHandlers{Audit: d.Audit, Config: d.Config} + aiSettings := &SettingsHandlers{Secrets: d.Settings} mux := http.NewServeMux() @@ -189,8 +199,10 @@ func NewRouter(d Deps) http.Handler { // this repo's own logic — cryden has no concept of a provider being // reachable, and no bulk way to read stored hash algorithms. // - // Every endpoint here is read-only, and has to stay that way: see - // CLAUDE.md's hard rule about the admin surface. + // Every endpoint here is read-only, and has to stay that way — the one + // exception is the /v1/admin/settings/* block at the bottom of this + // table, which is a settings save and not an action any AI tool can + // reach. See CLAUDE.md's hard rule about the admin surface. mux.HandleFunc("GET /v1/admin/oauth/health", RequireAdmin(engine, oauthHealth.Health)) mux.HandleFunc("GET /v1/admin/security/hash-migration", RequireAdmin(engine, security.HashMigration)) @@ -247,5 +259,30 @@ func NewRouter(d Deps) http.Handler { // settings path. See TuningHandlers and CLAUDE.md's hard rule. mux.HandleFunc("GET /v1/admin/config-tuning", RequireAdmin(engine, tuning.ConfigTuning)) + // The AI settings surface — the only writes under /v1/admin, and the + // "human saves it" half of pre-fill-never-auto-apply. + // + // The read-only rule above is about the AI *tools*, which are what the + // engine's interfaces make read-only by carrying no method that can + // act. These endpoints are the ordinary settings save path those tools' + // output is allowed to pre-fill, and they are where the credentials + // behind the tools live; they do not accept a suggestion, and no AI + // feature holds a reference to this handler. See SettingsHandlers. + // + // PUT database-provider is not a plain write: it connects with the + // supplied credentials and confirms the server refuses a write before + // storing anything, so a role that can modify the database is rejected + // at the form rather than trusted. That is why the endpoint is + // noticeably slower than its neighbours. + mux.HandleFunc("GET /v1/admin/settings/llm-provider", RequireAdmin(engine, aiSettings.LLMProvider)) + mux.HandleFunc("PUT /v1/admin/settings/llm-provider", RequireAdmin(engine, aiSettings.PutLLMProvider)) + mux.HandleFunc("DELETE /v1/admin/settings/llm-provider", RequireAdmin(engine, aiSettings.DeleteLLMProvider)) + mux.HandleFunc("GET /v1/admin/settings/database-provider", RequireAdmin(engine, aiSettings.DatabaseProvider)) + mux.HandleFunc("PUT /v1/admin/settings/database-provider", RequireAdmin(engine, aiSettings.PutDatabaseProvider)) + mux.HandleFunc("DELETE /v1/admin/settings/database-provider", RequireAdmin(engine, aiSettings.DeleteDatabaseProvider)) + mux.HandleFunc("GET /v1/admin/settings/ask-ai-widget", RequireAdmin(engine, aiSettings.AskAIWidget)) + mux.HandleFunc("PUT /v1/admin/settings/ask-ai-widget", RequireAdmin(engine, aiSettings.PutAskAIWidget)) + mux.HandleFunc("DELETE /v1/admin/settings/ask-ai-widget", RequireAdmin(engine, aiSettings.DeleteAskAIWidget)) + return mux } diff --git a/httpapi/settings_handlers.go b/httpapi/settings_handlers.go new file mode 100644 index 0000000..caec50c --- /dev/null +++ b/httpapi/settings_handlers.go @@ -0,0 +1,375 @@ +package httpapi + +import ( + "encoding/json" + "errors" + "net/http" + + "github.com/crydensync/api/aiprovider" + "github.com/crydensync/api/settings" +) + +// The errors these handlers add to the map. Declared here rather than in +// errors.go because they belong to this surface and mapError reaches them +// through errors.Is like everything else — see writeErr. +// errSettingsNotConfigured means SETTINGS_ENCRYPTION_KEY is unset, so +// there is nowhere to put a credential. A wiring fact, not a fault: 404 +// not_configured, the same shape every unconfigured feature in this api +// uses. Declared here rather than in errors.go because it belongs to this +// surface, and mapError reaches it through errors.Is like everything else +// — see writeErr. +var errSettingsNotConfigured = errors.New("the AI settings endpoints are not configured on this deployment") + +// SettingsHandlers serves the runtime configuration behind the AI-assisted +// admin features: which LLM provider to call, and which read-only database +// the ask-ai widget queries. +// +// This is the one admin surface in this repo that WRITES, and that is not +// a hole in the read-only rule — it is the rule's other half. CLAUDE.md's +// constraint is that the AI *tools* must stay read-only, and the decision +// recorded in NEXT.md is pre-fill, never auto-apply: a tuning suggestion +// pre-fills a settings field and a human saves it. These endpoints are +// that save. They store what an operator typed into the form; they do not +// accept a suggestion from any AI feature, and no AI feature can reach +// them — the tuning advisor's handler is built with no reference to this +// one at all. +// +// Every write here is also gated on RequireAdmin in the router, and every +// stored credential is sealed by settings.Secrets before it reaches the +// table. Nothing on this surface ever returns a credential back. +type SettingsHandlers struct { + Secrets *settings.Secrets +} + +// llmProviderDTO is the PUT body. Separate from settings.LLMProviderConfig +// rather than reusing it, so the wire shape and the stored shape can +// differ — and so a future field on the stored config is not silently +// accepted from a request that predates it. +type llmProviderDTO struct { + Kind string `json:"kind"` + Model string `json:"model"` + APIKey string `json:"api_key"` + MaxTokens int `json:"max_tokens"` +} + +// LLMProvider returns the stored provider with its credential removed. +// +// The API key is never in the response in any form — not masked, not +// truncated to its last four characters. A console does not need it: it +// renders "saved" from api_key_set. Anything more would put the +// credential in a browser's memory and a devtools panel, which is exactly +// what encrypting it at rest was meant to avoid. +func (h *SettingsHandlers) LLMProvider(w http.ResponseWriter, r *http.Request) { + if !h.Secrets.Configured() { + writeErr(w, errSettingsNotConfigured) + return + } + + plaintext, err := h.Secrets.Get(r.Context(), settings.KeyLLMProvider) + if errors.Is(err, settings.ErrNotFound) { + // Not an error: "nothing configured yet" is an answer a settings + // screen needs, and a 404 would make a console render a broken + // panel rather than an empty form. + writeData(w, http.StatusOK, settings.RedactedLLMProvider{}) + return + } + if err != nil { + writeErr(w, err) + return + } + + stored, err := settings.UnmarshalLLMProvider(plaintext) + if err != nil { + writeErr(w, err) + return + } + writeData(w, http.StatusOK, stored.Redacted()) +} + +// PutLLMProvider stores a provider configuration. +// +// The credential is required on every write rather than being optional +// with "blank means keep the existing one". That convention is the usual +// one and it is the wrong one here: an omitted field and a deliberately +// cleared one would be the same request, and the failure mode of getting +// it wrong is a settings form that appears to save an unchanged key while +// silently storing an empty one. Requiring it means the console always +// sends what it means. +func (h *SettingsHandlers) PutLLMProvider(w http.ResponseWriter, r *http.Request) { + if !h.Secrets.Configured() { + writeErr(w, errSettingsNotConfigured) + return + } + + var dto llmProviderDTO + if err := json.NewDecoder(r.Body).Decode(&dto); err != nil { + writeBadRequest(w, "malformed request body") + return + } + + // An omitted max_tokens is the common case — a console that does not + // ask, or asks with the field left empty — so it defaults here rather + // than being a validation failure. Any other value is passed through + // to Validate, which is the one place the bounds live. + if dto.MaxTokens == 0 { + dto.MaxTokens = settings.DefaultLLMMaxTokens + } + + config := settings.LLMProviderConfig{ + Kind: dto.Kind, + Model: dto.Model, + APIKey: dto.APIKey, + MaxTokens: dto.MaxTokens, + } + if err := config.Validate(); err != nil { + // The engine's message names the offending field and bound, which + // is what the operator needs to fix the form. Unlike an + // authentication error there is nothing secret in it, and the + // caller is already an operator. + writeErr(w, err) + return + } + + plaintext, err := settings.MarshalLLMProvider(config) + if err != nil { + writeErr(w, err) + return + } + if err := h.Secrets.Put(r.Context(), settings.KeyLLMProvider, plaintext); err != nil { + writeErr(w, err) + return + } + + writeData(w, http.StatusOK, config.Redacted()) +} + +// DeleteLLMProvider clears the stored provider, which switches the +// AI-assisted query features off until one is configured again. +// +// No key is needed to do this — see settings.Secrets.Delete — which is +// what makes it the way out for a deployment that has lost +// SETTINGS_ENCRYPTION_KEY and can no longer read what it stored. +func (h *SettingsHandlers) DeleteLLMProvider(w http.ResponseWriter, r *http.Request) { + if !h.Secrets.Configured() { + writeErr(w, errSettingsNotConfigured) + return + } + if err := h.Secrets.Delete(r.Context(), settings.KeyLLMProvider); err != nil { + writeErr(w, err) + return + } + writeData(w, http.StatusOK, settings.RedactedLLMProvider{}) +} + +// databaseProviderDTO is the PUT body for the read-only database. +type databaseProviderDTO struct { + Label string `json:"label"` + DSN string `json:"dsn"` + MaxRows int `json:"max_rows"` +} + +// DatabaseProvider returns the stored connection with its credential +// removed. The host and database come back so an operator can tell which +// one is configured; the user and password never do. +func (h *SettingsHandlers) DatabaseProvider(w http.ResponseWriter, r *http.Request) { + if !h.Secrets.Configured() { + writeErr(w, errSettingsNotConfigured) + return + } + + plaintext, err := h.Secrets.Get(r.Context(), settings.KeyDatabaseProvider) + if errors.Is(err, settings.ErrNotFound) { + writeData(w, http.StatusOK, settings.RedactedDatabaseProvider{}) + return + } + if err != nil { + writeErr(w, err) + return + } + + stored, err := settings.UnmarshalDatabaseProvider(plaintext) + if err != nil { + writeErr(w, err) + return + } + writeData(w, http.StatusOK, stored.Redacted()) +} + +// PutDatabaseProvider stores a connection, but only after checking that +// its role genuinely cannot write. +// +// This is the check NEXT.md calls for by name and the one cryden's own +// ai.QueryableStore interface demands: "the credential boundary, not just +// the allowlist, is the real safety guarantee." The allowlist in +// cryden's ai package is the first line; a role that cannot INSERT is +// what still holds when the first line has a bug. +// +// So the order matters and is the whole design of this handler: +// +// 1. Check the shape, which is cheap and local. +// 2. Connect with the supplied credentials and TRY TO WRITE — see +// aiprovider.CheckReadOnly. Nothing is stored until the server has +// refused a write on that connection, with those credentials, +// against that database. +// 3. Only then store it. +// +// A "read-only?" checkbox in the console would be a claim about a +// database made by whoever ticked it. This is the database answering for +// itself. It cannot be done on the client side either — a browser cannot +// open a Postgres connection, and if it could, its answer would be no +// more trustworthy than the form. +// +// The cost is that this endpoint is slower than the others: it opens a +// connection and runs two statements before it answers. That is the price +// of the guarantee, and it is paid once per save rather than per query. +func (h *SettingsHandlers) PutDatabaseProvider(w http.ResponseWriter, r *http.Request) { + if !h.Secrets.Configured() { + writeErr(w, errSettingsNotConfigured) + return + } + + var dto databaseProviderDTO + if err := json.NewDecoder(r.Body).Decode(&dto); err != nil { + writeBadRequest(w, "malformed request body") + return + } + if dto.MaxRows == 0 { + dto.MaxRows = settings.DefaultDatabaseMaxRows + } + + config := settings.DatabaseProviderConfig{ + Label: dto.Label, + DSN: dto.DSN, + MaxRows: dto.MaxRows, + } + if err := config.Validate(); err != nil { + writeErr(w, err) + return + } + + // The read-only check, on a live connection, before anything is + // stored. Its own timeout bounds it — see readOnlyProbeTimeout — so a + // host that black-holes the connection answers rather than hanging + // the request. + if err := aiprovider.CheckReadOnly(r.Context(), config.DSN); err != nil { + writeErr(w, err) + return + } + + plaintext, err := settings.MarshalDatabaseProvider(config) + if err != nil { + writeErr(w, err) + return + } + if err := h.Secrets.Put(r.Context(), settings.KeyDatabaseProvider, plaintext); err != nil { + writeErr(w, err) + return + } + + writeData(w, http.StatusOK, config.Redacted()) +} + +// DeleteDatabaseProvider clears the stored connection. Like the LLM +// provider's, it needs no encryption key. +func (h *SettingsHandlers) DeleteDatabaseProvider(w http.ResponseWriter, r *http.Request) { + if !h.Secrets.Configured() { + writeErr(w, errSettingsNotConfigured) + return + } + if err := h.Secrets.Delete(r.Context(), settings.KeyDatabaseProvider); err != nil { + writeErr(w, err) + return + } + writeData(w, http.StatusOK, settings.RedactedDatabaseProvider{}) +} + +// AskAIWidget returns the stored widget configuration. +// +// Unlike the other two GETs there is no redaction here, because there is +// nothing to redact: origins, entity names and copy are what the console +// shows an operator anyway. See settings.AskAIWidgetConfig for why that +// difference is real rather than an oversight. +// +// There is also no embed snippet in the response, deliberately. The +// snippet is markup the csax+ console renders into its own pages, and the +// URL in it would name an endpoint this repo does not serve yet — so +// generating one here would be handing a console a script tag pointing at +// a 404. What this endpoint owes the console is the configuration the +// snippet is built from, which is exactly what it returns. +func (h *SettingsHandlers) AskAIWidget(w http.ResponseWriter, r *http.Request) { + if !h.Secrets.Configured() { + writeErr(w, errSettingsNotConfigured) + return + } + + plaintext, err := h.Secrets.Get(r.Context(), settings.KeyAskAIWidget) + if errors.Is(err, settings.ErrNotFound) { + // A disabled widget is the honest default for "nothing has been + // configured": the zero value is Enabled=false, so a console that + // renders it shows the feature as off rather than showing a form + // that looks saved. + writeData(w, http.StatusOK, settings.AskAIWidgetConfig{}) + return + } + if err != nil { + writeErr(w, err) + return + } + + stored, err := settings.UnmarshalAskAIWidget(plaintext) + if err != nil { + writeErr(w, err) + return + } + writeData(w, http.StatusOK, stored) +} + +// PutAskAIWidget stores the widget configuration. +// +// This is the endpoint the config-tuning advisor's suggestions pre-fill, +// in the sense NEXT.md records: a suggestion about, say, scope ends up as +// a value in this form, and this handler is what writes it once a human +// has read it and pressed save. Nothing calls it automatically, and no +// AI-assisted handler in this repo holds a reference to it. +func (h *SettingsHandlers) PutAskAIWidget(w http.ResponseWriter, r *http.Request) { + if !h.Secrets.Configured() { + writeErr(w, errSettingsNotConfigured) + return + } + + var config settings.AskAIWidgetConfig + if err := json.NewDecoder(r.Body).Decode(&config); err != nil { + writeBadRequest(w, "malformed request body") + return + } + if err := config.Validate(); err != nil { + writeErr(w, err) + return + } + + plaintext, err := settings.MarshalAskAIWidget(config) + if err != nil { + writeErr(w, err) + return + } + if err := h.Secrets.Put(r.Context(), settings.KeyAskAIWidget, plaintext); err != nil { + writeErr(w, err) + return + } + + writeData(w, http.StatusOK, config) +} + +// DeleteAskAIWidget clears the stored configuration, which leaves the +// widget disabled — the zero value — until it is configured again. +func (h *SettingsHandlers) DeleteAskAIWidget(w http.ResponseWriter, r *http.Request) { + if !h.Secrets.Configured() { + writeErr(w, errSettingsNotConfigured) + return + } + if err := h.Secrets.Delete(r.Context(), settings.KeyAskAIWidget); err != nil { + writeErr(w, err) + return + } + writeData(w, http.StatusOK, settings.AskAIWidgetConfig{}) +} diff --git a/httpapi/settings_handlers_test.go b/httpapi/settings_handlers_test.go new file mode 100644 index 0000000..920ff7a --- /dev/null +++ b/httpapi/settings_handlers_test.go @@ -0,0 +1,608 @@ +package httpapi + +import ( + "context" + "encoding/json" + "net/http" + "net/http/httptest" + "strings" + "testing" + + "github.com/crydensync/cryden/v2" + "github.com/crydensync/cryden/v2/store/memory" + "github.com/crydensync/cryden/v2/token" + + "github.com/crydensync/api/settings" +) + +// testSettingsKey is the deployment key these tests seal with. Not a +// secret and not meant to look like one — the point of the encryption +// tests below is that a value written through the handler does not appear +// in the store as the bytes that went in, which any non-empty key shows. +const testSettingsKey = "test-settings-encryption-key" + +type settingsFixture struct { + store *settings.MemoryStore + router http.Handler + + adminToken string + userToken string +} + +// newSettingsFixture builds an engine on in-memory stores plus a settings +// store held directly, because the point of several of these tests is +// what did or did not reach the table — which only the store itself can +// answer. key is the deployment's SETTINGS_ENCRYPTION_KEY; "" is the +// unconfigured deployment, which is a case worth testing on its own. +func newSettingsFixture(t *testing.T, key string) settingsFixture { + t.Helper() + ctx := context.Background() + + var adminID string + engine, err := cryden.New(cryden.Config{ + JWTSecret: "test-secret", + Users: memory.NewUserStore(), + Sessions: memory.NewSessionStore(), + Audit: memory.NewAuditStore(), + Verifications: memory.NewVerificationStore(), + EmailSender: stubMailSender{}, + MagicLinkSender: stubMailSender{}, + AccessTokenClaims: token.ClaimsFunc(func(_ context.Context, userID string) (map[string]any, error) { + if userID == adminID { + return map[string]any{"role": "admin"}, nil + } + return nil, nil + }), + }) + if err != nil { + t.Fatalf("cryden.New on the in-memory stores: %v", err) + } + + admin, err := cryden.SignUp(ctx, engine, "operator@example.com", testPassword, "203.0.113.1") + if err != nil { + t.Fatalf("signup (operator): %v", err) + } + adminID = admin.ID + adminTokens, err := cryden.Login(ctx, engine, "operator@example.com", testPassword, "203.0.113.1", chromeOnMacOS) + if err != nil { + t.Fatalf("login (operator): %v", err) + } + + if _, err := cryden.SignUp(ctx, engine, "dana@example.com", testPassword, "203.0.113.2"); err != nil { + t.Fatalf("signup (user): %v", err) + } + userTokens, err := cryden.Login(ctx, engine, "dana@example.com", testPassword, "203.0.113.2", chromeOnMacOS) + if err != nil { + t.Fatalf("login (user): %v", err) + } + + store := settings.NewMemoryStore() + secrets, err := settings.NewSecrets(store, key) + if err != nil { + t.Fatalf("settings.NewSecrets: %v", err) + } + + return settingsFixture{ + store: store, + router: NewRouter(Deps{Engine: engine, Settings: secrets}), + adminToken: adminTokens.AccessToken, + userToken: userTokens.AccessToken, + } +} + +func (f settingsFixture) do(t *testing.T, method, path, token, body string) *httptest.ResponseRecorder { + t.Helper() + var reader *strings.Reader + if body == "" { + reader = strings.NewReader("") + } else { + reader = strings.NewReader(body) + } + req := httptest.NewRequest(method, path, reader) + if token != "" { + req.Header.Set("Authorization", "Bearer "+token) + } + rec := httptest.NewRecorder() + f.router.ServeHTTP(rec, req) + return rec +} + +// raw returns what is actually in the store under key, which is the only +// way to assert that a credential was sealed rather than trusted to a +// handler's good intentions. +func (f settingsFixture) raw(t *testing.T, key string) ([]byte, bool) { + t.Helper() + return f.store.Raw(key) +} + +const llmProviderPath = "/v1/admin/settings/llm-provider" +const databaseProviderPath = "/v1/admin/settings/database-provider" +const widgetPath = "/v1/admin/settings/ask-ai-widget" + +// Every route on this surface writes or reads a credential or a scope, so +// all nine are behind the same gate as the rest of /v1/admin. +func TestSettingsRoutesAreGatedByRequireAdmin(t *testing.T) { + f := newSettingsFixture(t, testSettingsKey) + + routes := []struct { + method string + path string + }{ + {http.MethodGet, llmProviderPath}, + {http.MethodPut, llmProviderPath}, + {http.MethodDelete, llmProviderPath}, + {http.MethodGet, databaseProviderPath}, + {http.MethodPut, databaseProviderPath}, + {http.MethodDelete, databaseProviderPath}, + {http.MethodGet, widgetPath}, + {http.MethodPut, widgetPath}, + {http.MethodDelete, widgetPath}, + } + + for _, route := range routes { + t.Run(route.method+" "+route.path, func(t *testing.T) { + if rec := f.do(t, route.method, route.path, "", `{}`); rec.Code != http.StatusUnauthorized { + t.Errorf("no token: status = %d, want 401", rec.Code) + } + if rec := f.do(t, route.method, route.path, "not-a-real-token", `{}`); rec.Code != http.StatusUnauthorized { + t.Errorf("garbage token: status = %d, want 401", rec.Code) + } + if rec := f.do(t, route.method, route.path, f.userToken, `{}`); rec.Code != http.StatusForbidden { + t.Errorf("ordinary user: status = %d, want 403 (body %s)", rec.Code, rec.Body.String()) + } + }) + } +} + +// A deployment with no SETTINGS_ENCRYPTION_KEY must refuse everything +// rather than store a credential it cannot seal. This is the case the +// whole Secrets type exists for, so it is asserted against every route +// and against the store, not just against the status code. +func TestSettingsRoutesRefuseWhenNoEncryptionKeyIsConfigured(t *testing.T) { + f := newSettingsFixture(t, "") + + routes := []struct{ method, path, body string }{ + {http.MethodGet, llmProviderPath, ""}, + {http.MethodPut, llmProviderPath, `{"kind":"anthropic","model":"claude-opus-5","api_key":"sk-ant-x"}`}, + {http.MethodDelete, llmProviderPath, ""}, + {http.MethodGet, databaseProviderPath, ""}, + {http.MethodPut, databaseProviderPath, `{"label":"replica","dsn":"postgres://ro@127.0.0.1:5432/db"}`}, + {http.MethodDelete, databaseProviderPath, ""}, + {http.MethodGet, widgetPath, ""}, + {http.MethodPut, widgetPath, `{"enabled":true}`}, + {http.MethodDelete, widgetPath, ""}, + } + + for _, route := range routes { + t.Run(route.method+" "+route.path, func(t *testing.T) { + rec := f.do(t, route.method, route.path, f.adminToken, route.body) + if rec.Code != http.StatusNotFound { + t.Fatalf("status = %d, want 404 (body %s)", rec.Code, rec.Body.String()) + } + if !strings.Contains(rec.Body.String(), "not_configured") { + t.Errorf("body = %s, want the not_configured code", rec.Body.String()) + } + }) + } + + keys, err := f.store.Keys(context.Background()) + if err != nil { + t.Fatalf("listing the store: %v", err) + } + if len(keys) != 0 { + t.Errorf("store holds %v, want nothing — an unconfigured deployment must not write a credential", keys) + } +} + +// The credential must not survive as the bytes that were submitted. This +// is the CLAUDE.md rule about treating this key like JWT_SECRET, asserted +// where it can actually be observed: in the table. +func TestLLMProviderKeyIsSealedBeforeItReachesTheStore(t *testing.T) { + f := newSettingsFixture(t, testSettingsKey) + + const apiKey = "sk-ant-super-secret-value" + body := `{"kind":"anthropic","model":"claude-opus-5","api_key":"` + apiKey + `"}` + + if rec := f.do(t, http.MethodPut, llmProviderPath, f.adminToken, body); rec.Code != http.StatusOK { + t.Fatalf("PUT: status = %d, want 200 (body %s)", rec.Code, rec.Body.String()) + } + + stored, ok := f.raw(t, settings.KeyLLMProvider) + if !ok { + t.Fatal("nothing was stored under the llm provider key") + } + if strings.Contains(string(stored), apiKey) { + t.Errorf("the stored value contains the api key in the clear: %s", stored) + } + if strings.Contains(string(stored), "claude-opus-5") { + t.Errorf("the stored value contains the plaintext config: %s", stored) + } +} + +// The credential never comes back out, in any form. Not masked, not +// truncated — the response type has no field for it at all, and this +// asserts the wire shape rather than the type, because the type is only a +// promise until something serialises it. +func TestLLMProviderGetNeverReturnsTheCredential(t *testing.T) { + f := newSettingsFixture(t, testSettingsKey) + + const apiKey = "sk-ant-super-secret-value" + body := `{"kind":"anthropic","model":"claude-opus-5","api_key":"` + apiKey + `","max_tokens":4096}` + if rec := f.do(t, http.MethodPut, llmProviderPath, f.adminToken, body); rec.Code != http.StatusOK { + t.Fatalf("PUT: status = %d (body %s)", rec.Code, rec.Body.String()) + } + + rec := f.do(t, http.MethodGet, llmProviderPath, f.adminToken, "") + if rec.Code != http.StatusOK { + t.Fatalf("GET: status = %d (body %s)", rec.Code, rec.Body.String()) + } + if strings.Contains(rec.Body.String(), apiKey) { + t.Errorf("GET returned the api key: %s", rec.Body.String()) + } + if strings.Contains(rec.Body.String(), "sk-ant") { + t.Errorf("GET returned part of the api key: %s", rec.Body.String()) + } + + var got struct { + Data struct { + Kind string `json:"kind"` + Model string `json:"model"` + MaxTokens int `json:"max_tokens"` + APIKeySet bool `json:"api_key_set"` + Configured bool `json:"configured"` + } `json:"data"` + } + if err := json.Unmarshal(rec.Body.Bytes(), &got); err != nil { + t.Fatalf("decoding the response: %v", err) + } + if !got.Data.Configured || !got.Data.APIKeySet { + t.Errorf("api_key_set = %v, configured = %v, want both true", got.Data.APIKeySet, got.Data.Configured) + } + if got.Data.Model != "claude-opus-5" || got.Data.MaxTokens != 4096 { + t.Errorf("model = %q, max_tokens = %d, want the stored values back", got.Data.Model, got.Data.MaxTokens) + } +} + +// "Nothing configured yet" is an answer, not a 404. A console renders an +// empty form from it; a 404 would make it render a broken panel. +func TestLLMProviderGetBeforeAnySaveIsAnEmptyAnswer(t *testing.T) { + f := newSettingsFixture(t, testSettingsKey) + + rec := f.do(t, http.MethodGet, llmProviderPath, f.adminToken, "") + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200 (body %s)", rec.Code, rec.Body.String()) + } + + var got struct { + Data struct { + Configured bool `json:"configured"` + APIKeySet bool `json:"api_key_set"` + } `json:"data"` + } + if err := json.Unmarshal(rec.Body.Bytes(), &got); err != nil { + t.Fatalf("decoding the response: %v", err) + } + if got.Data.Configured || got.Data.APIKeySet { + t.Errorf("configured = %v, api_key_set = %v, want both false", got.Data.Configured, got.Data.APIKeySet) + } +} + +// The credential is required on every write rather than "blank keeps the +// existing one" — see PutLLMProvider's comment for why the usual +// convention is the wrong one here. +func TestPutLLMProviderRequiresTheCredentialEveryTime(t *testing.T) { + f := newSettingsFixture(t, testSettingsKey) + + first := `{"kind":"anthropic","model":"claude-opus-5","api_key":"sk-ant-first"}` + if rec := f.do(t, http.MethodPut, llmProviderPath, f.adminToken, first); rec.Code != http.StatusOK { + t.Fatalf("first PUT: status = %d (body %s)", rec.Code, rec.Body.String()) + } + + second := `{"kind":"anthropic","model":"claude-opus-5","api_key":""}` + rec := f.do(t, http.MethodPut, llmProviderPath, f.adminToken, second) + if rec.Code != http.StatusBadRequest { + t.Fatalf("second PUT: status = %d, want 400 (body %s)", rec.Code, rec.Body.String()) + } + if !strings.Contains(rec.Body.String(), "invalid_llm_provider") { + t.Errorf("body = %s, want invalid_llm_provider", rec.Body.String()) + } + + // And the refusal changed nothing: the first credential is still what + // is stored, so a console that retried with an empty field has not + // silently cleared a working key. + if _, ok := f.raw(t, settings.KeyLLMProvider); !ok { + t.Error("the refused write cleared the stored provider") + } +} + +func TestDeleteLLMProviderClearsWithoutNeedingTheCredential(t *testing.T) { + f := newSettingsFixture(t, testSettingsKey) + + body := `{"kind":"anthropic","model":"claude-opus-5","api_key":"sk-ant-x"}` + if rec := f.do(t, http.MethodPut, llmProviderPath, f.adminToken, body); rec.Code != http.StatusOK { + t.Fatalf("PUT: status = %d (body %s)", rec.Code, rec.Body.String()) + } + + if rec := f.do(t, http.MethodDelete, llmProviderPath, f.adminToken, ""); rec.Code != http.StatusOK { + t.Fatalf("DELETE: status = %d (body %s)", rec.Code, rec.Body.String()) + } + if _, ok := f.raw(t, settings.KeyLLMProvider); ok { + t.Error("DELETE left the row in place") + } +} + +// A DSN that cannot be connected to must be refused AND not stored. This +// is the ordering that matters: the read-only check runs before the write, +// so a failure in it leaves the deployment exactly as it was rather than +// half-configured with a connection nothing has verified. +// +// The target is a closed port on the loopback interface, which refuses +// immediately and needs no server — the point is the failure path, and +// CheckReadOnly reports it as unverifiable rather than as a pass or a +// refusal. That distinction is what this asserts. +func TestPutDatabaseProviderRefusesAnUnverifiableConnection(t *testing.T) { + f := newSettingsFixture(t, testSettingsKey) + + body := `{"label":"reporting replica","dsn":"postgres://reader:pw@127.0.0.1:1/cryden?sslmode=disable&connect_timeout=2"}` + rec := f.do(t, http.MethodPut, databaseProviderPath, f.adminToken, body) + if rec.Code != http.StatusBadRequest { + t.Fatalf("status = %d, want 400 (body %s)", rec.Code, rec.Body.String()) + } + if !strings.Contains(rec.Body.String(), "database_role_unverified") { + t.Errorf("body = %s, want database_role_unverified — a connection that could not be made is not a pass", rec.Body.String()) + } + + if _, ok := f.raw(t, settings.KeyDatabaseProvider); ok { + t.Error("a connection the check could not verify was stored anyway") + } +} + +// A DSN that is not a Postgres connection string is caught before any +// connection is attempted, so a malformed form is a fast failure rather +// than a ten-second timeout. +func TestPutDatabaseProviderRejectsAMalformedDSNWithoutConnecting(t *testing.T) { + f := newSettingsFixture(t, testSettingsKey) + + cases := []struct{ name, body string }{ + {"no dsn", `{"label":"replica"}`}, + {"wrong scheme", `{"label":"replica","dsn":"mysql://reader@db.example.com/cryden"}`}, + {"no host", `{"label":"replica","dsn":"postgres://"}`}, + {"not a url or keyword string", `{"label":"replica","dsn":"just some words"}`}, + {"no label", `{"label":" ","dsn":"postgres://reader@db.example.com/cryden"}`}, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + rec := f.do(t, http.MethodPut, databaseProviderPath, f.adminToken, tc.body) + if rec.Code != http.StatusBadRequest { + t.Fatalf("status = %d, want 400 (body %s)", rec.Code, rec.Body.String()) + } + if !strings.Contains(rec.Body.String(), "invalid_database_provider") { + t.Errorf("body = %s, want invalid_database_provider", rec.Body.String()) + } + }) + } + + if _, ok := f.raw(t, settings.KeyDatabaseProvider); ok { + t.Error("a malformed configuration reached the store") + } +} + +// The stored DSN is a credential like the API key and is sealed the same +// way, and GET describes it without returning it. +func TestDatabaseProviderDSNIsSealedAndNeverReturned(t *testing.T) { + f := newSettingsFixture(t, testSettingsKey) + + // Sealed directly rather than through the handler, because the handler + // refuses to store anything it could not connect to and there is no + // Postgres in this test. What is under test here is the storage and + // the response shape, both of which sit behind that check. + const dsn = "postgres://reader:hunter2@db.internal:5432/cryden_ro?sslmode=require" + config := settings.DatabaseProviderConfig{Label: "reporting replica", DSN: dsn, MaxRows: 50} + if err := config.Validate(); err != nil { + t.Fatalf("the fixture config does not validate: %v", err) + } + plaintext, err := settings.MarshalDatabaseProvider(config) + if err != nil { + t.Fatalf("marshalling: %v", err) + } + secrets, err := settings.NewSecrets(f.store, testSettingsKey) + if err != nil { + t.Fatalf("settings.NewSecrets: %v", err) + } + if err := secrets.Put(context.Background(), settings.KeyDatabaseProvider, plaintext); err != nil { + t.Fatalf("storing: %v", err) + } + + stored, ok := f.raw(t, settings.KeyDatabaseProvider) + if !ok { + t.Fatal("nothing was stored") + } + if strings.Contains(string(stored), "hunter2") { + t.Errorf("the stored value contains the password in the clear: %s", stored) + } + + rec := f.do(t, http.MethodGet, databaseProviderPath, f.adminToken, "") + if rec.Code != http.StatusOK { + t.Fatalf("GET: status = %d (body %s)", rec.Code, rec.Body.String()) + } + for _, secret := range []string{"hunter2", "reader:", dsn} { + if strings.Contains(rec.Body.String(), secret) { + t.Errorf("GET returned %q: %s", secret, rec.Body.String()) + } + } + + var got struct { + Data struct { + Label string `json:"label"` + Host string `json:"host"` + Database string `json:"database"` + DSNSet bool `json:"dsn_set"` + Configured bool `json:"configured"` + } `json:"data"` + } + if err := json.Unmarshal(rec.Body.Bytes(), &got); err != nil { + t.Fatalf("decoding the response: %v", err) + } + // The host and database come back so an operator can tell which + // connection is stored without being shown the credential. + if got.Data.Host != "db.internal:5432" || got.Data.Database != "cryden_ro" { + t.Errorf("host = %q, database = %q, want the two harmless halves back", got.Data.Host, got.Data.Database) + } + if !got.Data.DSNSet || !got.Data.Configured || got.Data.Label != "reporting replica" { + t.Errorf("dsn_set = %v, configured = %v, label = %q", got.Data.DSNSet, got.Data.Configured, got.Data.Label) + } +} + +func TestAskAIWidgetRoundTrips(t *testing.T) { + f := newSettingsFixture(t, testSettingsKey) + + body := `{ + "enabled": true, + "allowed_origins": ["https://console.example.com"], + "entities": ["sessions", "audit_events"], + "greeting": "Ask about your account", + "placeholder": "When did I last log in?" + }` + if rec := f.do(t, http.MethodPut, widgetPath, f.adminToken, body); rec.Code != http.StatusOK { + t.Fatalf("PUT: status = %d (body %s)", rec.Code, rec.Body.String()) + } + + rec := f.do(t, http.MethodGet, widgetPath, f.adminToken, "") + if rec.Code != http.StatusOK { + t.Fatalf("GET: status = %d (body %s)", rec.Code, rec.Body.String()) + } + + var got struct { + Data settings.AskAIWidgetConfig `json:"data"` + } + if err := json.Unmarshal(rec.Body.Bytes(), &got); err != nil { + t.Fatalf("decoding the response: %v", err) + } + if !got.Data.Enabled { + t.Error("enabled came back false") + } + if len(got.Data.Entities) != 2 || got.Data.Entities[0] != "sessions" { + t.Errorf("entities = %v, want the two that were saved, in order", got.Data.Entities) + } + if got.Data.Greeting != "Ask about your account" { + t.Errorf("greeting = %q", got.Data.Greeting) + } +} + +// Before anything is saved the widget reads as off, which is what the +// zero value means — a console shows the feature as disabled rather than +// showing a form that looks as though someone filled it in. +func TestAskAIWidgetBeforeAnySaveIsDisabled(t *testing.T) { + f := newSettingsFixture(t, testSettingsKey) + + rec := f.do(t, http.MethodGet, widgetPath, f.adminToken, "") + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200 (body %s)", rec.Code, rec.Body.String()) + } + if !strings.Contains(rec.Body.String(), `"enabled":false`) { + t.Errorf("body = %s, want enabled:false", rec.Body.String()) + } +} + +// The scope list is the one part of this setting with teeth, so what it +// refuses matters as much as what it accepts. +func TestPutAskAIWidgetRefusesAnOutOfScopeConfiguration(t *testing.T) { + f := newSettingsFixture(t, testSettingsKey) + + cases := []struct{ name, body string }{ + { + "entity outside cryden's allowlist", + `{"enabled":true,"allowed_origins":["https://c.example.com"],"entities":["password_hashes"],"greeting":"hi"}`, + }, + { + "wildcard origin", + `{"enabled":true,"allowed_origins":["*"],"entities":["sessions"],"greeting":"hi"}`, + }, + { + "origin with a path", + `{"enabled":true,"allowed_origins":["https://c.example.com/widget"],"entities":["sessions"],"greeting":"hi"}`, + }, + { + "enabled with no entities", + `{"enabled":true,"allowed_origins":["https://c.example.com"],"greeting":"hi"}`, + }, + { + "enabled with no origins", + `{"enabled":true,"entities":["sessions"],"greeting":"hi"}`, + }, + { + "enabled with no greeting", + `{"enabled":true,"allowed_origins":["https://c.example.com"],"entities":["sessions"]}`, + }, + { + "duplicate entity", + `{"enabled":true,"allowed_origins":["https://c.example.com"],"entities":["sessions","sessions"],"greeting":"hi"}`, + }, + } + + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + rec := f.do(t, http.MethodPut, widgetPath, f.adminToken, tc.body) + if rec.Code != http.StatusBadRequest { + t.Fatalf("status = %d, want 400 (body %s)", rec.Code, rec.Body.String()) + } + if !strings.Contains(rec.Body.String(), "invalid_ask_ai_widget") { + t.Errorf("body = %s, want invalid_ask_ai_widget", rec.Body.String()) + } + }) + } + + if _, ok := f.raw(t, settings.KeyAskAIWidget); ok { + t.Error("a rejected configuration reached the store") + } +} + +// Switching the widget off must be savable without filling in the fields +// that are about to stop mattering. +func TestPutAskAIWidgetAcceptsADisabledEmptyConfiguration(t *testing.T) { + f := newSettingsFixture(t, testSettingsKey) + + rec := f.do(t, http.MethodPut, widgetPath, f.adminToken, `{"enabled":false}`) + if rec.Code != http.StatusOK { + t.Fatalf("status = %d, want 200 (body %s)", rec.Code, rec.Body.String()) + } +} + +func TestDeleteAskAIWidgetLeavesItDisabled(t *testing.T) { + f := newSettingsFixture(t, testSettingsKey) + + body := `{"enabled":true,"allowed_origins":["https://c.example.com"],"entities":["sessions"],"greeting":"hi"}` + if rec := f.do(t, http.MethodPut, widgetPath, f.adminToken, body); rec.Code != http.StatusOK { + t.Fatalf("PUT: status = %d (body %s)", rec.Code, rec.Body.String()) + } + if rec := f.do(t, http.MethodDelete, widgetPath, f.adminToken, ""); rec.Code != http.StatusOK { + t.Fatalf("DELETE: status = %d (body %s)", rec.Code, rec.Body.String()) + } + + rec := f.do(t, http.MethodGet, widgetPath, f.adminToken, "") + if rec.Code != http.StatusOK { + t.Fatalf("GET: status = %d (body %s)", rec.Code, rec.Body.String()) + } + if !strings.Contains(rec.Body.String(), `"enabled":false`) { + t.Errorf("body = %s, want the widget back to disabled", rec.Body.String()) + } +} + +// A malformed body is a 400 with a message about the body, not an invalid +// *configuration* — the two are different mistakes and an operator fixing +// one should not be told about the other. +func TestSettingsEndpointsRejectAMalformedBody(t *testing.T) { + f := newSettingsFixture(t, testSettingsKey) + + for _, path := range []string{llmProviderPath, databaseProviderPath, widgetPath} { + t.Run(path, func(t *testing.T) { + rec := f.do(t, http.MethodPut, path, f.adminToken, `{"enabled":`) + if rec.Code != http.StatusBadRequest { + t.Fatalf("status = %d, want 400 (body %s)", rec.Code, rec.Body.String()) + } + if !strings.Contains(rec.Body.String(), "malformed request body") { + t.Errorf("body = %s, want the malformed-body message", rec.Body.String()) + } + }) + } +} diff --git a/main.go b/main.go index 9b05914..41113cd 100644 --- a/main.go +++ b/main.go @@ -20,6 +20,7 @@ import ( "github.com/crydensync/api/digest" "github.com/crydensync/api/httpapi" "github.com/crydensync/api/operator" + "github.com/crydensync/api/settings" "github.com/crydensync/api/shiplog" "github.com/crydensync/api/templates" "github.com/crydensync/api/usermeta" @@ -96,6 +97,21 @@ func main() { digestStore = digest.NewStore(db) } + // The AI settings: which LLM provider and which read-only database back + // the AI-assisted admin features. Always constructed — NewSecrets + // treats an unset SETTINGS_ENCRYPTION_KEY as "this feature is off" + // rather than as a startup failure, and the handlers then answer 404 + // not_configured, matching how every other optional feature in this api + // behaves. Only the key being *malformed* is fatal, and that is a + // configuration mistake worth refusing to boot on. + settingsSecrets, err := settings.NewSecrets(settings.NewStore(db), cfg.SettingsEncryptionKey) + if err != nil { + log.Fatalf("invalid SETTINGS_ENCRYPTION_KEY: %v", err) + } + if settingsSecrets.Configured() { + log.Printf("AI settings endpoints enabled (llm-provider, database-provider)") + } + engineCfg := cryden.Config{ JWTSecret: cfg.JWTSecret, Users: users, @@ -355,6 +371,8 @@ func main() { Shipped: shippedLog, Digests: digestStore, + + Settings: settingsSecrets, }) limiter := httpapi.NewEdgeRateLimiter(cfg.EdgeRateLimit, cfg.EdgeRateLimitWindow) handler := httpapi.WithCORS(cfg.CORSOrigins, httpapi.WithEdgeRateLimit(limiter, router)) From 10cea8d1ad80988b2ea7dc51ceb28e2d955a816f Mon Sep 17 00:00:00 2001 From: Raymond Nicholas Date: Wed, 16 Sep 2026 00:12:19 +0100 Subject: [PATCH 10/12] feat: add the ask-ai widget settings and its entity scope The widget config endpoint stores which entities the widget answers over, and aiprovider.ScopedProvider is what makes that setting real: cryden's widget.Ask scopes every intent to the calling user but does so over all of ai.AllowedEntities, so narrowing that is a host decision and belongs here. Co-Authored-By: Claude Code --- aiprovider/scoped.go | 83 +++++++++++++++ aiprovider/scoped_test.go | 102 +++++++++++++++++++ settings/widget.go | 209 ++++++++++++++++++++++++++++++++++++++ settings/widget_test.go | 194 +++++++++++++++++++++++++++++++++++ 4 files changed, 588 insertions(+) create mode 100644 aiprovider/scoped.go create mode 100644 aiprovider/scoped_test.go create mode 100644 settings/widget.go create mode 100644 settings/widget_test.go diff --git a/aiprovider/scoped.go b/aiprovider/scoped.go new file mode 100644 index 0000000..6a921cf --- /dev/null +++ b/aiprovider/scoped.go @@ -0,0 +1,83 @@ +package aiprovider + +import ( + "context" + "errors" + + crydenai "github.com/crydensync/cryden/v2/ai" +) + +// ErrEntityOutOfScope means the parsed intent named an entity the +// deployment has not made available to the ask-ai widget. +var ErrEntityOutOfScope = errors.New("aiprovider: entity is outside the configured ask-ai scope") + +// ScopedProvider narrows an ai.LLMProvider to a configured set of +// entities. +// +// It sits in front of the provider the widget uses, not in front of the +// admin one. cryden's widget.Ask already forces every intent to the +// calling end user's own rows — it discards whatever identity filter the +// model produced and substitutes the real one — and that is the security +// boundary, which this type does not touch or replace. What it adds is +// the layer above: which entities a deployment is willing to answer +// questions about at all, from a public-facing surface. +// +// The distinction matters because the two are decided by different +// people. cryden decides what can be scoped safely; the operator decides +// what this deployment offers. An operator who wants the widget to answer +// "when did I last log in" but not "what has been recorded against me" +// has no way to say so through cryden's fixed allowlist, and asking +// cryden to grow a per-host policy knob would be putting a host decision +// in the engine — the boundary CLAUDE.md draws. +// +// Wrapping ParseQueryIntent is the only place this can be done. By the +// time widget.Ask has an Answer, the query has already run, and the +// intent itself never leaves the package: Ask parses, scopes and executes +// in one call. Refusing at parse time is the one point where the entity +// is still visible to the host and nothing has been executed yet. +type ScopedProvider struct { + inner crydenai.LLMProvider + entities map[string]bool +} + +// NewScopedProvider wraps inner, refusing any entity not in entities. +// +// An empty entities set produces a provider that refuses everything, +// which is the correct reading of "no scope configured": a settings form +// that was never filled in should answer no questions rather than all of +// them. The caller is expected to check the widget's own Enabled flag +// before it gets this far; this is the second lock on the same door. +func NewScopedProvider(inner crydenai.LLMProvider, entities []string) *ScopedProvider { + allowed := make(map[string]bool, len(entities)) + for _, entity := range entities { + allowed[entity] = true + } + return &ScopedProvider{inner: inner, entities: allowed} +} + +var _ crydenai.LLMProvider = (*ScopedProvider)(nil) + +// ParseQueryIntent defers to the wrapped provider and then refuses an +// entity outside the configured scope. +// +// The refusal happens after the parse rather than before it, because the +// entity is the model's output and does not exist until then. That costs +// a model call on a question that will be rejected, which is the honest +// price: the alternative would be a second model call asking the model to +// classify its own question first, which is more expensive, less +// reliable, and still untrusted input. +// +// The error names neither the entity nor the scope. It reaches an end +// user through the widget, and telling them which entities this +// deployment does have configured would be describing the console's +// schema to whoever is typing questions at it. +func (p *ScopedProvider) ParseQueryIntent(ctx context.Context, naturalLanguage string) (crydenai.QueryIntent, error) { + intent, err := p.inner.ParseQueryIntent(ctx, naturalLanguage) + if err != nil { + return crydenai.QueryIntent{}, err + } + if !p.entities[intent.Entity] { + return crydenai.QueryIntent{}, ErrEntityOutOfScope + } + return intent, nil +} diff --git a/aiprovider/scoped_test.go b/aiprovider/scoped_test.go new file mode 100644 index 0000000..c2c082d --- /dev/null +++ b/aiprovider/scoped_test.go @@ -0,0 +1,102 @@ +package aiprovider + +import ( + "context" + "errors" + "strings" + "testing" + + crydenai "github.com/crydensync/cryden/v2/ai" +) + +// fixedProvider returns whatever intent it was built with, which is what +// makes it usable as the inner provider here: the test controls exactly +// what the model "produced" without any network call. +type fixedProvider struct { + intent crydenai.QueryIntent + err error + + calls int +} + +func (p *fixedProvider) ParseQueryIntent(context.Context, string) (crydenai.QueryIntent, error) { + p.calls++ + if p.err != nil { + return crydenai.QueryIntent{}, p.err + } + return p.intent, nil +} + +func TestScopedProviderPassesThroughAConfiguredEntity(t *testing.T) { + inner := &fixedProvider{intent: crydenai.QueryIntent{Entity: "sessions"}} + scoped := NewScopedProvider(inner, []string{"sessions", "audit_events"}) + + intent, err := scoped.ParseQueryIntent(context.Background(), "when did I last log in") + if err != nil { + t.Fatalf("ParseQueryIntent: %v", err) + } + if intent.Entity != "sessions" { + t.Errorf("entity = %q, want it passed through", intent.Entity) + } +} + +// The whole point of the type: an entity the deployment has not made +// available is refused, even though cryden's own allowlist permits it and +// widget.Ask would happily scope it. +func TestScopedProviderRefusesAnEntityOutsideTheConfiguredScope(t *testing.T) { + inner := &fixedProvider{intent: crydenai.QueryIntent{Entity: "audit_events"}} + scoped := NewScopedProvider(inner, []string{"sessions"}) + + _, err := scoped.ParseQueryIntent(context.Background(), "what has been recorded against me") + if !errors.Is(err, ErrEntityOutOfScope) { + t.Fatalf("error = %v, want ErrEntityOutOfScope", err) + } +} + +// A scope that was never configured answers nothing rather than +// everything. A settings form nobody filled in is not permission. +func TestScopedProviderWithNoEntitiesRefusesEverything(t *testing.T) { + for _, entities := range [][]string{nil, {}} { + scoped := NewScopedProvider(&fixedProvider{intent: crydenai.QueryIntent{Entity: "users"}}, entities) + + _, err := scoped.ParseQueryIntent(context.Background(), "who am I") + if !errors.Is(err, ErrEntityOutOfScope) { + t.Errorf("scope %v: error = %v, want ErrEntityOutOfScope", entities, err) + } + } +} + +// The inner provider's own failure is passed through unchanged: this +// wrapper narrows a scope, it does not reinterpret a model call that +// failed. +func TestScopedProviderPassesThroughTheInnerError(t *testing.T) { + sentinel := errors.New("the provider refused") + scoped := NewScopedProvider(&fixedProvider{err: sentinel}, []string{"sessions"}) + + _, err := scoped.ParseQueryIntent(context.Background(), "anything") + if !errors.Is(err, sentinel) { + t.Errorf("error = %v, want the inner provider's error", err) + } +} + +// The refusal has to name nothing. It reaches an end user through the +// widget, and an error that listed the deployment's configured entities +// would be describing the console's schema to whoever is typing +// questions at it. +func TestScopedProviderRefusalNamesNeitherEntityNorScope(t *testing.T) { + scoped := NewScopedProvider( + &fixedProvider{intent: crydenai.QueryIntent{Entity: "audit_events"}}, + []string{"sessions"}, + ) + + _, err := scoped.ParseQueryIntent(context.Background(), "anything") + if err == nil { + t.Fatal("nothing was refused") + } + message := err.Error() + for _, leak := range []string{"audit_events", "sessions"} { + if strings.Contains(message, leak) { + t.Errorf("error = %q, want it to name neither the refused entity nor the configured scope", message) + } + } +} diff --git a/settings/widget.go b/settings/widget.go new file mode 100644 index 0000000..939e9af --- /dev/null +++ b/settings/widget.go @@ -0,0 +1,209 @@ +package settings + +import ( + "encoding/json" + "errors" + "fmt" + "net/url" + "strings" + + crydenai "github.com/crydensync/cryden/v2/ai" +) + +// ErrInvalidAskAIWidget means the supplied widget configuration cannot be +// stored: a scope that names something the query surface cannot bound, an +// origin that is not an origin, or copy too long for a settings form. +var ErrInvalidAskAIWidget = errors.New("settings: invalid ask-ai widget configuration") + +// Bounds on the stored copy and the two lists. Small on purpose: this is +// text a widget renders inside a chat bubble and a list an operator ticks +// by hand, so a value beyond these is a mistake or an attempt to use the +// settings table as storage rather than a configuration anyone wants. +const ( + maxWidgetCopyLength = 200 + maxWidgetOrigins = 20 + maxWidgetEntities = 8 +) + +// AskAIWidgetConfig is what an operator configures about the ask-ai +// widget: whether it is offered, which origins may embed it, which +// entities it will answer over, and the copy it shows. +// +// This is the one setting in this package that is NOT a credential, and +// that is why it has no Redacted counterpart — the other two would put an +// API key and a database password in a response if a handler reached for +// the wrong type, and this one is safe to return exactly as stored. The +// distinction is worth keeping visible rather than papering over with a +// Redacted method that returns a copy of itself. +// +// What is genuinely load-bearing here is Entities, and it is not +// decorative: cryden's widget.Ask force-scopes every parsed intent to the +// calling end user's own rows, but it scopes over the whole of +// ai.AllowedEntities. Narrowing that further is a host decision, so this +// repo enforces the configured subset in front of the provider — see +// aiprovider.ScopedProvider. A "scope" setting that nothing consulted +// would be worse than no setting at all. +type AskAIWidgetConfig struct { + // Enabled is whether the widget is offered to end users at all. A + // disabled widget is a 404 from whatever serves it, not an empty + // answer. + Enabled bool `json:"enabled"` + // AllowedOrigins are the sites permitted to embed the widget, as + // origins ("https://console.example.com"), never as patterns. + AllowedOrigins []string `json:"allowed_origins"` + // Entities are the ai.AllowedEntities the widget will answer over. + // Always a subset of cryden's own allowlist; cryden's list is the + // vocabulary and this is the host narrowing it. + Entities []string `json:"entities"` + // Greeting and Placeholder are the copy the embed shows. + Greeting string `json:"greeting"` + Placeholder string `json:"placeholder"` +} + +// Validate checks a configuration about to be stored. +// +// A disabled widget is allowed to be otherwise empty: an operator +// switching the feature off should be able to save the form without first +// filling in fields that are about to stop mattering. Everything else is +// checked in full, because "disabled" is the only state where an unset +// scope is not a promise the widget would fail to keep. +func (c AskAIWidgetConfig) Validate() error { + if len(c.Greeting) > maxWidgetCopyLength { + return fmt.Errorf("%w: greeting is longer than %d characters", ErrInvalidAskAIWidget, maxWidgetCopyLength) + } + if len(c.Placeholder) > maxWidgetCopyLength { + return fmt.Errorf("%w: placeholder is longer than %d characters", ErrInvalidAskAIWidget, maxWidgetCopyLength) + } + + if !c.Enabled { + // Still checked, so that a form which fills in the fields and then + // ticks the box cannot store something that was never validated + // and would be rejected the moment it was switched on. + if err := validateWidgetEntities(c.Entities); err != nil { + return err + } + return validateWidgetOrigins(c.AllowedOrigins) + } + + if strings.TrimSpace(c.Greeting) == "" { + return fmt.Errorf("%w: greeting is required when the widget is enabled", ErrInvalidAskAIWidget) + } + if err := validateWidgetEntities(c.Entities); err != nil { + return err + } + if len(c.Entities) == 0 { + return fmt.Errorf("%w: at least one entity is required when the widget is enabled — the query surface has nothing to answer over otherwise", ErrInvalidAskAIWidget) + } + if err := validateWidgetOrigins(c.AllowedOrigins); err != nil { + return err + } + if len(c.AllowedOrigins) == 0 { + return fmt.Errorf("%w: at least one allowed origin is required when the widget is enabled", ErrInvalidAskAIWidget) + } + return nil +} + +// validateWidgetEntities refuses an entity cryden does not allowlist. +// +// The check is against cryden's own map rather than a list restated here, +// for the same reason aiprovider.intentSchema builds its enums from those +// maps: a second copy of the vocabulary is a second thing to forget to +// update, and the failure mode of a stale copy is silent — a valid entity +// refused, or a refused one stored. +// +// One asymmetry is accepted knowingly. cryden's widget.scopeToOwner is a +// private switch over today's three entities, so an entity added to +// ai.AllowedEntities without a matching case there would pass here and +// then fail at Ask time with ErrEntityNotAvailable. That is the safe +// direction: the widget refuses the question rather than answering it +// unscoped, which is exactly the fail-closed default that package +// documents. +func validateWidgetEntities(entities []string) error { + if len(entities) > maxWidgetEntities { + return fmt.Errorf("%w: at most %d entities may be selected, got %d", + ErrInvalidAskAIWidget, maxWidgetEntities, len(entities)) + } + seen := make(map[string]bool, len(entities)) + for _, entity := range entities { + if !crydenai.AllowedEntities[entity] { + return fmt.Errorf("%w: %q is not an entity the AI query surface may touch", ErrInvalidAskAIWidget, entity) + } + if seen[entity] { + return fmt.Errorf("%w: %q is listed twice", ErrInvalidAskAIWidget, entity) + } + seen[entity] = true + } + return nil +} + +// validateWidgetOrigins checks that each entry is an origin and nothing +// more. A path, a query or an embedded credential in this list would be +// silently ignored by any origin comparison, so accepting one would be +// storing a value that does not mean what the operator who typed it +// thinks it means. +func validateWidgetOrigins(origins []string) error { + if len(origins) > maxWidgetOrigins { + return fmt.Errorf("%w: at most %d allowed origins may be listed, got %d", + ErrInvalidAskAIWidget, maxWidgetOrigins, len(origins)) + } + for _, origin := range origins { + if err := validateWidgetOrigin(origin); err != nil { + return err + } + } + return nil +} + +func validateWidgetOrigin(origin string) error { + if strings.TrimSpace(origin) != origin || origin == "" { + return fmt.Errorf("%w: an allowed origin must be non-empty and have no surrounding whitespace", ErrInvalidAskAIWidget) + } + // A wildcard is refused by name so the message can say why, rather + // than falling through to "not a parseable URL" and reading as a + // formatting nit. It is not a formatting nit: this widget answers + // questions about the signed-in end user's own sessions and audit + // events, so an origin list of "*" means any page on the internet can + // ask those questions through a visitor's browser. + if origin == "*" { + return fmt.Errorf("%w: \"*\" is not accepted — list the origins that may embed the widget, because this one answers questions about the signed-in user", ErrInvalidAskAIWidget) + } + + parsed, err := url.Parse(origin) + if err != nil { + return fmt.Errorf("%w: %q is not a parseable origin: %v", ErrInvalidAskAIWidget, origin, err) + } + if parsed.Scheme != "http" && parsed.Scheme != "https" { + return fmt.Errorf("%w: %q has scheme %q, want http or https", ErrInvalidAskAIWidget, origin, parsed.Scheme) + } + if parsed.Host == "" { + return fmt.Errorf("%w: %q names no host", ErrInvalidAskAIWidget, origin) + } + // "/" is what url.Parse produces for "https://x.example.com/" and is + // the same origin, so it is allowed; anything else is a path. + if parsed.Path != "" && parsed.Path != "/" { + return fmt.Errorf("%w: %q carries a path — an origin is scheme, host and port only", ErrInvalidAskAIWidget, origin) + } + if parsed.RawQuery != "" || parsed.Fragment != "" || parsed.User != nil { + return fmt.Errorf("%w: %q carries a query, fragment or credentials — an origin is scheme, host and port only", ErrInvalidAskAIWidget, origin) + } + return nil +} + +// MarshalAskAIWidget and UnmarshalAskAIWidget are the encode/decode pair +// for storage, the same shape as the other two settings so "JSON, then +// sealed" has one spelling per setting. Unlike those two the payload is +// not a credential, but it is stored through the same Secrets wrapper and +// so is sealed by the same key — one storage path with one rule about +// what reaches the table is worth more than saving a decryption on a +// value nobody reads often. +func MarshalAskAIWidget(c AskAIWidgetConfig) ([]byte, error) { + return json.Marshal(c) +} + +func UnmarshalAskAIWidget(plaintext []byte) (AskAIWidgetConfig, error) { + var c AskAIWidgetConfig + if err := json.Unmarshal(plaintext, &c); err != nil { + return AskAIWidgetConfig{}, fmt.Errorf("settings: decoding the stored ask-ai widget config: %w", err) + } + return c, nil +} diff --git a/settings/widget_test.go b/settings/widget_test.go new file mode 100644 index 0000000..c0a32c0 --- /dev/null +++ b/settings/widget_test.go @@ -0,0 +1,194 @@ +package settings + +import ( + "errors" + "strings" + "testing" + + crydenai "github.com/crydensync/cryden/v2/ai" +) + +// The accepting cases, including the two shapes that look wrong and are +// not: a trailing slash is the same origin, and a port is part of one. +func TestValidateWidgetOriginAcceptsRealOrigins(t *testing.T) { + for _, origin := range []string{ + "https://console.example.com", + "https://console.example.com/", + "https://console.example.com:8443", + "http://localhost:3000", + } { + if err := validateWidgetOrigin(origin); err != nil { + t.Errorf("validateWidgetOrigin(%q) = %v, want nil", origin, err) + } + } +} + +// Every one of these would be silently ignored by an origin comparison, +// so accepting one would store a value that does not mean what the +// operator who typed it thinks it means. +func TestValidateWidgetOriginRefusesAnythingThatIsNotAnOrigin(t *testing.T) { + cases := []struct{ name, origin string }{ + {"wildcard", "*"}, + {"empty", ""}, + {"surrounding whitespace", " https://console.example.com"}, + {"a path", "https://console.example.com/widget"}, + {"a query", "https://console.example.com?tenant=1"}, + {"a fragment", "https://console.example.com#top"}, + {"embedded credentials", "https://user:pw@console.example.com"}, + {"no scheme", "console.example.com"}, + {"a non-http scheme", "ftp://console.example.com"}, + {"no host", "https://"}, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + if err := validateWidgetOrigin(tc.origin); err == nil { + t.Errorf("validateWidgetOrigin(%q) was accepted, want a refusal", tc.origin) + } else if !errors.Is(err, ErrInvalidAskAIWidget) { + t.Errorf("error = %v, want it to wrap ErrInvalidAskAIWidget", err) + } + }) + } +} + +// The wildcard gets its own message, because "not a parseable origin" +// would read as a formatting nit and this is not one. +func TestWildcardOriginExplainsItself(t *testing.T) { + err := validateWidgetOrigin("*") + if err == nil { + t.Fatal("the wildcard origin was accepted") + } + if !strings.Contains(err.Error(), "signed-in user") { + t.Errorf("error = %q, want it to say why a wildcard is refused", err) + } +} + +// The entity list is checked against cryden's own map rather than a copy +// of it, so every entity the engine allowlists has to pass here. A +// restated list would drift, and this is what would catch it. +func TestValidateWidgetEntitiesAcceptsEverythingCrydenAllowlists(t *testing.T) { + for entity := range crydenai.AllowedEntities { + if err := validateWidgetEntities([]string{entity}); err != nil { + t.Errorf("validateWidgetEntities(%q) = %v, want nil — cryden allowlists it", entity, err) + } + } +} + +func TestValidateWidgetEntitiesRefusesUnknownAndRepeated(t *testing.T) { + if err := validateWidgetEntities([]string{"password_hashes"}); err == nil { + t.Error("an entity cryden does not allowlist was accepted") + } + if err := validateWidgetEntities([]string{"sessions", "sessions"}); err == nil { + t.Error("a repeated entity was accepted") + } + if err := validateWidgetEntities([]string{""}); err == nil { + t.Error("an empty entity was accepted") + } +} + +// A disabled widget may be otherwise empty — an operator switching the +// feature off should not have to fill in fields that are about to stop +// mattering. Anything that IS filled in is still validated, so a form +// cannot store a value that was never checked and would be rejected the +// moment it was switched on. +func TestValidateAcceptsADisabledEmptyWidgetButNotADisabledInvalidOne(t *testing.T) { + if err := (AskAIWidgetConfig{Enabled: false}).Validate(); err != nil { + t.Errorf("an empty disabled configuration = %v, want nil", err) + } + if err := (AskAIWidgetConfig{Enabled: false, Entities: []string{"password_hashes"}}).Validate(); err == nil { + t.Error("a disabled configuration carrying an out-of-scope entity was accepted") + } + if err := (AskAIWidgetConfig{Enabled: false, AllowedOrigins: []string{"*"}}).Validate(); err == nil { + t.Error("a disabled configuration carrying a wildcard origin was accepted") + } +} + +// An enabled widget has to say what it is and who may embed it; the zero +// values of those fields are not usable answers. +func TestValidateRequiresTheEnabledWidgetToBeComplete(t *testing.T) { + complete := AskAIWidgetConfig{ + Enabled: true, + AllowedOrigins: []string{"https://console.example.com"}, + Entities: []string{"sessions"}, + Greeting: "Ask about your account", + } + if err := complete.Validate(); err != nil { + t.Fatalf("a complete configuration = %v, want nil", err) + } + + for _, tc := range []struct { + name string + mutate func(*AskAIWidgetConfig) + }{ + {"no entities", func(c *AskAIWidgetConfig) { c.Entities = nil }}, + {"no origins", func(c *AskAIWidgetConfig) { c.AllowedOrigins = nil }}, + {"blank greeting", func(c *AskAIWidgetConfig) { c.Greeting = " " }}, + {"greeting too long", func(c *AskAIWidgetConfig) { c.Greeting = strings.Repeat("x", maxWidgetCopyLength+1) }}, + {"placeholder too long", func(c *AskAIWidgetConfig) { c.Placeholder = strings.Repeat("x", maxWidgetCopyLength+1) }}, + } { + t.Run(tc.name, func(t *testing.T) { + config := complete + tc.mutate(&config) + if err := config.Validate(); err == nil { + t.Error("accepted, want a refusal") + } + }) + } +} + +// The two lists are bounded so a settings form cannot be used as storage. +func TestValidateBoundsTheWidgetLists(t *testing.T) { + tooManyOrigins := AskAIWidgetConfig{Enabled: true, Entities: []string{"sessions"}, Greeting: "hi"} + for i := 0; i <= maxWidgetOrigins; i++ { + tooManyOrigins.AllowedOrigins = append(tooManyOrigins.AllowedOrigins, "https://c.example.com") + } + if err := tooManyOrigins.Validate(); err == nil { + t.Errorf("accepted %d origins, want at most %d", len(tooManyOrigins.AllowedOrigins), maxWidgetOrigins) + } + + // Distinct entities, because the duplicate rule would fire first. + tooManyEntities := AskAIWidgetConfig{ + Enabled: true, + AllowedOrigins: []string{"https://console.example.com"}, + Greeting: "hi", + } + for entity := range crydenai.AllowedEntities { + tooManyEntities.Entities = append(tooManyEntities.Entities, entity) + } + for len(tooManyEntities.Entities) <= maxWidgetEntities { + tooManyEntities.Entities = append(tooManyEntities.Entities, "sessions"+strings.Repeat("x", len(tooManyEntities.Entities))) + } + if err := tooManyEntities.Validate(); err == nil { + t.Errorf("accepted %d entities, want at most %d", len(tooManyEntities.Entities), maxWidgetEntities) + } +} + +func TestAskAIWidgetMarshalRoundTrip(t *testing.T) { + config := AskAIWidgetConfig{ + Enabled: true, + AllowedOrigins: []string{"https://console.example.com"}, + Entities: []string{"sessions", "audit_events"}, + Greeting: "Ask about your account", + Placeholder: "When did I last log in?", + } + + encoded, err := MarshalAskAIWidget(config) + if err != nil { + t.Fatalf("marshalling: %v", err) + } + decoded, err := UnmarshalAskAIWidget(encoded) + if err != nil { + t.Fatalf("unmarshalling: %v", err) + } + if len(decoded.Entities) != 2 || decoded.Entities[0] != "sessions" { + t.Errorf("entities = %v, want them back in order", decoded.Entities) + } + if decoded.Greeting != config.Greeting || decoded.Placeholder != config.Placeholder { + t.Errorf("copy did not survive the round trip: %+v", decoded) + } +} + +func TestUnmarshalAskAIWidgetRefusesGarbage(t *testing.T) { + if _, err := UnmarshalAskAIWidget([]byte("not json")); err == nil { + t.Error("garbage was accepted as a stored configuration") + } +} From 4d5d03fd7a47c853f53a59ecbb6d625bd5178ee7 Mon Sep 17 00:00:00 2001 From: Raymond Nicholas Date: Wed, 16 Sep 2026 16:04:24 +0100 Subject: [PATCH 11/12] docs: spec the AI settings surface and write it up in the README MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit openapi goes to 1.5 — the first version with admin endpoints that write, and the first whose responses depend on a non-cryden deployment setting. Three paths, five schemas, and a README section on why the read-only database is checked by attempting a write rather than by a checkbox. Co-Authored-By: Claude Code --- README.md | 46 +++- openapi/spec.yaml | 567 +++++++++++++++++++++++++++++++++++++++++++++- 2 files changed, 610 insertions(+), 3 deletions(-) diff --git a/README.md b/README.md index 4dbedee..21a0cc4 100644 --- a/README.md +++ b/README.md @@ -370,6 +370,48 @@ Setting `DIGEST_INTERVAL_HOURS` (168 is weekly) turns on the schedule: a backgro - A failed run is **logged and swallowed**. This runs in a goroutine with nobody to hand an error to, and a scheduler that stopped at the first database blip would silently stop producing digests for the rest of the process's life. - Nothing on the HTTP surface can create a digest run. Only the scheduler writes, and it is a process component rather than a request handler — the read-only rule the whole admin surface follows. +## AI provider settings + +The AI-assisted admin features need two things only this repo can supply, because cryden defines them as interfaces the host implements: an `ai.LLMProvider` that turns a question into a `QueryIntent`, and an `ai.QueryableStore` that runs one. Three settings endpoints configure them, all operator-only: + +``` +GET|PUT|DELETE /v1/admin/settings/llm-provider +GET|PUT|DELETE /v1/admin/settings/database-provider +GET|PUT|DELETE /v1/admin/settings/ask-ai-widget +``` + +These are the admin surface's **only** writes, and they are the other half of the read-only rule rather than a hole in it. A tuning suggestion pre-fills one of these forms; an operator presses save; this is what handles that save. No AI-assisted handler in this repo holds a reference to any of them, and none accepts a suggestion as input. + +All three answer `404 not_configured` when `SETTINGS_ENCRYPTION_KEY` is unset — without a key there is nowhere safe to put a credential, so the API refuses rather than storing one in the clear. + +### Credentials + +The LLM API key and the database connection string are sealed with **AES-256-GCM before they reach the table**, keyed from `SETTINGS_ENCRYPTION_KEY`. Treat it like `JWT_SECRET`: set it, keep it out of source control, and expect a rotation to need the old value for as long as rows written under it exist. + +- **Separate from cryden's `ENCRYPTION_KEY` on purpose.** The two seal different things with different lifetimes — cryden's covers what the engine stores (TOTP secrets), this one covers what the API stores — so one leaking or rotating need not touch the other. Same reasoning as `CLOUD_LOG_HASH_KEY`. +- **Never returned, in any form.** Not masked, not truncated to the last four characters: `api_key_set` and `dsn_set` booleans are what a console renders "saved" from, and returning any part of the value would put it in a browser's memory and a devtools panel. `GET` on the database provider returns the host and database name only, so an operator can tell which connection is stored without being shown a password. +- **A changed key is `409 setting_undecryptable`, not `404`.** The row is still there, and reporting it as missing would send an operator to re-enter a credential that is fine. `DELETE` needs no key at all, which is what makes it the way out for a deployment that has lost one. +- **The credential is required on every `PUT`**, rather than optional with "blank means keep the existing one". That convention is the usual one and it is wrong here: an omitted field and a deliberately cleared one would be the same request, and getting it wrong means a form that appears to save a key while silently storing an empty one. + +### The read-only database requirement + +`PUT /v1/admin/settings/database-provider` **connects with the supplied credentials and attempts a write before storing anything.** cryden's own position is that for this feature "the credential boundary, not just the allowlist, is the real safety guarantee" — a role that cannot `INSERT` cannot `INSERT` whatever the query builder does with its input. So the order is: validate the shape, prove the role cannot write, and only then store. + +- **An attempted write, not a reading of the role's attributes.** A role with `rolsuper` set, or a connection string containing the word "readonly", or a "read-only?" checkbox in the console, are all claims. Only the server's refusal is evidence, and it is evidence about the actual role, on the actual database, through the actual credentials. It cannot be done client-side either — a browser cannot open a Postgres connection. +- **The probe writes to `pg_temp`**, the session's own temporary schema. The table lives only for the life of that connection and is dropped when it closes, so a probe that fails leaves nothing for an operator to clean up. The pool is capped at one connection so the `CREATE` and the `INSERT` share the session that owns the temp table — a pool that split them would have the `INSERT` fail on a missing table, which looks like a refusal and is not one. +- **Three outcomes, deliberately distinct.** A refused write is a pass. A successful write is `400 database_role_not_read_only`. Anything else — no connection, a timeout, a `CREATE` that failed for a reason other than privilege — is `400 database_role_unverified`, **which is not a pass**. Treating "could not find out" as success would make the check pass exactly when it is least able to tell. Only Postgres' own SQLSTATE `42501` (`insufficient_privilege`) counts as a refusal, and it is matched by code rather than by message, since the message is localized and reworded between major versions. +- **It is slower than its neighbours**, because it opens a connection and runs statements before answering, bounded by a ten-second timeout. That is paid once per save, not per query. + +### The ask-ai widget + +`GET|PUT /v1/admin/settings/ask-ai-widget` stores the widget's enabled flag, the origins allowed to embed it, the entities it answers over, and its copy. It is the one setting here that is **not** a credential, so there is nothing to redact. + +- **`entities` has teeth.** cryden's `widget.Ask` force-scopes every parsed intent to the calling end user's own rows — it discards whatever identity filter the model produced and substitutes the real one, rather than validating and rejecting, so there is no oracle — but it scopes over the whole of `ai.AllowedEntities`. Narrowing that further is a host decision, so `aiprovider.ScopedProvider` enforces the configured subset in front of the provider. A scope setting nothing consulted would be worse than no setting at all. The list is validated against cryden's own allowlist rather than a copy of it. +- **`"*"` as an origin is refused by name**, with the reason in the message: this widget answers questions about the signed-in user's sessions and audit events, so a wildcard origin would let any page on the internet ask them through a visitor's browser. +- **The refusal names neither the entity nor the scope.** That error reaches an end user through the widget; listing the configured entities would be describing the console's schema to whoever is typing questions at it. +- **A disabled widget may be otherwise empty**, so switching the feature off does not require filling in fields that are about to stop mattering. Anything that *is* filled in is still validated, so a form cannot store a value that was never checked and would be rejected the moment it was switched on. +- **No embed snippet is returned.** The snippet is markup the console renders into its own pages, and the URL in it would name an endpoint this API does not serve yet — returning one would hand the console a script tag pointing at a 404. What this endpoint owes the console is the configuration a snippet is built from. + ## Design notes - `CORS_ORIGINS` is required, no wildcard default — an API handling auth tokens should never allow every origin. @@ -379,8 +421,8 @@ Setting `DIGEST_INTERVAL_HOURS` (168 is weekly) turns on the schedule: a backgro - A paused login is a `200`, not an error: nothing failed, the caller just has one more step. `httpapi/second_factor.go` is the one place that response shape is written. - `DELETE /v1/passkeys/{credentialID}` takes a JSON body (`{"password": "..."}`) — the password is re-confirmation, so a stolen access token alone cannot weaken an account's own auth requirements. - Passkey ceremony options and the browser's credential response travel as raw JSON (an object, not a JSON-encoded string), since that is exactly what `navigator.credentials.create()`/`.get()` produce and consume. -- **Three of this repo's tables are not cryden's and never will be**: `user_metadata`, `webhook_deliveries`, `shipped_log_events`. cryden calls an interface and moves on; it keeps no queryable history of what a sender or a logger did, and `store.User` has no metadata concept on purpose. Each lives in its own package (`usermeta/`, `webhook/`, `shiplog/`) with a Postgres store and an in-memory double behind one interface, mirroring the `store/interfaces.go` + `store/memory` + `store/postgres` split cryden itself uses — which is what makes an endpoint over them testable with no database. -- **The admin surface is read-only by construction.** `GET /v1/admin/webhooks/deliveries` and `GET /v1/admin/logging/recent` report; neither offers a "retry this delivery" button, a "replay this event", or any way to write a log record or a delivery row. That is the same rule cryden's AI admin tools are built under, carried across the repo boundary: an operator reads the state of the system, and every change to it goes through the explicit path that owns that change (or through the receiving system, for a delivery). Adding a write here is a design change, not a convenience. +- **Five of this repo's tables are not cryden's and never will be**: `user_metadata`, `webhook_deliveries`, `shipped_log_events`, `digest_runs` and `settings`. cryden calls an interface and moves on; it keeps no queryable history of what a sender or a logger did, no schedule, no run record, and no configuration storage — and `store.User` has no metadata concept on purpose. Each lives in its own package (`usermeta/`, `webhook/`, `shiplog/`, `digest/`, `settings/`) with a Postgres store and an in-memory double behind one interface, mirroring the `store/interfaces.go` + `store/memory` + `store/postgres` split cryden itself uses — which is what makes an endpoint over them testable with no database. +- **The admin surface is read-only by construction, with one named exception.** `GET /v1/admin/webhooks/deliveries` and `GET /v1/admin/logging/recent` report; neither offers a "retry this delivery" button, a "replay this event", or any way to write a log record or a delivery row. That is the same rule cryden's AI admin tools are built under, carried across the repo boundary: an operator reads the state of the system, and every change to it goes through the explicit path that owns that change (or through the receiving system, for a delivery). Adding a write here is a design change, not a convenience. **The exception is `/v1/admin/settings/*`**, which is a settings save — the "a human still saves it" half of the pre-fill rule, not an action any AI tool can reach. Its credentials are encrypted at rest, and it is the only place in this API that stores one. If you are adding a write under `/v1/admin` that is not a settings save, the answer is no. - `webhook_deliveries.id` is a `BIGSERIAL` surrogate key rather than the natural key you might expect. The event id it corresponds to **can be empty** — cryden generates it with `crypto/rand` and deliberately delivers an event without one rather than dropping it — and a delivery log whose primary key could be blank is a log that loses exactly the rows you would most want to see. The engine's own id is recorded beside it as `event_id` and is used for the receiver's idempotency. - This repo has **no graceful shutdown**, and as of this tier that is a stated gap rather than an unnoticed one: `main.go` ends at `log.Fatal(http.ListenAndServe(...))`, so the webhook worker's context is never cancelled and the shipped-events sink has no flush-and-exit path. Both were built so that adding one later is a change to `main.go` alone — the worker takes a `context.Context`, which today is `context.Background()`. The sink writes synchronously for the same reason: a buffered sink with no shutdown path drops its last records on a crash. diff --git a/openapi/spec.yaml b/openapi/spec.yaml index 13db5af..7519775 100644 --- a/openapi/spec.yaml +++ b/openapi/spec.yaml @@ -1,7 +1,7 @@ openapi: 3.0.3 info: title: CrydenSync API - version: "1.4" + version: "1.5" description: > A self-hosted HTTP wrapper around the CrydenSync auth engine. Every response follows one of two envelope shapes: {"data": ...} @@ -39,6 +39,24 @@ info: nothing else and its suggestions pre-fill a settings field for a human to save rather than being applied, and only a background job this API runs writes the digest history. + + 1.5 is additive, and is the first version with admin endpoints that + WRITE: the AI provider settings (GET/PUT/DELETE + /admin/settings/llm-provider, /admin/settings/database-provider and + /admin/settings/ask-ai-widget). No existing path, field or status + code changed. The writes are the "human saves it" half of the + pre-fill rule 1.4 describes — a tuning suggestion pre-fills one of + these forms and an operator saves it — and they are the only place + a credential is stored. PUT database-provider additionally verifies + against the database itself that the supplied role cannot write, + and refuses to store the connection until the server has rejected + one; see that path's description for why the check is done there + and not in the console. + + 1.5 is also the first version whose responses depend on a + deployment setting that is not a cryden one: all three paths answer + 404 not_configured unless SETTINGS_ENCRYPTION_KEY is set, since + without that key there is nowhere safe to put a credential. servers: - url: http://localhost:8080/v1 description: Local dev @@ -403,6 +421,173 @@ components: has been applied, and nothing that reads this response can apply it either. + LLMProviderInput: + type: object + description: > + The PUT body for the LLM provider. Separate from what is stored + rather than the same shape, so the wire contract and the stored + record can diverge — and so a field added to the stored config + later is not silently accepted from a client that predates it. + required: [kind, model, api_key] + properties: + kind: + type: string + description: > + Which provider implementation to build. One value today: + "anthropic". + model: + type: string + description: > + The model id to call, e.g. "claude-opus-5". Validated for + shape, not against a list: this API does not decide which + models a deployment may use, and a list here would go stale + between SDK releases. + api_key: + type: string + description: > + The provider credential. Required on every write — see PUT + on this path for why "blank keeps the existing one" is + refused. Sealed before storage and never returned. + max_tokens: + type: integer + minimum: 256 + maximum: 8192 + description: > + Ceiling on a single model response. Defaults when omitted + (0), rather than failing validation. + + RedactedLLMProvider: + type: object + description: > + What GET, PUT and DELETE return. There is no field for the API + key in this schema and no mask of one — a console renders + "saved" from api_key_set. All fields are false/empty when + nothing is configured. + properties: + kind: + type: string + description: The stored provider kind, or empty when none is stored. + model: + type: string + description: The stored model id, or empty when none is stored. + max_tokens: + type: integer + description: The stored ceiling, or 0 when none is stored. + api_key_set: + type: boolean + description: Whether a credential is stored. Says nothing about whether it still works. + configured: + type: boolean + description: > + Whether anything is stored at all. A console branches on + this to decide between an empty form and a saved one. + + DatabaseProviderInput: + type: object + description: The PUT body for the read-only database. + required: [label, dsn] + properties: + label: + type: string + description: > + What a console shows instead of the connection string, so an + operator can tell two configurations apart without reading a + password. + dsn: + type: string + description: > + The full Postgres connection string, including the + credential. Either a postgres:// or postgresql:// URL or a + lib/pq keyword/value string carrying a host. Checked for + shape here, then used to connect and attempt a write before + anything is stored. + max_rows: + type: integer + minimum: 1 + maximum: 500 + description: > + Ceiling on rows from a single AI-driven query. Defaults when + omitted (0). Bounded by cryden's own maximum, so this cannot + become a setting that does not do what it appears to. + + RedactedDatabaseProvider: + type: object + description: > + What GET, PUT and DELETE return. The DSN has no field here at + all, so a handler cannot serialise it by reaching for the wrong + type. Host and database come back because an operator needs them + to recognise which connection is stored; the user and password + never do. + properties: + label: + type: string + description: The operator-supplied name for this connection. + max_rows: + type: integer + description: The stored row ceiling. + dsn_set: + type: boolean + description: Whether a connection string is stored. + host: + type: string + description: > + Host and port from the stored connection string, for + display. Empty when nothing is stored or the string cannot + be parsed — this decorates a response and never fails one. + database: + type: string + description: Database name from the stored connection string, for display. + configured: + type: boolean + description: Whether anything is stored at all. + + AskAIWidgetConfig: + type: object + description: > + The ask-ai widget's configuration, and both the PUT body and the + GET response — this is the one setting on this surface that is + not a credential, so there is nothing to redact. All fields are + false/empty when nothing is configured, which is what makes the + unconfigured state read as "the widget is off". + properties: + enabled: + type: boolean + description: > + Whether the widget is offered to end users at all. A + disabled widget is a 404 from whatever serves it, not an + empty answer. + allowed_origins: + type: array + maxItems: 20 + items: { type: string } + description: > + Origins permitted to embed the widget — scheme, host and + optional port, never a path or a pattern. "*" is refused by + name: this widget answers questions about the signed-in + user's own sessions and audit events, so a wildcard would + let any page on the internet ask them through a visitor's + browser. + entities: + type: array + maxItems: 8 + items: { type: string } + description: > + Which of cryden's allowlisted entities the widget will + answer over — the host narrowing the engine's vocabulary. + Validated against cryden's own allowlist rather than a copy + of it. This is the one field here with teeth: the API + enforces the configured subset in front of the provider, + since cryden's widget.Ask scopes to the calling user but + does so over every entity it allowlists. + greeting: + type: string + maxLength: 200 + description: The copy the embed shows before a question is asked. Required when enabled. + placeholder: + type: string + maxLength: 200 + description: The input's placeholder text. Optional. + responses: BadRequest: description: Malformed request body @@ -1414,3 +1599,383 @@ paths: content: application/json: schema: { $ref: '#/components/schemas/ErrorResponse' } + + /admin/settings/llm-provider: + get: + summary: The configured LLM provider, without its credential + description: > + Admin only. Returns which model and token ceiling back the + AI-assisted admin features, and whether a credential is stored — + never the credential itself, in any form. It is not masked and + not truncated to its last few characters: a console does not need + it (it renders "saved" from `api_key_set`), and returning any + part of it would put the secret in a browser's memory and a + devtools panel, which is what encrypting it at rest was for. + + Nothing configured yet is a 200 with `configured: false`, not a + 404: an empty settings form is an answer a console renders, while + a 404 would make it render a broken panel. Compare GET + /admin/digest/history, where the 404 is right because there the + absence means the feature is off rather than unconfigured. + + Answers 404 not_configured when SETTINGS_ENCRYPTION_KEY is unset, + which is this deployment saying it has nowhere safe to put a + credential. + security: [{ bearerAuth: [] }] + responses: + '200': + description: > + The stored provider with the credential replaced by + `api_key_set`, or an all-false body when none is stored. + content: + application/json: + schema: + type: object + properties: + data: { $ref: '#/components/schemas/RedactedLLMProvider' } + '401': { $ref: '#/components/responses/Unauthorized' } + '403': { $ref: '#/components/responses/Forbidden' } + '404': + description: not_configured. See this path's description. + content: + application/json: + schema: { $ref: '#/components/schemas/ErrorResponse' } + '409': + description: > + setting_undecryptable — a credential is stored but this + deployment's key cannot open it, which almost always means + SETTINGS_ENCRYPTION_KEY changed. Deliberately not reported + as 404: the setting is there, and telling an operator it is + missing would send them to re-enter a credential that is + fine. DELETE clears the row without needing the key. + content: + application/json: + schema: { $ref: '#/components/schemas/ErrorResponse' } + put: + summary: Store an LLM provider + description: > + Admin only. The credential is required on every write rather than + being optional with "blank means keep the existing one". That + convention is the usual one and it is the wrong one here: an + omitted field and a deliberately cleared one would be the same + request, and the failure mode of getting it wrong is a form that + looks as though it saved a key while silently storing an empty + one. + + The stored value is sealed with AES-256-GCM before it reaches the + table, keyed from SETTINGS_ENCRYPTION_KEY — a separate secret + from cryden's ENCRYPTION_KEY on purpose, since the two have + different lifetimes and blast radii. This credential is treated + with the same care as JWT_SECRET: it is never logged, never + returned, and never written in the clear. + + `max_tokens` defaults when omitted (0) rather than failing + validation, since a console that does not ask is the common case. + security: [{ bearerAuth: [] }] + requestBody: + required: true + content: + application/json: + schema: { $ref: '#/components/schemas/LLMProviderInput' } + responses: + '200': + description: > + The stored provider, redacted. Echoed as stored rather than + as submitted, minus the credential. + content: + application/json: + schema: + type: object + properties: + data: { $ref: '#/components/schemas/RedactedLLMProvider' } + '400': + description: > + invalid_llm_provider — an unknown kind, a missing or + malformed credential, or max_tokens outside the accepted + range. The message names the offending field and bound: the + caller is already an operator and none of it is secret. + content: + application/json: + schema: { $ref: '#/components/schemas/ErrorResponse' } + '401': { $ref: '#/components/responses/Unauthorized' } + '403': { $ref: '#/components/responses/Forbidden' } + '404': + description: not_configured. See GET on this path. + content: + application/json: + schema: { $ref: '#/components/schemas/ErrorResponse' } + delete: + summary: Clear the stored LLM provider + description: > + Admin only. Switches the AI-assisted query features off until a + provider is configured again. + + Needs no encryption key, which is what makes it the way out for a + deployment that has lost SETTINGS_ENCRYPTION_KEY and can no + longer read what it stored: clearing is always possible, even + when reading is not. + security: [{ bearerAuth: [] }] + responses: + '200': + description: Always the all-false redacted body. + content: + application/json: + schema: + type: object + properties: + data: { $ref: '#/components/schemas/RedactedLLMProvider' } + '401': { $ref: '#/components/responses/Unauthorized' } + '403': { $ref: '#/components/responses/Forbidden' } + '404': + description: not_configured. See GET on this path. + content: + application/json: + schema: { $ref: '#/components/schemas/ErrorResponse' } + + /admin/settings/database-provider: + get: + summary: The configured read-only database, without its credential + description: > + Admin only. Returns the label, the row ceiling and the two + harmless halves of the connection string — host and database — + so an operator can tell which connection is stored without being + shown a password. The user, the password and the full DSN are not + returned in any form and have no field in the response schema. + + Nothing configured yet is a 200 with `configured: false`; 404 + not_configured when SETTINGS_ENCRYPTION_KEY is unset. Same + reasoning as GET /admin/settings/llm-provider. + security: [{ bearerAuth: [] }] + responses: + '200': + description: The stored connection, described rather than returned. + content: + application/json: + schema: + type: object + properties: + data: { $ref: '#/components/schemas/RedactedDatabaseProvider' } + '401': { $ref: '#/components/responses/Unauthorized' } + '403': { $ref: '#/components/responses/Forbidden' } + '404': + description: not_configured. See this path's description. + content: + application/json: + schema: { $ref: '#/components/schemas/ErrorResponse' } + '409': + description: setting_undecryptable. See GET /admin/settings/llm-provider. + content: + application/json: + schema: { $ref: '#/components/schemas/ErrorResponse' } + put: + summary: Store a read-only database connection, after verifying it + description: > + Admin only. This endpoint is where cryden's "the credential + boundary, not just the allowlist, is the real safety guarantee" + requirement is enforced, and it enforces it in three steps whose + order is the whole design: + + 1. Validate the shape, which is cheap and local. + 2. Connect with the supplied credentials and ATTEMPT A WRITE — + create a temporary table, in the session's own pg_temp schema + so nothing is left behind. Nothing is stored unless the + database itself refuses. + 3. Only then store it, sealed. + + The check is an attempted write rather than a reading of the + role's attributes, and it cannot be done in the console: a + "read-only?" checkbox is a claim about a database made by whoever + ticked it, and a browser cannot open a Postgres connection at + all. Only the server's refusal is evidence, and it is evidence + about the actual role, on the actual database, through the actual + credentials. + + The cost is that this endpoint is slower than its neighbours: it + opens a connection and runs statements before answering, bounded + by a ten-second timeout. That is paid once per save, not per + query. A connection that cannot be made, times out, or fails for + any reason other than an insufficient-privilege refusal is + refused as unverified — not accepted. Treating "could not find + out" as a pass would make the check succeed exactly when it is + least able to tell. + security: [{ bearerAuth: [] }] + requestBody: + required: true + content: + application/json: + schema: { $ref: '#/components/schemas/DatabaseProviderInput' } + responses: + '200': + description: The stored connection, redacted. + content: + application/json: + schema: + type: object + properties: + data: { $ref: '#/components/schemas/RedactedDatabaseProvider' } + '400': + description: > + invalid_database_provider for a malformed DSN, a missing + label or max_rows out of range; database_role_not_read_only + when the role could write; database_role_unverified when the + check could not reach a conclusion. The last two are + different answers on purpose — the first is a grant to + change, the second is usually a connection string to fix. + content: + application/json: + schema: { $ref: '#/components/schemas/ErrorResponse' } + '401': { $ref: '#/components/responses/Unauthorized' } + '403': { $ref: '#/components/responses/Forbidden' } + '404': + description: not_configured. See GET on this path. + content: + application/json: + schema: { $ref: '#/components/schemas/ErrorResponse' } + delete: + summary: Clear the stored database connection + description: > + Admin only. Needs no encryption key, for the same reason + DELETE /admin/settings/llm-provider does not. + security: [{ bearerAuth: [] }] + responses: + '200': + description: Always the unconfigured redacted body. + content: + application/json: + schema: + type: object + properties: + data: { $ref: '#/components/schemas/RedactedDatabaseProvider' } + '401': { $ref: '#/components/responses/Unauthorized' } + '403': { $ref: '#/components/responses/Forbidden' } + '404': + description: not_configured. See GET on this path. + content: + application/json: + schema: { $ref: '#/components/schemas/ErrorResponse' } + + /admin/settings/ask-ai-widget: + get: + summary: The ask-ai widget's embed and scope configuration + description: > + Admin only. Returns the widget's enabled flag, the origins + allowed to embed it, the entities it will answer over, and its + copy. + + Unlike the other two settings paths there is no redaction here, + because there is nothing to redact: none of these fields is a + credential, and all of them are what the console shows an + operator anyway. + + There is deliberately no embed snippet in the response. The + snippet is markup the csax+ console renders into its own pages, + and the URL in it would name an endpoint this API does not serve + yet — so generating one here would hand a console a script tag + pointing at a 404. What this endpoint owes the console is the + configuration a snippet is built from, which is what it returns. + + Nothing configured yet is a 200 with `enabled: false`, the zero + value: a console then shows the feature as off rather than + showing a form that looks as though someone filled it in. + security: [{ bearerAuth: [] }] + responses: + '200': + description: The stored widget configuration. + content: + application/json: + schema: + type: object + properties: + data: { $ref: '#/components/schemas/AskAIWidgetConfig' } + '401': { $ref: '#/components/responses/Unauthorized' } + '403': { $ref: '#/components/responses/Forbidden' } + '404': + description: not_configured. See this path's description. + content: + application/json: + schema: { $ref: '#/components/schemas/ErrorResponse' } + '409': + description: setting_undecryptable. See GET /admin/settings/llm-provider. + content: + application/json: + schema: { $ref: '#/components/schemas/ErrorResponse' } + put: + summary: Store the ask-ai widget's embed and scope configuration + description: > + Admin only. This is the endpoint the config tuning advisor's + suggestions pre-fill, in the sense recorded for version 1.4: a + suggestion ends up as a value in this form and a human presses + save. Nothing calls it automatically, and no AI-assisted handler + in this API holds a reference to it. + + `entities` is the one field with teeth rather than decoration. + cryden's widget.Ask force-scopes every parsed intent to the + calling end user's own rows — it discards whatever identity + filter the model produced and substitutes the real one — but it + scopes over the whole of cryden's allowlist. Narrowing that + further is a host decision, so this API enforces the configured + subset in front of the provider. The list is validated against + cryden's own allowlist rather than a copy of it, so this API + cannot refuse an entity the engine permits. + + A disabled widget may be otherwise empty: an operator switching + the feature off should not have to fill in fields that are about + to stop mattering. Anything that IS filled in is still validated, + so a form cannot store a value that was never checked and would + be rejected the moment it was switched on. + security: [{ bearerAuth: [] }] + requestBody: + required: true + content: + application/json: + schema: { $ref: '#/components/schemas/AskAIWidgetConfig' } + responses: + '200': + description: The configuration as stored. + content: + application/json: + schema: + type: object + properties: + data: { $ref: '#/components/schemas/AskAIWidgetConfig' } + '400': + description: > + invalid_ask_ai_widget — an entity cryden does not allowlist, + a duplicated entity, a malformed or wildcard origin, or the + required fields missing while the widget is enabled. A + wildcard origin is refused by name with the reason, since + this widget answers questions about the signed-in user and + "*" would let any page on the internet ask them. + content: + application/json: + schema: { $ref: '#/components/schemas/ErrorResponse' } + '401': { $ref: '#/components/responses/Unauthorized' } + '403': { $ref: '#/components/responses/Forbidden' } + '404': + description: not_configured. See GET on this path. + content: + application/json: + schema: { $ref: '#/components/schemas/ErrorResponse' } + delete: + summary: Clear the ask-ai widget configuration + description: > + Admin only. Leaves the widget disabled — the zero value — until + it is configured again. Needs no encryption key, like the other + two DELETEs on this surface. + security: [{ bearerAuth: [] }] + responses: + '200': + description: Always the disabled, empty configuration. + content: + application/json: + schema: + type: object + properties: + data: { $ref: '#/components/schemas/AskAIWidgetConfig' } + '401': { $ref: '#/components/responses/Unauthorized' } + '403': { $ref: '#/components/responses/Forbidden' } + '404': + description: not_configured. See GET on this path. + content: + application/json: + schema: { $ref: '#/components/schemas/ErrorResponse' } From 5e4fe879c4ab3bea8cecebdf385c96a7f7d9e8e0 Mon Sep 17 00:00:00 2001 From: Raymond Nicholas Date: Wed, 16 Sep 2026 16:04:24 +0100 Subject: [PATCH 12/12] docs: record Tier 4 Stage 2 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit NEXT.md marks the three Stage 2 endpoints built with what was and was not done against each bullet; CURRENT-STATE gains the Stage 2 section; and PROGRESS carries the honest verification note — CheckReadOnly has never run against a real Postgres, the Anthropic provider has never called Anthropic, and 013_settings has never been applied. Co-Authored-By: Claude Code --- docs/development/CURRENT-STATE.md | 128 ++++++++++++++++---- docs/development/NEXT.md | 107 ++++++++++++++--- docs/development/PROGRESS.md | 193 ++++++++++++++++++++++++++++++ 3 files changed, 385 insertions(+), 43 deletions(-) diff --git a/docs/development/CURRENT-STATE.md b/docs/development/CURRENT-STATE.md index 6ce57b5..04efa0b 100644 --- a/docs/development/CURRENT-STATE.md +++ b/docs/development/CURRENT-STATE.md @@ -15,12 +15,18 @@ has its own `httpapi/apple.go` — see `NEXT.md` Tier 1). Tier 2 added one admin endpoint on top of those, the first in this repo — see below. Tier 3 added three more admin endpoints and this repo's first three tables of its own, plus the config that lights up Argon2id, -cloud logging and email templates — see below. Tier 4's Stage 1 added -three more admin endpoints and this repo's fourth table — the weekly +cloud logging and email templates — see below. Tier 4 added six more +admin endpoints and two more tables of its own: Stage 1 is the weekly digest and its recorded history, the support-ticket login diagnosis and -the config tuning advisor. Every admin endpoint in this repo is either -read-only or an explicit operator action on a named key; nothing on that -surface applies a suggestion by itself. +the config tuning advisor; Stage 2 is the AI provider settings — the LLM +provider, the read-only database and the ask-ai widget config. Tier 4 is +also where this repo stopped being purely a wrapper: it now ships a live +`ai.LLMProvider` over the Anthropic SDK and a live `ai.QueryableStore` +over a second database connection, neither of which is wired to a +consumer yet. Every admin endpoint here is read-only except the +`/v1/admin/settings/*` saves, which are the human half of the +pre-fill-never-auto-apply rule — nothing on that surface applies a +suggestion by itself. Tier 1 also added the second-factor surface: TOTP enroll/confirm/ disable, passkey registration/list/delete, magic-link request/complete, @@ -308,15 +314,19 @@ in-memory double, not against Postgres `FOR UPDATE SKIP LOCKED`, and that double cannot reproduce two workers racing. `PROGRESS.md` says all of this plainly. -## Tier 4 — AI-assisted admin endpoints (Stage 1): DONE +## Tier 4 — AI-assisted admin endpoints: DONE Built in two stages on `feat/tier4-ai-admin-endpoints`, for the same reason Tier 3 was: the three read-only reports below had their decisions -already made in `NEXT.md`, while Stage 2 needs two decisions that are -not a build session's to make (see the end of this section). `go build`, -`go vet`, `gofmt -l` and `go test ./...` are clean, `httpapi` is also -green under `-race`, and `PROGRESS.md` records what that does and does -not cover. +already made in `NEXT.md`, while Stage 2 needed two decisions that are +not a build session's to make. Those two were resolved by following +`NEXT.md`'s own instruction to make the reasonable call and record it — +see Stage 2 below. `go build`, `go vet`, `gofmt -l` and `go test ./...` +are clean, `httpapi` and the two new packages are also green under +`-race`, and `PROGRESS.md` records what that does and does not cover, +which is a lot. + +### Stage 1 — the three read-only reports Everything here is `RequireAdmin`, read-only, and buildable on the engine alone — no LLM, no second database connection, no outbound call: @@ -373,19 +383,89 @@ sandbox — and the digest schedule is a goroutine on `context.Background()`, because this repo still has no graceful shutdown. `PROGRESS.md` says both plainly. -## Tier 4 Stage 2, and Tier 5 +### Stage 2 — the providers and the widget config + +This is the half of the tier that needed an LLM, a second database +connection and an outbound call, so it is the half where this repo +stopped being purely a wrapper. Three settings endpoints, all +`RequireAdmin`, all in `httpapi/settings_handlers.go`, backed by +`settings/` and `migrations/013_settings`: + +- **`GET`/`PUT`/`DELETE /v1/admin/settings/llm-provider`** stores which + model and key back `ai.LLMProvider`. `DELETE` was added alongside the + specced pair: a settings screen with no way to clear a credential is + a screen an operator cannot leave. +- **`GET`/`PUT`/`DELETE /v1/admin/settings/database-provider`** stores + the connection `ai.QueryableStore` runs against. +- **`GET`/`PUT`/`DELETE /v1/admin/settings/ask-ai-widget`** stores the + widget's enabled flag, allowed origins, entity scope and copy. + +Both credentials are sealed with **AES-256-GCM before they reach the +table**, keyed from a new `SETTINGS_ENCRYPTION_KEY`. That key is +deliberately *not* cryden's `ENCRYPTION_KEY`: the two seal different +things with different lifetimes and blast radii, and this repo already +sets the precedent with `CLOUD_LOG_HASH_KEY`. An unset key is not a +startup failure — the three endpoints answer `404 not_configured`, like +every other optional feature here. The encryption itself is cryden's +`security.NewAESGCMEncryptor` rather than a second implementation of the +same primitive; see `settings/secrets.go`. + +Three things in this stage are worth reading before touching them: + +- **`PUT /database-provider` proves the role cannot write, then stores.** + Order is the whole design: validate the shape, connect with the + supplied credentials and attempt a write, and only store once the + server refuses. The probe targets `pg_temp`, so a failed probe leaves + nothing behind, and the pool is pinned to one connection so the + `CREATE` and the `INSERT` share the session owning that temp table. + Three outcomes are distinguished — refused is a pass, succeeded is + `400 database_role_not_read_only`, anything else is + `400 database_role_unverified` and **is not a pass**. Only SQLSTATE + `42501` counts as a refusal, matched by code rather than message. +- **`aiprovider.ScopedProvider` gives the widget's `entities` setting + teeth.** cryden's `widget.Ask` force-scopes every parsed intent to the + calling end user's own rows, overwriting whatever identity filter the + model produced rather than validating it — no oracle — but it does so + over all of `ai.AllowedEntities`. Narrowing that is a host decision, so + this repo refuses an out-of-scope entity in front of the provider. +- **`settings.AskAIWidgetConfig` is not a credential**, and that is why + it has no `Redacted` counterpart while the other two do. All three are + stored through the same `Secrets` wrapper anyway — one storage path + with one rule about what reaches the table is worth more than saving a + decryption. + +`aiprovider.NewAnthropic` is a real `ai.LLMProvider` over the official +Anthropic Go SDK, and `aiprovider.NewPostgresSnapshot` a real +`ai.QueryableStore`. **Nothing wires either from the stored config yet**: +the only consumer would be a widget serving endpoint, which does not +exist, so that glue lands with its first caller rather than being +written blind. `allowed_origins` is stored and validated but nothing +consults it at request time for the same reason, and the widget GET +carries no embed snippet because the URL in one would name a route this +repo does not serve. + +What Stage 2 does **not** have evidence for, and `PROGRESS.md` says in +full: `CheckReadOnly` has never run against a real Postgres (the tested +branch is the *unverifiable* one, not the pass), the Anthropic provider +has never called Anthropic (it is tested against a local fake in the +Messages API's wire shape), and `013_settings` has never been applied to +a database. + +**The read-only rule now has a named exception, and it is this one.** +`/v1/admin/settings/*` is the admin surface's first write. The reading +is that `CLAUDE.md`'s rule covers the AI *tools* — which cryden builds +through interfaces carrying no way to act — rather than every route +under `/v1/admin`, and that a settings save is exactly what `NEXT.md`'s +pre-fill-never-auto-apply decision names as the human half. No +AI-assisted handler holds a reference to these routes, and none accepts +a suggestion as input. The alternative readings (store the key in +cryden, or environment-only) are worse and one of them is explicitly +ruled out by `NEXT.md`, which says this repo owns that config storage. + +## Tier 5 Not started. See `NEXT.md` for the full, ordered, specced-in-detail -queue. Stage 2 is the LLM provider config, the read-only database -provider config and the ask-AI widget config; -`ai.LLMProvider`/`ai.QueryableStore` have no implementation in this repo -yet, so nothing in it has an endpoint. Two decisions are open and were -left for the user rather than guessed at: whether this repo ships a live -LLM client against a real vendor (an outbound integration, which this -repo has so far shipped none of), and where the at-rest encryption key -for the stored provider credential comes from (reusing the existing -`ENCRYPTION_KEY` is the obvious candidate and still a decision with a -blast radius). Tier 4 stays read-only by construction, with the -decision already made that an AI suggestion **pre-fills** a settings -form and never auto-applies. +queue — the users admin surface, which has no engine gap and is just +missing endpoints, plus the widget's own serving endpoint, which is what +the Stage 2 config above is waiting for. diff --git a/docs/development/NEXT.md b/docs/development/NEXT.md index 41b42e8..eed060c 100644 --- a/docs/development/NEXT.md +++ b/docs/development/NEXT.md @@ -288,24 +288,45 @@ Two details were decided rather than assumed, and are recorded in ## Tier 4 — AI-assisted admin endpoints (all behind `RequireAdmin`) -> **Status: Stage 1 is built on `feat/tier4-ai-admin-endpoints`.** -> The weekly digest and its schedule and history, the support-ticket -> assistant and the config tuning advisor all exist, are wired in -> `main.go`, and are tested end to end on the in-memory stores — -> `go build`/`go vet`/`go test ./...` clean, `httpapi` also green under -> `-race`. Stage 2 — the LLM provider config, the read-only database -> provider config and the ask-AI widget config — is **not started**: -> `ai.LLMProvider` and `ai.QueryableStore` still have no implementation -> in this repo, so nothing in Stage 2 has an endpoint yet. +> **Status: Stage 1 and Stage 2 are both built on +> `feat/tier4-ai-admin-endpoints`.** The weekly digest and its schedule +> and history, the support-ticket assistant, the config tuning advisor, +> the LLM provider config, the read-only database provider config and the +> ask-ai widget config all exist, are wired in `main.go`, and are tested +> end to end on the in-memory stores — `go build`/`go vet`/`go test ./...` +> clean, `httpapi` also green under `-race`. > -> What is still owed from Stage 1, said plainly: **migration -> `012_digest_runs` has never been applied to a database** (no Postgres -> in this sandbox — the same is true of `009`–`011`), so the digest -> history's real `PostgresStore` has only been reasoned about, not run; -> and the digest schedule is a goroutine on `context.Background()`, -> because this repo still has no graceful shutdown. +> What is still owed, said plainly, because none of it is a small +> caveat: > -> Two things Stage 1 changed that were not in the spec below, both +> - **No migration in this tier has ever been applied to a database.** +> There is still no Postgres in this sandbox, so `012_digest_runs` and +> `013_settings` have only been reasoned about, not run — the same is +> true of `009`–`011`. Every `PostgresStore` added here is unexercised. +> - **`aiprovider.CheckReadOnly` has never run against a real Postgres.** +> The probe is a `CREATE TEMP TABLE` plus an `INSERT`, and the branch +> that matters — SQLSTATE 42501 arriving as a `*pq.Error` — has only +> been tested against a closed port, which is the *unverifiable* +> outcome rather than the pass. The accepting path is the one no test +> here covers. +> - **The Anthropic provider has never called Anthropic.** It is tested +> against a local `httptest` server in the Messages API's wire shape, +> which pins the request this repo builds and the response it parses, +> but it is not evidence that the live service agrees. +> - **The ask-ai widget has no serving endpoint.** Stage 2 stores its +> embed and scope configuration and enforces the scope in +> `aiprovider.ScopedProvider`; nothing yet calls `widget.Ask`. So +> `allowed_origins` is recorded and validated but nothing consults it +> at request time, and the GET response deliberately carries no embed +> snippet, because the URL in one would name a route this repo does +> not serve. +> +> What is still owed from Stage 1: +> +> - the digest schedule is a goroutine on `context.Background()`, because +> this repo still has no graceful shutdown. +> +> Three things this tier changed that were not in the spec below, all > recorded because they are behaviour rather than plumbing: > > - **`LOCKOUT_THRESHOLD`/`LOCKOUT_DURATION_MINUTES` are now passed to @@ -322,9 +343,28 @@ Two details were decided rather than assumed, and are recorded in > write a row, so an operator hitting it twenty times does not fill > the history with twenty near-identical reports. Only the scheduled > job writes. - -Every endpoint in this tier stays read-only/surface-only, no -exceptions — see `CLAUDE.md`'s hard rule at the top. +> - **`/v1/admin/settings/*` is the admin surface's first write**, and +> the read-only rule below has been read as covering the AI *tools* +> rather than every route under `/v1/admin`. The reasoning is in +> `SettingsHandlers`' doc comment and in `CLAUDE.md`'s own wording: a +> settings save is what "a human still has to explicitly save that +> change through the normal config UI" names, and no AI-assisted +> handler holds a reference to it. The alternative reading — store the +> LLM key in cryden, or in the environment only — is worse: the spec +> below explicitly says this repo owns that config storage. +> +> The two decisions this tier had recorded as open were resolved by +> following this file's own instruction to make the reasonable call and +> note it: the live provider is built on the **official Anthropic Go +> SDK** rather than hand-rolled HTTP, and the settings credentials use a +> **dedicated `SETTINGS_ENCRYPTION_KEY`** rather than reusing cryden's +> `ENCRYPTION_KEY`, matching this repo's existing convention of +> purpose-specific keys (`CLOUD_LOG_HASH_KEY`). + +Every AI-assisted endpoint in this tier is read-only by construction — +see `CLAUDE.md`'s hard rule at the top. The settings routes at the end of +this list are not AI-assisted endpoints: they are the settings save those +tools' suggestions pre-fill. - **Weekly digest**: `GET /v1/admin/digest` → `cryden.WeeklyDigest`/ `DigestSince`. Plus **scheduling and history** (new, this repo's own @@ -357,6 +397,15 @@ exceptions — see `CLAUDE.md`'s hard rule at the top. at-rest encryption — treat this credential with the same care as `JWT_SECRET`). This repo then constructs the real `ai.LLMProvider` implementation from that stored config at startup or on change. + **Built, with one piece of this bullet not done.** The endpoints + exist, `DELETE` was added alongside `GET`/`PUT` (a settings screen + with no way to clear a credential is a screen an operator cannot + leave), and `aiprovider.NewAnthropic` is the real implementation, + built on the official Anthropic Go SDK. What is **not** built is the + last sentence: nothing reads the stored config and constructs a + provider from it, because nothing consumes one yet — the widget's + serving endpoint does not exist. The glue lands with its first + caller rather than before it, so it is not written blind. - **Database Provider config** (new): same shape, for pointing `ai.QueryableStore` at a read-only database role/connection string. **The read-only-role requirement is not optional** — cryden's own @@ -366,10 +415,30 @@ exceptions — see `CLAUDE.md`'s hard rule at the top. role is actually read-only before accepting it if there's any feasible way to check (e.g. attempt a write and confirm it's rejected), don't just trust a checkbox in the UI. + **Built.** `PUT` connects with the supplied credentials and refuses + to store anything until the server has rejected a write on that + connection — see `aiprovider.CheckReadOnly`. Two outcomes are + distinguished that the bullet does not mention, because they call + for different words: a role that *can* write, and a check that could + not reach a conclusion. The second is refused too, since treating it + as a pass would make the check succeed exactly when it is least able + to tell. `aiprovider.NewPostgresSnapshot` is the matching + `ai.QueryableStore`; like the provider above, nothing constructs it + from the stored config yet, for the same reason. - **Ask-AI widget embed/scope config** (new): once the two providers above exist, `widget.Ask` itself needs no new engine work — expose whatever embed snippet / scope configuration the csax+ console needs as its own settings endpoint. + **Built, with two deliberate departures.** The endpoint stores the + widget's enabled flag, origins, entity scope and copy. + `aiprovider.ScopedProvider` then *enforces* the entity scope — + cryden's `widget.Ask` scopes every intent to the calling end user but + does so over all of `ai.AllowedEntities`, so narrowing that is a host + decision and a scope setting nothing consulted would be worse than no + setting. And the response carries **no embed snippet**: the snippet is + markup the console renders into its own pages, and the URL in one + would name a route this repo does not serve. The console gets the + configuration a snippet is built from instead. --- diff --git a/docs/development/PROGRESS.md b/docs/development/PROGRESS.md index 42cae43..322ff4e 100644 --- a/docs/development/PROGRESS.md +++ b/docs/development/PROGRESS.md @@ -812,3 +812,196 @@ Both were left for the user rather than guessed at. is worse than one that refuses to start — but it means an existing deployment with a typo in one of them will fail to boot rather than run with a default. + +## 2026-09-16 — Tier 4, Stage 2 (LLM provider, read-only DB, ask-ai widget) + +Stage 2 of Tier 4, on the same branch. Three settings endpoints, two new +packages of this repo's own, and the point where this repo stopped being +purely an HTTP wrapper. + +Commits, in order: + +- `bf1abaa` — the settings store and its at-rest encryption + (`settings/`, `migrations/013_settings`, `SETTINGS_ENCRYPTION_KEY`). +- `bc7de0e` — `aiprovider.NewAnthropic`, a live `ai.LLMProvider` over + the official Anthropic Go SDK. +- `b31ef89` — `aiprovider.PostgresSnapshot` plus `CheckReadOnly`. +- `2815e90` — the LLM and database provider endpoints. +- `10cea8d` — the ask-ai widget config and `aiprovider.ScopedProvider`. +- plus an `openapi` bump to 1.5 and a README section. + +### The two decisions the Stage 1 entry left open were taken + +Both by following `NEXT.md`'s own instruction — "where something is +genuinely unspecified, make the most reasonable call consistent with +`CLAUDE.md`'s ownership rules and note the assumption in `PROGRESS.md`" +— rather than by asking, since the instruction to ask was absent and the +spec was explicit that it should not be needed. + +1. **A live LLM client, yes, and on the official SDK.** Implementing + `ai.LLMProvider` was unavoidable: the spec says the console + configures a provider, and a console configuring a provider nothing + can call is not useful. The SDK over hand-rolled HTTP because a + hand-rolled client would be a second thing to keep correct against a + moving API, and because it is the one part of this repo whose + correctness cannot be checked by reading it. +2. **A dedicated `SETTINGS_ENCRYPTION_KEY`, not a reuse of + `ENCRYPTION_KEY`.** The precedent is already in this repo: + `CLOUD_LOG_HASH_KEY` exists rather than reusing `JWT_SECRET`. The two + seal different things with different lifetimes and blast radii — + cryden's key covers what the engine stores, this one what the API + stores — so one rotating should not force the other. Reusing it would + also mean a TOTP-secret rotation and an API-key rotation cannot be + scheduled apart. + +The encryption itself is cryden's `security.NewAESGCMEncryptor`, not a +second AES-GCM implementation. That is the "if cryden already answers +the question, call it" rule applied to a primitive: there is nothing +about a provider API key that needs different treatment from a TOTP +secret, and two implementations of the same cipher is one more place for +a nonce to be reused. + +### The read-only check is the one piece of this tier worth reading twice + +`PUT /v1/admin/settings/database-provider` validates the DSN's shape, +then **connects with the supplied credentials and attempts a write**, +and stores nothing unless the server refuses. cryden's own interface +comment is the requirement ("a real credential-level guarantee, not just +a promise made in code, so a bug in validation still can't cause a +write"), and `NEXT.md` says not to trust a checkbox. Three details are +load-bearing: + +- The probe writes to `pg_temp`, the session's own temporary schema, so + a probe that fails leaves nothing for an operator to clean up. The + pool is pinned to one connection so the `CREATE` and the `INSERT` + share the session that owns the temp table — a pool that split them + would have the `INSERT` fail on a missing table, which looks exactly + like the refusal being tested for and is not one. +- Only SQLSTATE `42501` counts as a refusal, matched by code rather than + by message, because the message is localized and reworded between + major versions and this is the branch that decides acceptance. +- A third outcome is distinguished from both: a connection that never + opened, a timeout, or a `CREATE` that failed for a non-privilege + reason is `database_role_unverified` and **is refused**. Treating + "could not find out" as a pass would make the check succeed precisely + when it is least able to tell. + +The cost is that this endpoint is slow relative to its neighbours — +a connection and two statements — bounded by a ten-second timeout. That +is once per save, not once per query. + +### The widget's entity scope has teeth, on purpose + +`widget.Ask` force-scopes every parsed intent to the calling end user's +own rows, overwriting rather than validating the identity filter the +model produced (no oracle: every phrasing executes the same query). But +it scopes over all of `ai.AllowedEntities`. Narrowing that is a host +decision — cryden's allowlist is "what can be scoped safely", the +operator's is "what this deployment offers" — so this repo enforces the +configured subset in `aiprovider.ScopedProvider`, in front of the +provider, which is the only place the entity is still visible before +`Ask` parses, scopes and executes in one call. The alternative was +storing a scope setting nothing consulted, which is worse than not +having the setting. + +The list is checked against cryden's own `AllowedEntities` map rather +than a copy, so this repo cannot refuse an entity the engine permits. +One asymmetry is accepted knowingly and commented: `scopeToOwner` is a +private switch over today's three entities, so an entity added to +cryden's allowlist without a matching case there would pass validation +and then fail at `Ask` time with `ErrEntityNotAvailable`. That is the +safe direction — the widget refuses the question rather than answering +it unscoped. + +### Verification: what this does NOT cover + +The suite is green — `gofmt -l` clean, `go build ./...`, `go vet ./...`, +`go test -count=1 ./...` all pass, and `httpapi`, `settings` and +`aiprovider` also pass under `-race`. That is not the same as this +working, and three specific things are unproven: + +- **`aiprovider.CheckReadOnly` has never run against a real Postgres.** + There is no Postgres in this sandbox. `query_test.go` covers the + statement builder, the filter allowlist, the LIKE escaping and the + probe statements' shape, and `TestPutDatabaseProviderRefusesAnUnverifiableConnection` + drives a real connection attempt — but against a *closed port*, which + exercises the unverifiable branch. **The accepting path — 42501 + arriving as a `*pq.Error` and being read as a pass — is the branch no + test here covers**, and it is the branch the feature depends on. + Likewise nothing has confirmed that a `CREATE TEMP TABLE` is actually + refused by a `GRANT SELECT`-only role as opposed to failing some other + way, which is the assumption the probe is built on. +- **The Anthropic provider has never called Anthropic.** It is tested + against a local `httptest` server in the Messages API's wire shape, + which pins the request this repo builds and the response it parses. + That is a real test of this repo's half and no evidence at all about + the live service's half: a model id, a schema field name or a refusal + shape that differs in production would not be caught. +- **`012_digest_runs` and `013_settings` have never been applied to a + database**, the same as `009`–`011`. Every `PostgresStore` in + `settings/` and `digest/` is reasoned-about rather than run, so a + column type or a constraint error would surface at first deploy. + +Also unchanged from every previous tier: `internal/smoketest` has still +never been run. + +### Two things deliberately not built, rather than half-built + +- **Nothing constructs `ai.LLMProvider` or `ai.QueryableStore` from the + stored config.** `NEXT.md` asks for it ("this repo then constructs the + real implementation from that stored config at startup or on change"). + The glue's only possible consumer today is a widget serving endpoint, + which does not exist, so writing it now would mean writing the + consumer's half blind and then rewriting it. It lands with its first + caller. +- **The widget GET returns no embed snippet.** The snippet is markup the + console renders into its own pages, and its `