Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 16 additions & 0 deletions openspec/changes/hackathon-analysis/apply-progress.md
Original file line number Diff line number Diff line change
Expand Up @@ -1285,3 +1285,19 @@ Checked against the repo's generated types and the official docs (developers.clo
- **Diagnostics (`feat(hackathon-llm)`, 9cac8b0):** seam chosen: `LlmExtractor.extract` returns `{ value, meta }` (not a sentinel), so the never-throw-for-content contract and `value: null` for unparseable output stay, and the domain builds `attempts` from plain data. `meta` is `finishReason` (plain token <= 20 chars, else `other`), `contentLength` and `parseFailure` (`no-content|unterminated|prose-around|not-json|non-object`). The safe logger re-validates each key (pattern, finite number, fixed-code set); tests prove no content, snippet or hostile finish_reason reaches the log.
- **Tolerant extraction (`feat(hackathon-llm)`, c178c8c):** on a failed direct parse the first balanced `{...}` is parsed (string/escape-aware). Decision: a recovery reports `parseFailure: "prose-around"` AND `recovered: true`, so diagnostics still show the model wraps its JSON. Recovered values still pass through `validateExtraction` unchanged (tests: garbage and off-page snippets rejected). Fence, whitespace and `reasoning_content` tests are unchanged and green.
- **Docs:** design.md "Extraction Schema and Prompt" updated.

### Root-cause fix: production-faithful evidence (branch `fix/hackathon-extraction-root-cause`)

- **Evidence:** a harness running the REAL modules on Cloudflare (`wrangler dev --remote`, real AI and BROWSER, Cloudflare egress) against .../tokenized-stocks showed: (1) the static fetch gets 403, so production always uses the rendered path; its text has TABs (`innerText` table cells). (2) Qwen copies TAB-containing snippets into JSON strings; `JSON.parse` throws "Bad control character", reported as `parseFailure: not-json`; escaping in-string control chars parses all 11 keys. (3) GLM returns unfindable fields as `{value:null,snippet:null,confidence:0}` and the all-or-nothing shape check failed the whole response as `invalid-shape`. (4) `parsed:false` was reported for any `ok:false`, even when JSON parsing succeeded.
- **Decisions:** repair reported as a NEW code `parseFailure: "control-chars"` + `recovered: true` (one signal, allowlisted through the shared code list; inside prose it stays `prose-around`). A MISSING key counts as rejected (`wrong-shape`), not null: treating it as null would let `{"unrelated":true}` become a usable all-null success. `parsed` = JSON parsing succeeded (`value !== null` or `non-object`); `shape: "invalid"` only when parsed and the top level is not a plain object. `THIN_STATIC_TEXT_THRESHOLD` and `isUsable` are now exported so the harness shares the production rule.

| Work unit | Commit | RED | GREEN | Notes |
|---|---|---|---|---|
| 1 Page text normalization | c04cb08 | 6 failing (tabs in rendered and static output, `normalizePageText` missing) | 770 passed | triangulated: table tabs, control chars/CR, blank-line runs, TEXT_MAX cap |
| 2 Control-char tolerance | 34a6d42 | 6 failing (real TAB shape, CR/LF, prose-around, fence, non-object) | 76+ passed | also asserts structural whitespace untouched and no content in meta |
| 3 Per-field shape | 2841998 | 20 failing | 798 passed | GLM null object, array value, string teamSize, top-level array, majority; old whole-response tests rewritten to per-field; spec + design updated |
| 4 Diagnostics label | 7e077c9 | 4 failing | 803 passed | logger asserts fixed `shape` literal and no value leak |
| 5 Harness | 6dbbcd5 | N/A: needs live Cloudflare bindings (run by the orchestrator, not vitest) | typecheck clean, harness type-checked separately, `wrangler deploy --dry-run` bundles it | excluded from root tsconfig, vitest and the main bundle |

- **Work Unit Evidence:** focused commands `npx vitest run <files>` per unit (results above); runtime harness: N/A in this batch (run by the orchestrator against real URLs); rollback boundary: one commit per unit.
- **Verification:** `npm test` 803 passed (70 files), `npm run typecheck` clean.
11 changes: 10 additions & 1 deletion openspec/changes/hackathon-analysis/design.md
Original file line number Diff line number Diff line change
Expand Up @@ -74,6 +74,11 @@ index.queue ─ buildHackathonConsumer(env) ─ runHackathonJob(msg)
└─ transient error → attempts ≤ 2 ? retry(30 s) : as permanent ("temporary error")
```

### Worker egress reality

Verified with the harness (`npm run harness`, real modules on Cloudflare with real AI and BROWSER bindings) against https://www.bnbchain.org/en/hackathons/tokenized-stocks: from Cloudflare egress the static fetch gets HTTP 403, so production always takes the rendered (browser) path, whose text is 2.5k-3.4k chars. Rendered `innerText` separates table cells with TAB characters, and Qwen copies snippets containing raw TABs into JSON string literals, which `JSON.parse` rejects ("Bad control character in string literal"). Consequences in the design: (1) both fetch paths pass their text through `normalizePageText` (every ASCII control character except `
` becomes a space, space runs collapse per line, 3+ newlines collapse to a blank line; the 22,000 cap still applies), and (2) the extractor repairs control characters in string literals as a second line of defense. GLM returns fields it cannot find as `{ "value": null, "snippet": null, "confidence": 0 }`, handled by the per-field validation above.

The rendered fetcher calls `page.setRequestInterception(true)`. Each request goes through the pure `browserRequestPolicy`: http(s) only, the same host guard, images, fonts, media and stylesheets aborted, and at most 100 requests. After `goto`, it re-checks `page.url()`, reads `innerText` (capped), and calls `browser.close()` in `finally`. A 429 raises `BrowserQuotaExceededError`.

The static fetcher uses `redirect: "manual"`. It follows at most 3 hops, re-guarding each one, and streams up to 2 MB. It accepts only a 200 `text/html` or `text/plain` response. `HTMLRewriter` drops noise elements and keeps the title, the meta and OG description and `ld+json` (up to 4 KB). The text is capped at 22,000 chars.
Expand All @@ -88,7 +93,11 @@ These are unchanged:

## Extraction Schema and Prompt

`Field<T> = { value; snippet; confidence } | null`, where the prompt asks for a snippet of at most 160 characters copied verbatim from a single passage and for concise values, and validation accepts a snippet that is verbatim modulo whitespace (every whitespace run, including line breaks, tabs and NBSP, collapsed to one space and trimmed on both sides; the page is normalized once per response) with a 200-character cap on the normalized snippet, which is the form stored. Each rejection carries a reason code (`empty-snippet`, `snippet-too-long`, `not-verbatim`, `value-too-long`); `ExtractionFailedError` carries per-attempt `{ model, parsed, rejectedCount, rejected: [{ field, reason }] }` diagnostics that the job logs (names and codes only, never snippet text, values or the URL). Each attempt also carries safe parse metadata (`finishReason` capped to a plain token or `other`, `contentLength`, and `parseFailure` in `no-content|unterminated|prose-around|not-json|non-object`), returned by the extractor as `{ value, meta }` and re-validated by the safe logger; no model text is logged. Parsing is tolerant: when a direct parse (after fence stripping) fails, the first balanced JSON object (string- and escape-aware) is parsed instead and, if it is an object, used with `parseFailure: "prose-around"` and `recovered: true`; the recovered value still goes through `validateExtraction` unchanged. Unchanged:, covering name, format, location, team size, four dates, prizes, tracks and eligibility. Invalid fields become null, and so do fields whose snippet is not found in the page. The fallback model is tried when the output is unparseable or more than half of its fields are invalid. The page is framed as untrusted between `<<<PAGE`/`PAGE>>>` with those tokens stripped, and the call uses `temperature: 0` and `max_tokens: 2500` (both models are reasoning models whose thinking consumes completion tokens). The input is chat `messages`: fixed instructions in the system message and the framed page in the user message, because a bare `prompt` makes these models do raw text completion. GLM models additionally get `chat_template_kwargs: { enable_thinking: false }` (with thinking on, GLM-4.7-Flash exhausts max_tokens and returns truncated JSON; Qwen3-30B returns null content if thinking is disabled, so only GLM gets it).
`Field<T> = { value; snippet; confidence } | null`, where the prompt asks for a snippet of at most 160 characters copied verbatim from a single passage and for concise values, and validation accepts a snippet that is verbatim modulo whitespace (every whitespace run, including line breaks, tabs and NBSP, collapsed to one space and trimmed on both sides; the page is normalized once per response) with a 200-character cap on the normalized snippet, which is the form stored. Validation is per field: only a top-level value that is not a plain object fails the whole response (`invalid-shape`). A field object whose `value` is null (GLM returns `{ "value": null, "snippet": null, "confidence": 0 }` for what it cannot find) is a model-null and is not counted as rejected. A missing key, a non-object field, a wrong value type, a non-string snippet or a non-number confidence rejects that field alone with `wrong-shape` (nulled and counted in `rejectedCount`); a missing key counts as rejected, not null, so a garbage object such as `{ "unrelated": true }` cannot become a usable all-null success. Each rejection carries a reason code (`empty-snippet`, `snippet-too-long`, `not-verbatim`, `value-too-long`, `wrong-shape`); `ExtractionFailedError` carries per-attempt `{ model, parsed, rejectedCount, rejected: [{ field, reason }] }` diagnostics that the job logs (names and codes only, never snippet text, values or the URL). Each attempt also carries safe parse metadata (`finishReason` capped to a plain token or `other`, `contentLength`, and `parseFailure` in `no-content|unterminated|prose-around|control-chars|not-json|non-object`), returned by the extractor as `{ value, meta }` and re-validated by the safe logger; no model text is logged. Parsing is tolerant: when a direct parse (after fence stripping) fails, the first balanced JSON object (string- and escape-aware) is parsed instead and, if it is an object, used with `parseFailure: "prose-around"` and `recovered: true`; the recovered value still goes through `validateExtraction` unchanged. Control-character repair: when the direct parse fails, raw TAB/CR/LF inside JSON string literals are escaped by a string-aware scan (structural whitespace and existing escapes untouched) and the parse is retried once; success is reported as `parseFailure: "control-chars"` with `recovered: true` (a repair inside a prose-wrapped object stays `prose-around`). The safe logger allowlists the code from the same fixed list. Diagnostics: `parsed` means ONLY that JSON parsing succeeded; a top-level validation failure on a parsed value is reported as `shape: "invalid"` (logged as a fixed literal), and per-field problems appear in `rejected[]` with their reason code. Unchanged:, covering name, format, location, team size, four dates, prizes, tracks and eligibility. Invalid fields become null, and so do fields whose snippet is not found in the page. The fallback model is tried when the output is unparseable or more than half of its fields are invalid. The page is framed as untrusted between `<<<PAGE`/`PAGE>>>` with those tokens stripped, and the call uses `temperature: 0` and `max_tokens: 2500` (both models are reasoning models whose thinking consumes completion tokens). The input is chat `messages`: fixed instructions in the system message and the framed page in the user message, because a bare `prompt` makes these models do raw text completion. GLM models additionally get `chat_template_kwargs: { enable_thinking: false }` (with thinking on, GLM-4.7-Flash exhausts max_tokens and returns truncated JSON; Qwen3-30B returns null content if thinking is disabled, so only GLM gets it).

### Verification

Local tests fake the network, the AI binding and the browser, so they cannot show what production sees. Any change to the fetch, LLM or validation path MUST be checked with `npm run harness` (`scripts/hackathon-harness/`) against real URLs before merge. It costs Workers AI neurons and Browser Rendering time.

## Migration `0003_hackathon_analysis.sql`

Expand Down
20 changes: 17 additions & 3 deletions openspec/changes/hackathon-analysis/specs/llm-extraction/spec.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ Turns a fetched page's reduced text into a strict, nullable-field hackathon reco

### Requirement: Strict Schema Output

The system MUST request extraction against a fixed schema (format, team-size cap, dates, prizes, tracks, name) and MUST validate the model's response against that schema before it reaches the domain layer, rejecting any response that fails validation.
The system MUST request extraction against a fixed schema (format, team-size cap, dates, prizes, tracks, name) and MUST validate the model's response against that schema before it reaches the domain layer. A response that is not a JSON object is rejected as a whole; a malformed field is rejected individually.

#### Scenario: Well-formed response passes validation

Expand All @@ -18,11 +18,25 @@ The system MUST request extraction against a fixed schema (format, team-size cap

#### Scenario: Malformed response is rejected

- GIVEN the model returns a response that does not match the fixed schema (missing required shape, wrong types, or unparseable output)
- GIVEN the model returns a response whose top level is not a plain JSON object (unparseable output, an array, a string, ...)
- WHEN the adapter validates it
- THEN the system MUST treat this as an extraction failure
- THEN the system MUST treat this as an extraction failure (`invalid-shape`)
- AND MUST NOT pass partial or malformed data to the domain layer

#### Scenario: Malformed field is rejected individually

- GIVEN the response is a JSON object but one field is malformed (a missing key, a field that is not an object, a wrong value type, a non-string snippet or a non-number confidence)
- WHEN the adapter validates it
- THEN only that field MUST be rejected with the reason `wrong-shape`, stored as null and counted in `rejectedCount`
- AND the other fields MUST be validated normally
- AND the response is usable only while no more than half of the fields are rejected (otherwise the fallback model runs)

#### Scenario: A field the model reports as not found is null, not rejected

- GIVEN the model returns a field object whose `value` is null (for example `{ "value": null, "snippet": null, "confidence": 0 }`)
- WHEN the adapter validates it
- THEN the field MUST be null and MUST NOT count as rejected, whatever its snippet or confidence hold

### Requirement: Null Over Guess for Every Field

The system MUST represent a field the page does not clearly state as null rather than an invented value, for every field in the schema.
Expand Down
3 changes: 2 additions & 1 deletion package.json
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,8 @@
"test": "vitest run",
"test:watch": "vitest",
"typecheck": "tsc --noEmit",
"cf-typegen": "wrangler types"
"cf-typegen": "wrangler types",
"harness": "wrangler dev --remote --config scripts/hackathon-harness/wrangler.jsonc --port 8799"
},
"dependencies": {
"@cloudflare/puppeteer": "1.4.0",
Expand Down
48 changes: 48 additions & 0 deletions scripts/hackathon-harness/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
# Hackathon analysis harness

A dev-only Worker that runs the REAL `src/` modules (static fetcher, rendered
fetcher, Workers AI extractor, `validateExtraction`) on Cloudflare with real
`AI` and `BROWSER` bindings and Cloudflare egress. Unit tests fake all of
these, so they cannot show what production sees (for example, a static fetch
that gets HTTP 403 from Cloudflare egress, or a model copying raw TABs from
rendered page text). This harness can.

It is never deployed, and it is not part of the main Worker bundle, vitest or
`npm run typecheck` (`tsconfig.json` only includes `src`, `test` and
`vitest.config.ts`).

## Cost

Every request runs `wrangler dev --remote`, so it consumes real Workers AI
neurons (one call per model) and Browser Rendering time, billed to your
Cloudflare account. It needs `wrangler login` (or `CLOUDFLARE_API_TOKEN`; set
`CLOUDFLARE_ACCOUNT_ID` if you have several accounts). No DB, queue or
secrets are used.

## Run

```
npm run harness
# in another terminal:
curl "http://localhost:8799/?url=https://www.bnbchain.org/en/hackathons/tokenized-stocks"
```

Optional: `&models=@cf/model-a,@cf/model-b` overrides the model IDs (the
default is the pair in the root `wrangler.jsonc`).

## Output

JSON with:

- `static` / `rendered`: ok, text length, or the error kind and HTTP status.
- `chosen`: `static` or `rendered`, using the same rule as
`analyzeHackathon` (rendered when the static fetch is bot-walled or thin).
- `pageText`: length and the first 1500 characters of the chosen text.
- `attempts[]`, one per model: `finish` reason, extractor `meta`
(`parseFailure`, `recovered`, `contentLength`), the first 2500 characters of
the model content, the `validateExtraction` result and `usable`.

## When to use it

Any change to the fetch, LLM or validation path MUST be checked with
`npm run harness` against real URLs before merge.
Loading
Loading