Repository navigation
[Decision] published skills/** teaching errors are found by hand (21 of 109 lane cards in two weeks): run the skill evals or a schema check over skill examples in CI (a new gate), or keep maintainer-scored evals #22059
Description
Activity
objectstack-fleet commented
on Oct 7, 2026 ContributorAuthorMore actionsBox intake (triage):
domain:skills·priority:p3added. The card's shape is complete: 维护者速读, the option table with costs, the four axes,Prior rulings read:, the confidence gap and 裁后执行 are all present. The quoted rule text was checked onmain:SKILL.md:255carries 「新增门禁默认否」, andskills/README.md:67says theevals/fixtures are "used by maintainers to score the skill".Triage seat (objectstack-wide, seat post #6015) ·
session_01AavokzJ5DndAwitDXvKy4U· 2026-10-07T06:00Z. ⛔ Not a claim.- Why p3: today's path works. Teaching errors are caught by hand and close in a median of 9 hours, as the card measures, and no release waits on this ruling.
- For the director: the rule's default answer for an unnamed new gate is 「不加」 (C). B rests on reading it as a blind-spot fix of an existing gate, and the card puts that reading to the maintainer itself. Triage adds no recommendation.
- Nothing else changes. The card waits in the box for the director, beside decision: open v18 now and ship it in stages — release the last 17.x from main first without waiting for #21908's deny (A), skip the last 17.x (B), or keep #22009's order (C)? #22050 and [Decision] the >5000-line SIZE limb counts a pure deletion the maintainer already ruled (ruling 208), and the Tier H wait for an authorized APPROVED has no clock — PR #22002 idles green #22051.
objectstack-fleet commented
on Oct 8, 2026 ContributorAuthorMore actionsRuling: batch #287 item 2 · letter B (with the probe as its precondition) · maintainer 「批 #287 同意」 2026-10-08T02:21Z
Director seat, summon #35,
session_01VYToj6PQehTEKNrjGM9akg(GitHubos-zhuang; written asobjectstack-fleet[bot]via the relay). Batch #287 was presented in chat twice, the second time in the detailed form the maintainer asked for, with this card's options A / B / C / D and the question whether B counts as a new gate underSKILL.md:255(新增门禁默认否); the maintainer answered 「批 #287 同意」, which takes B as presented: the seat's framing of B as a blind-spot fix of the existingcheck-corpus-claim-driftgate, and, where that reading is contested, the maintainer's own word on this batch is the naming the rule asks for. Freshness gate: the body and the one comment (6031996518, triage's intake) were re-read before this record; nothing newer exists. Thread-read: 6031996518.The ruling
B, in two steps, the second conditional on the first.
- The probe. Extract the example code blocks from the published
skills/objectstack-*texts (TypeScript, JSON and YAML fences alike; a TypeScript example needs a transform into a shape the schemas accept, and whether that transform covers the catalog's fences is what the probe measures), validate them againstpackages/spec's schemas, and replay the validator over the fix diffs of the 21 teaching-error cards named in the card's premise 1. Report the hit rate. - The gate step, only if the hit rate is at least one half. The validation joins
check-corpus-claim-drift's own CI step (no new job, no model call, no inference cost), with its self-test and several of the 21 cases as fixed fixtures. A structural teaching error (a wrong key, a wrong nesting, a retired spelling) then fails the pull request loudly; a semantic one (a rule stated backwards) still needs a reader, as the card says plainly. If the hit rate is under one half, the ruling falls to C: no gate, the finding process as today, and the probe's readings stay on this card.
⛔ Not taken: A (a model scoring the
evals/fixtures in CI: a non-deterministic gate that adds inference cost and a new source of false reds), D (an offline report nobody reads, the fateevals/already has).Why B on the first axis: the spec is the declaration and a shipped example must parse under it; the check reuses the schemas and the corpus gate that exist, where A would add a part the project then has to keep calibrated.
State
needs-user-decision→pm:queuein this act;tooling·domain:skills·priority:p3unchanged, no assignee: thedomain:skillsseat claims the probe from its queue; the gate step follows on this card only on a passing probe. Body:Ruled:line added in the same act.
Generated by Claude Code
- The probe. Extract the example code blocks from the published
- added and removed
on Oct 8, 2026 objectstack-fleet commented
on Oct 8, 2026 ContributorAuthorMore actionsClaim: PM loop round 2
Session:session_01CXydFDyiQwNbGFkmwrcRQq
Account:huangyiirene(the seat's linked user asGET /useranswers it; always the card's assignee)
Branch:claude/issue-22059-skill-example-schema-probe
Worktree:objectstack-issue-22059
Domain:domain:skills
Seat:domain:skills#1
File surface: step ① of the ruling (the probe, this dispatch):scripts/check-corpus-claim-drift.mjs— a report-only probe mode beside its existing batteries, or a sibling module underscripts/that it imports — with--self-testcases; the probe READS the publishedskills/objectstack-*/**/*.mdtexts and the fix diffs of the 21 teaching-error cards (REST reads) and reports the hit rate on this card. ⛔ In step ① no edit toskills/**,packages/spec/**,.github/workflows/**,package.json,scripts/corpus-claim-drift-baseline.jsonorscripts/pm/dispatch-gates.mjs(FROZEN). Step ② (the gate step: the validation joinscheck-corpus-claim-drift's existing CI step with fixtures) follows on this card and branch only on a hit rate of at least one half and only after this seat reads the probe report. Not a governed surface: a code PR, green + ACCEPT → queue landing. Stop on breach; explain in the report
Container & model:M,mode:subagent,model: opus—dispatch-gates --tier --repo objectstack-ai/objectstack scripts/check-corpus-claim-drift.mjsprints no path-derived mandate ("the surface hits none of the 3 declared glob(s) … floor sonnet · default opus · ceiling fable"); reason: the default judgment tier for a measurement-first probe whose extraction and transform choices decide a gate ruling, judged from the card content
Clause-②: no
Responsibility:n/a — not a defect card(a ruled decision: letter B, ruling 6050857613, maintainer 「批 #287 同意」 at 2026-10-08T02:21Z)
Thread-read: 6050857613
Serial constraints cleared: no open PR touchesscripts/check-corpus-claim-drift.mjs, its baseline orlint.yml(nine open PRs' file lists read at 2026-10-08T02:26Z); same-day churn none on the file (last touchcc1b8ac94dat 2026-09-13T08:08Z); file-disjoint from #22123 (same batch,skills/objectstack-upgrade/SKILL.md);scripts/pm/dispatch-gates.mjsFROZEN and not in this surface; verify lock free, queue empty (arrival depth 1 at this read); the lane's open P0/P1: none ⇒ p3 is a candidate (NORTH-STAR 「优先级」 3 read). Ruling-ref: 6050857613 (retrieved in this act, with itsGoverning text:items). Readings taken at 2026-10-08T02:28Z.
Generated by Claude Code
objectstack-fleet commented
on Oct 8, 2026 ContributorAuthorMore actionsos-dev-report
{ "issue": 22059, "status": "done", "branch": "claude/issue-22059-skill-example-schema-probe", "pr": "https://github.com/objectstack-ai/objectstack/pull/22171", "session": "session_01CXydFDyiQwNbGFkmwrcRQq", "premise_still_valid": false, "summary": "Step one of ruling B is delivered as a report-only probe: `node scripts/check-corpus-claim-drift.mjs --probe-examples` (module `scripts/skill-example-probe.mjs`, exit 0 on any reading, wired into no check script, workflow or baseline), with four new self-test batteries (65 cases). HIT RATE 2/19 = 10.5% over the cards filed before #22059 that have a landed published-catalog repair (2/20 = 10.0% over all 20 the listing matched at filing; 2/22 = 9.1% over today's 22). The 2 hits are #20173 (page:header actions; ComponentPropsMap) and #21392 (select with no options; FieldSchema refinement), both in os:check blocks the TS type gate passes. Misses over the 22: prose 10, comment-or-string 4, neither 2, transform 2, not-in-corpus 1 (#18964), no-fix 1 (#22123). Ceiling with a perfect transform: 6/19 = 31.6%, because 9 of the 19 repairs touched no example fence. Transform coverage: 283 population fences (262 typescript, 16 ts, 1 tsx, 2 json, 2 jsonc, 0 yaml). Of the 279 TS fences, the transform produced a value for 203 (72.8%) and 115 (41.2%) were read and mapped whole. Control on today's catalog: 110 green, 3 red. Two of the reds are real teaching errors (`skills/objectstack-ui/rules/actions.md:77` and `:176`, both throw under the built `defineAction`; the TS gate `check:skill-examples` is green over :77). The third is the upgrade skill's deliberate protocol-16 example. Premise corrections: the listing gives 20 cards at filing, not 21 (#22120 and #22123 were created later; #18964 is a seat process finding). Four gates already read examples opt-in (check:skill-examples, check:yaml-examples, check:doc-formula-expressions, check:doc-security-posture). By the ruling's own terms the readings sit under one half; the step-two or C decision is the seat's.", "tests": "Head e6d00e1421. (1) `pnpm check:corpus-claim-drift`: exit 0, self-test OK with the Probe batteries at 14/18/21/12 cases, gate 'OK, no new claim sites' over 252 files. (2) `node scripts/check-corpus-claim-drift.mjs --probe-examples --json F`: exit 0. Apart from its head line, the output is byte-identical on every run from before the first commit to e6d00e1421. (3) Ablations via scripts/ablation-replace.mjs, each landed (anchor 1 to 0, blob moved) and each restored (blob == HEAD, git diff HEAD empty): opaque suppression off gives self-test exit 1 (2 cases); the HIT rule off gives exit 1 (1 case); the probe registration call off gives exit 1 (all 4 Probe batteries 'DID NOT RUN'). (4) `node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstack` derived 30 commands; all 30 were run on e6d00e1421 and all exited 0, `pnpm check:pm-dispatch-gates` included ('1976 cases pass', EXIT=0). --ran reconciliation: 30 derived, 30 run, 0 UNRUN. Roster families under scripts/ also run: check-published-list-mirrors (both), check:dts-closure, check:engine-double-contract, check:i18n-stale-fill and check-dts-references --self-test exit 0. NOT MEASURED: check:published-readme-exports, reason: PREREQUISITE NOT MET, exit 3, 45 packages unbuilt. check:console-injection, reason: no console dist, skipped. Test Core CI shard steps and the type-check lanes are CI's own; no package TypeScript is touched. An earlier battery run on 4a648ffec1 failed 1 of 1976 cases (compound self-test function names unledgered). The names were fixed in e6d00e1421 rather than ledgered. (5) eslint narrowed to the 2 touched files: `eslint --no-inline-config --format json` gives 2 files, 0 errors, 0 warnings. The population is eslint.config.mjs `files: ['**/*.{ts,tsx,mts,cts,js,jsx,mjs,cjs}']`. Invariance: no parserOptions.project or typed rules, so the diff cannot move a verdict on an untouched file. (6) Control: `pnpm --filter @objectstack/spec run check:skill-examples` exits 0 ('262 prose examples type-check across 3 surface(s)') with actions.md:77 in it. The spec and the client-react closure were built first through os-verify-lock. (7) pnpm check:nul-bytes exits 0; the control-byte self-scan of both files is empty.", "mcp_calls": "0 — none", "api_writes": "3 relay dispatches — POST /repos/objectstack-ai/objectstack/dispatches ×3, executed as objectstack-fleet[bot]: (a) pr_create, i.e. POST /pulls (draft #22171; read back 13547/13547 bytes identical); (b) label-write, i.e. POST /issues/22171/labels (skip-changeset) plus POST /issues/22171/assignees (huangyiirene); read back matching; size/xl came from another actor; (c) this os-dev-report comment, i.e. POST /issues/22059/comments. git push is not REST. Reads (REST GET, not writes): issue 22059 and its comments (2), the domain:skills listing (3 pages), 22 card timelines, PR 22171 read-back (1).", "open_questions": [], "out_of_scope_findings": [ "class: c · reach: copying the published example verbatim makes the built `defineAction(...)` throw at module load. Measured at e6d00e1421; os validate and the metadata save parse with the same ActionSchema. · evidence: skills/objectstack-ui/rules/actions.md:77 (os:check block 'Modal-typed action'): `type: 'modal'` plus `body` is refused by the ActionSchema refinement on `body` ('`body` only runs for `type: 'script'` …'). actions.md:176 ('opensInNewTab + newTabUrl'): `type: 'url'` with no `target` is refused: 'Action 'target' is required when type is 'url', 'flow', 'modal', 'api', or 'form'.' The TS gate check:skill-examples is green over :77. One family (one file, one schema): file as one card. · Seam: spec:ActionSchema refinements (body vs type, target vs type) → runtime:defineAction / metadata save · dedupe words: actions.md modal body · defineAction target required url · opensInNewTab newTabUrl target · Modal-typed action body script", "carrier: step two of #22059 if it lands (an expect-red convention), else 承接者:无 · noted, not filed — skills/objectstack-upgrade/references/examples-upgrade.md:7 reads red (connectors[].fieldMappings, label, type) because it is the deliberate protocol-16 FROM shape the upgrade skill migrates; a probe gap, not a teaching error (PR Acceptance notes)" ] }
Generated by Claude Code
objectstack-fleet commented
on Oct 8, 2026 ContributorAuthorMore actionsClose-out — ruling B step ① delivered, readings under one half ⇒ C · skills seat 1,
session_01CXydFDyiQwNbGFkmwrcRQq, 2026-10-08T04:02ZACCEPT of the probe readings (dev report 6051826214; PR #22171, head
e6d00e1421), reviewed against the PR body's replay table and the tree, not the report's narrative:- Hit rate 2 / 19 = 10.5% over the cards filed before this one that have a landed published-catalog repair; 2 / 20 over the listing at filing; 2 / 22 over today's. Ceiling with a perfect transform 6 / 19 = 31.6%, because 9 of the 19 repairs touched no example fence and 4 more changed only comments or strings inside one. Spot-check by this seat: replay row skills(objectstack-ui): the dashboards rule tells authors to override a compareTo series with
series.dashArray/series.opacity— a dashboard widget'schartConfig.seriesis a tombstone the schema refuses by name #21211 (class "prose", 0 / 0 fences) against its landing PR docs(skills): the dashboards rule stops teaching a per-series override on the compareTo comparison series #21283: one file, +2 / −2, no fence opener changed. Holds. - Positive controls: the two HITs are the cards' own known errors (skills(objectstack-ui): pages.md and its eval teach full Action objects in
page:header.properties.actions; the contract is action ids (ruling B on #11592) andos validateflags the taught shape #20173page:headeractions as objects; finding(skills): the objectstack-upgrade skill's "v17-canonical shapes, compiled" block authors aselectwith no options, whichFieldSchemarefuses once PR #21390 lands #21392selectwith no options), both inos:checkblocks the TypeScript gate passes. Catalog control: 110 green, 3 red; the twoactions.mdreds re-read by this seat onorigin/mainec8f37c890(sources and refinements) are real teaching errors and are filed as finding(skills): objectstack-ui rules/actions.md teaches two action shapes the ActionSchema refinements refuse — a modal carrying abody(:77) and aurlaction with notarget(:176); copying either makesdefineActionthrow #22176; the third is the upgrade skill's deliberate protocol-16 shape, a probe gap, not an error. - Not re-run by this seat: the probe itself (it needs a built
packages/spec/dist; the review stayed read-only). The ruling's bar is one half; no reading, denominator or ceiling comes near it, so the decision does not turn on the measurement's last digit.
Decision, from the ruling's own text: "If the hit rate is under one half, the ruling falls to C: no gate, the finding process as today, and the probe's readings stay on this card." So: no step ②, no gate, no change to
check-corpus-claim-drift's CI step. The readings stay here (the report above and PR #22171's body). PR #22171 is closed without merging (its note says why; the branch stays ate6d00e1421as the probe's source of record). The card closescompleted: the decision is made and executed.Premise corrections, acknowledged in public as the dev made them: the premise's listing matched 20 cards at filing, not 21 (#22120 and #22123 were created after this card; #18964 is a seat process finding, not in the published corpus). Four existing gates already read example blocks opt-in (
check:skill-examples,check:yaml-examples,check:doc-formula-expressions,check:doc-security-posture), against the body's "none reads what an example teaches": the correct sentence is that none reads a Zod refinement, which is where both HITs and both new reds sit. No follow-up card: nothing further turns on either correction once C is taken.Carriers for the dev's acceptance notes: the expect-red convention for the upgrade skill's FROM shapes, the 86 ownership misses and the
Pluginhold-out were gaps a gate would have had to close; with no gate they have no carrier and are recorded here only.State:
pm:dispatchedand the assignee come off in this act; closedcompleted. Reads and writes of the delivery:mcp_calls0;api_writes3, all relay dispatches asobjectstack-fleet[bot].
Generated by Claude Code
- Hit rate 2 / 19 = 10.5% over the cards filed before this one that have a landed published-catalog repair; 2 / 20 over the listing at filing; 2 / 22 over today's. Ceiling with a perfect transform 6 / 19 = 31.6%, because 9 of the 19 repairs touched no example fence and 4 more changed only comments or strings inside one. Spot-check by this seat: replay row skills(objectstack-ui): the dashboards rule tells authors to override a compareTo series with
Ruled: 6050857613 · letter B (probe first) · 2026-10-08T02:24Z
Filing gate: ② a decision only the maintainer can make — whether the published
skills/**catalog gets a CI gate over its teaching content (theevals/fixtures scored in CI, or a schema check over the skill texts' example blocks), which SKILL.md:255 makes a maintainer-named act (新增门禁默认否), or stays maintainer-scored offline asskills/README.md:67says today. Filed bydomain:skillsseat 2 (seat post #19287,session_0181E4ZeZmWyknawnauxD2CE) on the maintainer's instruction in this session, verbatim: 「创建卡片,暂时不派发。」 ⛔ Not a claim, ⛔ not a dispatch.Reader: the maintainer (the
needs-user-decisioninbox); after the ruling, thedomain:skillsseat executes per 「裁后执行」 below.Dedupe: REST listings, closed included (
labels=toolingsince 2026-09-07: 417;labels=needs-user-decisionsince 2026-09-07: 2; every issue updated since 2026-10-01: 625;domain:skillssince 2026-09-23: 109; union 1,030) grepped forevals? (in|as|to) CI|published skills|skills/\*\*|skills catalog|skill eval→ 3, all closed programs: #13658 (factual sweep of every behavioural claim inskills/**), #14292 (skills catalog optimization program), #19715 (the four SKILL.md files over 500 lines). None asked for a gate. Existing mechanical coverage read frompackage.jsonandscripts/: five structural gates (check:skill-identifier-liveness,check:skill-frame-sync,check:skill-frame-freshness,check:skill-compatibility,check:skill-top-level-keys) and the lexical ratchetcheck-corpus-claim-drift.mjs(#13582) — none reads what an example teaches.维护者速读
问题:近两周 skills 车道 109 张卡里有 21 张是发给客户的
skills/objectstack-*教错(ui 7 张,ai、api、i18n、upgrade、data 各若干;其中 p1 四张,如 #21567 把公开表单的三个共享键只写了两个、#21537 把容器表单写在不存在的formViews.default下)。每一张都是人读出来、立 finding、再派人修,中位 9 小时关单。十个 skill 都带evals/夹具(12 个 JSON、prompt→expected_output 对),但 README 写明它们是「维护者离线打分用」,CI 不跑;五个结构门只查 frontmatter、标识符与版本号,不查示例教的对不对。推荐:B——不跑 evals,而是在既有门
check-corpus-claim-drift的盲区上补一步:把 skills 文本里的示例代码块按packages/spec的 schema 校验,spec 一变就红。第一步先做一次探针:用这个校验回扫这 21 张卡的修复 diff,数一数能抓住几张;抓不到一半就改选 C(不加门)。A(模型在 CI 里给 evals 打分)不荐:不确定性门禁是随机红的新来源。你要做的:回一个字母 A / B / C / D。B 或 A 都算「新增门禁」,按规则要你在回批里点名;C 本卡即关。
Governing text
.claude/skills/pm-dispatch/SKILL.md:255: 「新增门禁默认否:门禁是只减不增的零件,例外只有维护者点名;荐新增门禁须引其点名原话,否则荐「不加」。」 — this card has no naming to quote, so its default recommendation would be 「不加」; B is offered as a blind-spot fix of an existing gate and the maintainer decides whether that reading holds..claude/skills/pm-dispatch/SKILL.md〈升级与决策〉 不升级类: 「既有门禁内部参数与盲区修复(加强,非削弱,非新增)」 — the clause B leans on..claude/skills/pm-dispatch/references/filing-gate.md:44: 「北极星第 4 条只认出货面(skills/**、产品文档),⛔ 不含.claude/**内部协议。它认的是一句说错的话,⛔ 不是一句缺席的话。」skills/README.md:67: "evals/— skill eval fixtures — used by maintainers to score the skill".scripts/pm/check-skill-line-ratchet.mjs:2938-2941: "Extending coverage to the published catalog is a policy change — it lands with a maintainer ruling" (said of the line ratchet; the same posture is read here for any new coverage ofskills/**).前提(每条带 re-check 命令)
domain:skillscards updated in the two weeks before this filing concern published skill text; the ids read: finding(skills): the objectstack-ui skill authors a container's create/edit form underformViews.default/formViews.edit, which PR #21535 serves as a named form, never the default #21537, skills(objectstack-api): the published public-form opt-in names two of the three sharing keys the anonymous form endpoints require — an AI following it authors a form both endpoints answer 404 #21567, finding(skills): objectstack-ai's Common Pitfalls says "memoryis declared only — no runtime reads it"; once PR #21413 lands, the cloud AI runtime enforces it and the spec refuses its old shapes #21415, skills/objectstack-i18n: the_actionsrow and thetranslation-target-unknownparagraph omit the action-entry keysos validatejudges (param options,outcomeMessages,resultDialogleaves and fields) #21583, skills/objectstack-ui: thekind: 'html'page example writes<object-metric aggregate="count">, a string the html tier only warns on and the tile cannot draw #21627, skills(objectstack-ui): the dashboards rule says Postgres buckets with date_trunc; the SQL driver groups by to_char(... AT TIME ZONE UTC) on Postgres #21588, skills(objectstack-ui): the dashboards rule tells authors to override a compareTo series withseries.dashArray/series.opacity— a dashboard widget'schartConfig.seriesis a tombstone the schema refuses by name #21211, finding(skills, docs): published AI-author text still describesguardrailsas unenforced andoutputSchemaas never validated; the cloud AI runtime now enforces both (objectstack-ai SKILL.md, actions.mdx) #21288, finding(skills): the objectstack-upgrade skill's "v17-canonical shapes, compiled" block authors aselectwith no options, whichFieldSchemarefuses once PR #21390 lands #21392, skills(objectstack-ui): dashboards.md teaches the measure-currency chain with nocurrencyMode: 'fixed'condition and cites ADR-0053 — remainder (a) of #20126 #20196, skills(objectstack-ui): pages.md and its eval teach full Action objects inpage:header.properties.actions; the contract is action ids (ruling B on #11592) andos validateflags the taught shape #20173, finding(skills,docs): the currency teaching infield-types.mdandfield-types.mdxcontradicts the spec:dynamicis taught as a per-record{ value, currency }, and the resolution order has no mode #20090 and nine more. —gh api 'repos/objectstack-ai/objectstack/issues?state=all&labels=domain:skills&since=YYYY-MM-DDThh:mmZ&per_page=100' --jq '[.[] | select(.title|test("skills\\(objectstack-|skills/objectstack-|finding\\(skills"))] | length'(the window's start instant, 2026-09-23T00:00Z, goes insince=)find skills -path '*/evals/*' -type f | wc -landgrep -rn "evals" package.json .github/workflows | wc -lcheck:skill-*gates read structure (frontmatter keys, identifiers, versions), not example semantics. —grep -c '"check:skill-' package.jsoncheck-corpus-claim-drift.mjsis a shrink-only lexical ratchet over the teaching corpus,skills/**included, pinned to the [finding] The teaching corpus has no lexical anti-drift pin, so a filter operator's prose can contradict its driver forever — the reachable half of "no gate can catch it" #13582 claim family. —sed -n '1,40p' scripts/check-corpus-claim-drift.mjsclosed_by_pull_requests) listskills/**paths only.一句话问题
发给客户的 skills 教错了,是继续靠人读出来,还是让机器在 spec 每次变化时去撞一下示例。
选项 × 真实代价
evals/夹具交给模型在 CI 里打分,低分即红check-corpus-claim-drift的盲区上加一步:skills 文本的示例代码块按packages/specschema 校验;先探针回扫 21 张卡的修复 diff 验证命中率业务含义直译
A = 「让 AI 批改 AI 写的教材,红了就拦」;B = 「教材里的示例必须能被规范解析」;C = 「客户先撞上,我们再改」;D = 「有个体检表,没人看」。
四轴论证(从业务立场)
推荐 + 回退 + 置信缺口
推荐 B,带前提:先做探针(用校验器回扫 21 张卡的修复 diff),命中率低于一半即改荐 C。回退 C。置信缺口:未测校验器对这 21 张的命中率;未测示例代码块的提取是否能覆盖 skills 文本里的所有代码围栏形态(TS、JSON、YAML 混用);不知道维护者是否把 B 读作「新增门禁」。
自检行:只看①选 B;②③④ 是否翻转:否。
终态句:两年后每份发布的 skill 示例都能被平台自己的 schema 解析,教材与规范是一个源;参照主流做法,文档代码块进 CI 做 doctest 是常态(Rust doctest、mdBook test)。
os-decision-facets
① 项目长远合理性:B 让示例服从 spec 声明,复用既有门,不新造机制;A 养一个非确定性门。
② 实际业务拉动:两周 21 张出货面教错卡、p1 四张,每张都可能被客户的 AI 照抄。
③ 防 AI 犯错:schema 校验在 PR 期响亮拒绝结构性教错;模型打分把真错与噪声混在一起。
④ 创业阶段不扩散:B 不加 job 不加推理成本;A 加新零件;D 是声明不兑现。
Prior rulings read: skills,evals,gate,corpus,teaching,published,catalog → 101 ADR / 47 AGENTS.md / 920 spec term hits, 10 ADR candidates, none on the published catalog; ADR none; thread: #13658, #14292, #19715, #13582
推荐 B(带探针前提);选项 A / B / C / D;自检「只看①选 B;②③④ 是否翻转:否」。置信缺口:命中率未测;围栏形态覆盖未测。
裁后执行
domain:skills卡,两步:① 探针——提取 skills 文本示例代码块、按packages/specschema 校验、回扫 21 张卡的修复 diff,报命中率;② 命中率过半才接着把校验接进check-corpus-claim-drift的同一 CI 步(不加 job),含自测与 21 张中的若干作为固定夹具。completed关闭;出货面教错继续走 finding 立卡(北极星第 4 条)。scripts/check-*.mjs --report。Related: #13658 · #14292 · #19715 · #13582 · the 21 cards named in 前提 1.
Generated by Claude Code