Skip to content

test(spec): automation test titles state each cited decision in words instead of a tracker number (stage 10) - #21713

Merged
objectstack-fleet[bot] merged 2 commits into
mainfrom
claude/issue-20749-test-strings-a
Oct 4, 2026
Merged

objectstack-fleet[bot] merged 2 commits into
mainfrom
claude/issue-20749-test-strings-a

Conversation

@objectstack-fleet

Copy link
Copy Markdown
Contributor

Part of #20749
Clause-②: no

Stage 10 of this card, and the first area of class (e): the test strings shipped under packages/spec/src, as ruled in 5902360492 on #20513. This stage takes the whole automation/ directory. Its 105 test-title and test-message literals carried 110 tracker ids citing 55 records. Each id now either states what its record decided, in words (form D), or is dropped where the title already says it. Text only: no assertion, fixture value, test count or code comment changes.

Census at the base (7e0066af7a, the claim's base)

Instrument: stage 9's census.cjs (md5 6e42a45a926d375013c32d62f16a296e, byte-identical), plus one added classification pass. A literal counts as a test title when its folded message is argument 0 of a describe / it / test call, .each / .skip / .only chains included. Everything else is an "other" string.

Reference: the same instrument reads 1803 messages / 1919 ids in 425 files at 9b8c7f38d7, stage 9's reading exactly. Since then, #21699 added 1 / 1 (ui/component-props-unknown-members.pin.test.ts:143) and #21700 moved 3 / 3 (the two reader literals it re-anchored), which gives 1801 / 1917 at the base.

directory files titles msg / ids other msg / ids total msg / ids excluded files
data/ 95 445 / 475 23 / 26 468 / 501
ui/ 81 374 / 397 18 / 18 392 / 415 report.test.ts 3 / 3
api/ 40 181 / 193 8 / 8 189 / 201
system/ 34 128 / 138 26 / 27 154 / 165 job.test.ts 4 / 4
(files directly in src/) 30 117 / 119 1 / 1 118 / 120
automation/ (this PR) 21 101 / 106 4 / 4 105 / 110
kernel/ 36 92 / 96 10 / 10 102 / 106
shared/ 21 73 / 81 12 / 14 85 / 95
contracts/ 25 59 / 70 4 / 4 63 / 74
conversions/ 9 34 / 34 0 34 / 34
security/ 8 28 / 28 0 28 / 28
ai/ 9 13 / 15 5 / 5 18 / 20
identity/ 6 14 / 14 1 / 1 15 / 15
integration/ 4 13 / 13 1 / 1 14 / 14
migrations/ 2 9 / 12 0 9 / 12
marketplace/, meta-spelling/, studio/ 5 7 / 7 0 7 / 7
total 426 1688 / 1798 113 / 119 1801 / 1917 7 / 7

How the area was chosen

The directories are ranked by id count, and a stage takes whole directories up to about 100 ids. The four busiest each exceed that bound alone: data/ (501), ui/ (415), api/ (201) and system/ (165). The files directly in src/ (120) are 20% over. automation/ (110) is the busiest whole directory within about 10% of the bound, so it is this stage. The rule picked it before any card was read.

Its 55 records (53 in this repository, 2 in objectui) were all readable in one pass. 54 answer 200. #6362 answers 404, and its decision was read from its landing commit b5404f496.

Named for the next stages (by directory, from the table): data/ (about five stages, by subdirectory or file group; data/driver/ alone is 52), ui/ (about four), api/ (two), system/ (two), the files directly in src/ (one), kernel/ (one), shared/ (one), contracts/ with conversions/ (one, 108), and security/, ai/, identity/, integration/, migrations/, marketplace/, meta-spelling/ and studio/ together (one, 96). The seven excluded ids join system/ and ui/ once their owners land.

What each id became

38 literals (40 ids) now state a decision in words. 67 literals (70 ids) drop a citation the title already explains. Each record was read with its comments through REST, and where a record has no comments, from what landed.

record ids result
#3508 2 APPROVER_VALUE_BINDINGS: "an approver value is picked from the records the engine resolves". APPROVER_VALUE_SOURCES (the follow-up): "where each picker finds its candidates, published on the wire".
objectui#2955 1 "for the decision dialog to render and enforce": both decision entry points collect the typed outputs, and required is enforced.
#3810 1 "names the match-everything-write hazard": a filter emptied by interpolation matches every row, and the node is refused instead. The test pins the words match-everything write.
#4001 13 "strict as of #4001 批 9" becomes "an unknown key is refused, not stripped" (5 titles). Also "refused, not stripped", "an unknown key is refused, per shape" and "the unknown-key gate". Dropped from 5 titles that already read "unknown keys are rejected, not stripped".
objectui#2670 1 "so the flow designer renders it as a template": the designer's loop and region rendering reads this marker.
#4396 2 "a function that writes says so; pure is the default", and "the authoring surface where a function declares its effect" (landed eb4204b).
#4697 2 defaultValue: "a declared variable is bound on every path" (ruling A). Dropped once.
#3896 3 "outputSchema retired: declared, never validated" (the audit close-out's reason, as flow.zod.ts's tombstone records it). Dropped twice.
#4247 1 "maxRetries — one default, and no zero-attempt “retry”" (landed a648e96).
#16134, #15713 7 "a region node reusing a top-level id is refused — one node-id space now spans every region" (ruling, batch 61). Dropped from 5 titles that state uniqueness.
#9205 5 "template reference — notify content localized through an email template" (the ruled emailTemplates route). Three titles say "the template path" where they said "pre-#9205". Dropped once.
#7086 1 "severity — the closed info | warning | critical vocabulary" (the enum route).
#4415 3 "FlowNodeSchema parses its own regions" (ruling 2026-08-07, direction 1). Dropped from 2 titles that already say it.
#4347 1 "collectFlowGraphs — every region is walked, not only the top level" (landed 31e0be9).
#4401 2 "FLOW_REGION_SLOTS, the one declaration of where regions live" (landed 4bfd455). Dropped once.
#4414 1 "condition is pointed at the out-edges, NOT given the one-edit rename" (one working routing model, landed 5293114).
#15429 1 "taking every true branch must be declared" (ruling item 2: explicit inclusive).
#14149 3 "every entry older than the value role" (ruling A: a value-role CEL slot on assignment). Dropped twice.
#19938 4 "the CRUD fields.* value slots added exactly two rows". Dropped three times.
#15572 3 "predicateSlotRefusal — a predicate slot holds bare CEL text", and "the other two doors refuse it". Dropped once.
#15662 1 "structuralConditionRefusal — a structural condition is CEL text or an expression".
#15792, #15807 4 "REFUSES an ast-only envelope — admitted at first, refused once an evaluated slot required a source". Dropped from 2 titles that state the rule.
#17493 3 Table row "a blank string — blanks are refused" (ruling A). Dropped twice.
#14945 4 "EndConfigSchema — the end node contract: it may refuse the run with a message" (ruling 2′). Dropped three times.
#15617 4 "failed is the fold INCLUDING what a delegating node rolled up from its child — it answers what the run caused" (ruling option 1). Dropped three times.
dropped only 36 #3196, #3266, #4158, #4277, #4343 (2), #4389, #4525, #4738, #4964 (2), #6758, #7085, #9106, #12278, #14964, #15430, #15646 (3), #16752, #17306, #17852, #18102, #18112 (2), #18847, #19151, #19961 (4), #20316 (2): each title already states the pinned decision. For #4988 and #6414, the two expect messages already say what was retired.
#6362 (404) 1 Dropped. The title "PRESERVES all seven envelope keys — measured, not assumed" carries the decision recorded in b5404f496 (webhook was measured, and all seven keys survive).

Readers

  • Test-name filters: none. A tracked-tree search for -t and --testNamePattern finds only packages/qa/dogfood/README.md:142 (-t "owner-scoped"), which is unrelated.
  • Snapshots: none. automation/ has no __snapshots__ and no toMatchSnapshot.
  • Titles by substring: every old title, plus a window around each id (250 needles), was searched across the tracked tree outside its own file. No gate, doc or script matches one. The hits are other files' own titles with the same words: identity/, security/ and automation/ siblings, a later stage's lot. There are also two code comments in flow.zod.ts and flow-function.zod.ts, which belong to the comment lane.
  • Twin tables, not readers: packages/lint/src/validate-expressions.test.ts:4442/4445 and packages/services/service-automation/src/decision-branch-expression-absent.test.ts:61/64 repeat the flow-decision-branch-expression-absent table's row names. Nothing compares them mechanically: the "same table" parity is prose in the headers, and it covers assertions, not names. They are outside this stage's surface (see Acceptance notes).

Text-only proof

A scratch tool (textonly10.cjs) compares base and head file by file on three legs:

  1. Skeleton: the full AST, with string pieces masked. It must be identical.
  2. Comments: every comment, byte-equal.
  3. Strings: each string leaf that changed must sit in a test-call title position, or on one of the 4 declared lines (flow-decision-branch-expression-absent.test.ts:63 and :66, table name values that feed $name titles; sync-retirement.test.ts:120 and :158, expect messages). It must carry a tracker id before and no # plus digits after.
  • Result: 21 of 21 files SAME, 105 changed (101 title, 4 declared), on all three legs.
  • Diff hunks: exactly the 105 planned lines, with every file keeping its line count.
  • Controls (10 of 10 as predicted, on scratch copies, each anchor hit once): identifier rename DIFF; numeric literal DIFF; comment edit COMMENT DIFF; a non-title string with an id VIOLATION; a rewritten title given a new id VIOLATION; a title that was id-free at base edited VIOLATION; one title reverted to base SAME (104 changed); a declared string keeping an id VIOLATION; an undeclared expect message changed VIOLATION; a title re-split into a + chain DIFF.

Test counts: the 21 files run at the base (in a separate base worktree) and at the head: 801 / 801 tests on both sides, with the same count and status sequence per file in 21 of 21. 560 full test names change, and each equals the base name with the planned replacements applied.

Changeset: skip-changeset

Measured, not assumed:

  • npm pack --dry-run of @objectstack/spec lists 2068 files, under files[] (dist, src/**/*.zod.ts and the rest). 0 of the 21 touched files are in it, and 0 *.test.ts at all. The control src/automation/flow.zod.ts is in it.
  • In dist/, three new phrases and three old ones each read in 0 files. The control A predicate slot holds BARE CEL TEXT reads in 2.

So this PR publishes nothing, and no changeset is added.

Verification (at f3dc3fab03)

  • pnpm turbo run build over all packages: 71 / 71.
  • @objectstack/spec: vitest run --project local, 612 files and 18185 passed, 1 todo. typecheck exit 0, including check:test-typecheck, whose program holds all 21 touched files.
  • Gates: dispatch-gates --commands derived 79 families, and all 79 exit 0. --ran reconciles: 79 derived, 79 run, 0 NOT-MEASURED, 0 UNRUN.
  • ESLint, a proven narrowing: --no-inline-config over the 21 files, 0 errors and 0 warnings. The population comes from ESLint's own config: 21 configured, 0 ignored. No parserOptions.project or projectService, so no untouched file's verdict can move.
  • check-governed-merges --test: NOT governed, 210 changed lines.

Acceptance notes


Generated by Claude Code

claude added 2 commits October 4, 2026 09:00
… instead of a tracker number

The 105 test-title and test-message literals under packages/spec/src/automation
that carried a tracker id (110 ids, 55 records) now either state what the cited
record decided, or drop a citation the title already explains. Text only: no
assertion, fixture value, test count or code comment changes.

Claude-Session: https://claude.ai/code/session_01T9u38rswFp5Rw8DswRUReJ
Co-authored-by: Claude <noreply@anthropic.com>
vitest prints a table row's `$name` truncated past about forty characters, so
the rewritten row now states its decision in a name that prints whole.

Claude-Session: https://claude.ai/code/session_01T9u38rswFp5Rw8DswRUReJ
Co-authored-by: Claude <noreply@anthropic.com>
@objectstack-fleet objectstack-fleet Bot added the skip-changeset PR has no user-facing published change; bypasses the changeset gate label Oct 4, 2026
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

📓 Docs Drift Check

Nothing in this diff resolved to a documentable surface (no symbol, route or SDK anchor derived from 0 changed package(s)), so this run has no opinion about the docs.

What this run could not see

Coarse fallback — 0 page(s) merely mention a changed package (the pre-#9192 predicate, kept for the deliberately-wide backstop): node scripts/docs-audit/affected-docs.mjs --json 5259a35e696f56ae84b342e0864b76efba921d8a → packageMentionDocs.

@objectstack-fleet
objectstack-fleet Bot marked this pull request as ready for review October 4, 2026 10:24
@objectstack-fleet
objectstack-fleet Bot enabled auto-merge October 4, 2026 10:24
@objectstack-fleet
objectstack-fleet Bot added this pull request to the merge queue Oct 4, 2026
Merged via the queue into main with commit 3fa850c Oct 4, 2026
41 checks passed
@objectstack-fleet
objectstack-fleet Bot deleted the claude/issue-20749-test-strings-a branch October 4, 2026 10:49
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

⛔ merge queue 构建失败 — 先分诊,再决定要不要重排

队列构建 37195380707 红了。队列跑的是全量套件(PR 侧 CI 只跑 affected 子集),
所以失败的测试可能在本 PR 没碰过的包里 —— 那不是重排能修的。每次盲目重排都会让排在后面的所有 PR 重建一轮。

失败的 job(日志抽取,best effort):

  • Console Pin Gate — 失败步骤: Build the Console SPA at the pinned objectui SHA

    ✗ Built console still carries the PUBLISHED @objectstack/spec.
    

↳ 失败原因 是判读的关键:超时(Test timed out in … / Hook timed out in …)多半是负载/时序,不是本 PR 的回归;
断言(AssertionError: …)才指向真实的行为改变。两者的 FAIL 行长得一模一样,只有这一行能区分。

⚠️ 断言这一侧有一类例外,判据是断言在测什么,不是它是不是 AssertionError。 断言的对象是产品行为(一个值、一个形状、一次拒收)⇒ 照上面读:真实的行为改变,去查,⛔ 不要重排掉;
断言的对象是这次实验自身的有效性前提(跑完的耗时、负载下的先后、任何只在时间预算内才成立的条件)⇒ 它跟超时是同一类,同样对负载敏感,重排一次是合法的判别手段。
识别是机械的:断言的消息或它比较的值本身点名了一段时长、一个时间戳、一个耗时计数。实测过的一对 —— AssertionError: SecurityPlugin.init() ran: expected false to be true 测的是产品行为(真回归);
AssertionError: this run took over a second, so second-precision stamps could have differed too: expected 1006 to be less than 1000 测的是实验前提:它守护的那条不变式当时是绿的,同一个 head 原样重排一次即成功。
穿着 AssertionError 外衣的时间测量,仍然是时间测量。(⛔ 这只改「怎么读一次红」,不改「哪些测试可以重排」——后者由别处管。)

跨 PR 相同签名(24h,按失败测试文件聚合):

  • ⚠️ 本次没有可用的聚合签名(日志里没有能解析出测试文件名的 FAIL 行)—— 这不是「没有同签名的其他 PR」,是这一轮没测到。跨 PR 聚合本次不可用,请手工比对其他 PR 的同类评论。
  • ⚠️ 24h 评论账本没读完(超过 5 页仍未读到窗口尽头),所以上面的「不同 PR 数」是下界,不是全量。

历史信号:

  • 本 PR 过去 24h 无队列失败记录(首次)。
  • 过去 24h 队列共有 12 个失败构建(不含本次)。

分诊清单:

  1. 失败测试在本 PR 改动的包里 → 真回归,修 PR。
  2. 失败测试与本 PR 无关 → 看上面的「跨 PR 相同签名」;已有汇总 issue ⇒ flaky/环境问题实锤,去那张 issue 上谈,修好前重排只会再烧一轮全队列。
  3. 两者都不是 → 可能与同组 PR 语义冲突;等前面的 PR 落地或失败出队后再重排一次即可,不要连续重排。

Generated by Claude Code · merge-queue-triage workflow (#4859)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/m skip-changeset PR has no user-facing published change; bypasses the changeset gate tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants