-
Notifications
You must be signed in to change notification settings - Fork 0
feat(cli): support explicit finite invocation execution budgets #741
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
Show all changes
9 commits
Select commit
Hold shift + click to select a range
84cb272
test: record explicit CLI execution budget regressions
logbie 68c4665
feat(cli): allow explicit finite invocation execution budgets
logbie 01416f4
test(cli): capture invocation deadline review regressions
logbie 370ec22
test(cli): assert terminal waits cannot outlive execution budgets
logbie a5988f1
test(cli): preserve server stream limits under short invocation overr…
logbie 028ef59
test(cli): capture duration wait resource cleanup regressions
logbie 5d78133
test(cli): observe leaked wait resources before driver cleanup
logbie 30ed946
fix(cli): enforce waits and preserve server operation budgets
logbie 358dc9e
test: bind trusted proxy fixtures to owned ephemeral ports
logbie File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
207 changes: 207 additions & 0 deletions
207
Engineering/evidence/2026-09-20-cli-execution-budget.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,207 @@ | ||
| # Finite invocation execution budgets | ||
|
|
||
| Risk class: **R3** (resource limits, CLI compatibility and subprocess lifecycle). | ||
|
|
||
| The Scriptorium full Linux suite reached the existing CLI's 300-second execution | ||
| deadline on official WFL 26.9.14 after about 34 suites, then subsequent child | ||
| launches failed against the same expired budget. The configured runner timeout | ||
| could not extend this limit. The same published runtime completed the Windows | ||
| consumer suite at `07e8adcb` in about 160 seconds: 41 suites passed and its one | ||
| deliberately failing gate test correctly made the invocation exit 1. That local | ||
| consumer log is `target/full-official-windows-red.log` in Scriptorium. This is a | ||
| real capacity failure, not a | ||
| reason to hide suite failures or exempt a batch runner with `main loop`. | ||
| The consumer Red is [Scriptorium CI35507634634](https://github.com/WebFirstLanguage/Scriptorium/actions/runs/35507634634) | ||
| at `07e8adcb`, using the published runtime rather than a development build. | ||
|
|
||
| The published source `8d82d785ea59300834de1c48b04a7ba0e187a1cd` | ||
| unconditionally applies `timeout_seconds.min(300)` in `src/main.rs`. | ||
| Configuration accepts larger values, but there is no existing file-entrypoint | ||
| override. The July 10 bind-address diary records retaining the historical cap | ||
| for compatibility. The new option leaves that default/config behavior intact. | ||
|
|
||
| ## Red evidence | ||
|
|
||
| New scenarios, fixtures and drivers are WFL. Before implementation, the three | ||
| `TestPrograms/cli_budget/*.test.wfl` programs were run with the official Windows | ||
| 26.9.14 executable (SHA-256 | ||
| `da109d5926f6af4a2f24c45150140764c406c055aef2dd7e43e45b85278f1dfe`). | ||
| The argument suite failed 5/5 expectations, deadlines passed its unchanged | ||
| configuration baseline and failed the four override cases, and server-policy | ||
| failed its one override case. All programs parsed and ran as WFL tests. The new | ||
| flag is unavailable on that release; these are CLI capability failures, not | ||
| claims that the old runtime implemented the new flag incorrectly. The existing | ||
| consumer 300-second failure establishes the underlying semantic limitation. | ||
|
|
||
| The initial draft had an invalid reserved variable name and expected the wrong | ||
| timeout wording. Those fixture mistakes were corrected before the recorded Red | ||
| run; they are not counted as product failures. Logs are retained only under | ||
| ignored `target/reports/cli-budget/red-*.log`. | ||
|
|
||
| ## Acceptance and design | ||
|
|
||
| `--execution-timeout SECONDS`, before the source filename, selects one finite | ||
| invocation deadline from 1 through 31,536,000 whole seconds. It changes only | ||
| `BudgetLimits.max_duration` before the shared budget starts. It does not change | ||
| `WflConfig`, default/config caps, request/stream limits, explicit subprocess wait | ||
| timeouts, subprocess permissions, operation/depth/size ceilings, or separately | ||
| launched WFL child configuration. Ordinary foreground operations continue to | ||
| share the invocation deadline; their duration therefore grows with an explicit | ||
| larger invocation budget. Included and executed files share that same deadline. | ||
|
|
||
| Fast WFL suites cover operands, routing, test mode, cumulative deadlines, child | ||
| configuration, parent-expiration cleanup, and unchanged main-loop HTTP timeout, | ||
| operation limits and subprocess permissions. | ||
| The cleanup test first proves the child can write under its own longer budget. | ||
| `tests/fixtures/cli_budget/long-run.test.wfl` is a real 305-second ordinary run, | ||
| invoked explicitly by both integration jobs with a 330-second budget. | ||
| It belongs outside the recursive 30-second program sweep and must not be skipped. | ||
|
|
||
| ## Green verification and review | ||
|
|
||
| Red commit `84cb272c` preserves the original executable capability tests before | ||
| the product implementation. Green adds the CLI parsing and budget selection in | ||
| `src/main.rs`; no interpreter or configuration policy implementation changes. | ||
| Further cases cover disallowed CLI modes and unchanged operation/permission | ||
| limits. While bringing the fixtures to Green, the shared-file fixture gained | ||
| the required `execute file at` syntax and the child assertion fixture gained | ||
| `--test`. These fixture corrections are not product defects or semantic Red | ||
| evidence. The later resource-policy cases also fail against the official runtime | ||
| because the flag is absent, rather than because the old policies are incorrect. | ||
|
|
||
| The release candidate reports version 26.9.15 and has SHA-256 | ||
| `1185f150c8214d982a27431d4b88b94f7f20d6ce6c718161c35120deb817650f`. | ||
| Local Windows verification at the reviewed source: | ||
|
|
||
| - `cargo build --release --locked`: passed. | ||
| - Four fast WFL suites: arguments 6/6, deadlines 5/5, server policy 1/1, | ||
| resource policy 3/3; all 15 passed. | ||
| - `wfl --execution-timeout 330 --test tests/fixtures/cli_budget/long-run.test.wfl`: | ||
| passed 1/1 after the actual 305-second wait and checkpoint. This run was not | ||
| shortened, skipped or made lifetime-exempt. | ||
| - Existing Rust compatibility suites: 139/139 passed across execution budget, | ||
| CLI help/version, lint CLI, bind-address CLI, outbound HTTP budget, subprocess, | ||
| subprocess cleanup and subprocess security. | ||
| - `cargo fmt --all -- --check` and | ||
| `cargo clippy --all-targets --all-features -- -D warnings`: passed. | ||
| - Existing documentation validation: 36/36 passed. | ||
| - Static repository hygiene after staging all new files, and `git diff --check`: | ||
| passed. | ||
| - `cargo check --locked --manifest-path fuzz/Cargo.toml -j 1 --verbose`: | ||
| passed, including `fuzz_frontend`, in 5m22s. This is compilation evidence, | ||
| not a sustained fuzz campaign. | ||
|
|
||
| The first default-parallel fuzz compilation failed while compiling unchanged | ||
| `crypto-common 0.2.2`: rustc exited 1 without an explaining compiler diagnostic. | ||
| The successful single-worker diagnostic build establishes that the candidate | ||
| fuzz workspace compiles, but does not identify the initial failure's cause. | ||
| Both logs are retained and the unresolved observation is tracked in | ||
| [issue #740](https://github.com/WebFirstLanguage/wfl/issues/740). No dependency, | ||
| lockfile or source change was made between those builds; the issue does not | ||
| waive any check. The host used stable x86_64-pc-windows-msvc Rust 1.98.1. | ||
|
|
||
| Logs remain under ignored `target/reports/cli-budget/`. An independent reviewer | ||
| checked flag operands and placement, CLI modes and child argv, shared deadlines, | ||
| configured HTTP/operation/permission policies and cancellation cleanup. The | ||
| review's two fixture findings were fixed: the owned child now has its own longer | ||
| budget plus a positive control, and the child assertion program is launched in | ||
| test mode. Final source and fixture review identified no blocking finding. | ||
|
|
||
| Exact-head CI must still complete, including both real five-minute boundary | ||
| steps and the ordinary Linux/Windows program sweeps. No merge or release is | ||
| authorized by this evidence record itself. | ||
|
|
||
| ## Review follow-up Red | ||
|
|
||
| Automated review of `68c46650` exposed three additional semantic cases, verified | ||
| locally before their fixes in `TestPrograms/cli_budget/review-boundaries.test.wfl`. | ||
| The WFL suite passed its existing main-loop wait exemption baseline and failed | ||
| three real expectations: an ordinary five-second wait outlived a one-second | ||
| deadline, a one-second invocation override shortened a ten-second server HTTP | ||
| policy, and dump modes accepted a misplaced timeout after the source filename. | ||
| The log is `target/reports/cli-budget/red-review-boundaries.log` (1/4 passed). | ||
| The wait gap predates the CLI option; it matters to an explicit finite deadline | ||
| and cannot be hidden by adding a later operation checkpoint to the assertion. | ||
| These findings supersede the earlier source-review verdict until corrected. | ||
| An additional last-statement wait case then confirmed successful exit after an | ||
| expired deadline without any following checkpoint; the expanded Red was 1/5 | ||
| (`red-review-boundaries-final-wait.log`). Existing embedded-runtime tests also | ||
| establish that a manually supplied 250ms budget bounds main-loop HTTP requests, | ||
| so the remedy must distinguish a CLI invocation override without changing that | ||
| existing API contract. | ||
|
|
||
| Additional WFL streamed-response tests failed 0/2 on the frozen `68c46650` | ||
| binary, covering delayed headers and delayed body reads under config 10s / CLI | ||
| 1s. Their first draft used a reserved parameter name and an ungrouped list | ||
| expression; those fixture parse errors were corrected before the recorded | ||
| semantic Red. An independently authored resource suite also failed its ordinary | ||
| wait and active WebSocket wait cases while passing the main-loop exemption | ||
| baseline (1/3). Red commits are `01416f4d`, `370ec22a`, `a5988f13`, and `028ef59e`. | ||
|
|
||
| The final remedy retains the original main-loop operation duration privately in | ||
| `ExecutionBudget`. Only `with_invocation_timeout` replaces run lifetime; existing | ||
| explicit-budget constructors retain their behavior. HTTP continues choosing the | ||
| minimum of that operation duration and the configured/remaining stream limit. | ||
| Duration waits now check eagerly before/after the wait and each WebSocket pump | ||
| iteration. Passive sleeping and receiving use at most 10ms intervals; handler | ||
| dispatch is awaited normally so WFL finally blocks and interpreter state unwind. | ||
| A reviewed intermediate outer-select design was rejected because dropping an | ||
| arbitrary running handler could skip that cleanup. Its successful checks are | ||
| not final-source acceptance evidence. Dump modes scan only for a misplaced new | ||
| option, preserving the handling of unrelated trailing arguments. | ||
|
|
||
| The resource regression was strengthened in `5d781336` after review found that | ||
| driver-side process reaping/closing could mask a surviving descendant. The | ||
| fixture now establishes writer readiness, releases it after the owner's | ||
| deadline, and observes the marker before any parent poll/reap. Exact-port | ||
| WebSocket rebind also happens before the driver's cleanup. The final frozen | ||
| old-binary result is 1/3: the late-write assertion fails with an actual pre-reap | ||
| write, the ordinary WebSocket wait outlives the deadline, and the exempt server | ||
| case passes. The unchanged final-source suite passes 3/3 in about eight seconds. | ||
| This covers an idle active WebSocket receiver; queued handler traffic remains | ||
| covered by the existing Rust WebSocket suite, not by this new WFL fixture. | ||
|
|
||
| The final reviewed no-drop candidate SHA-256 is | ||
| `c0619c544b09551ce7988f58a564f50ff04e48a5e94a6f903c08a141b911c516`. | ||
| At this source, all seven fast WFL suites pass 25/25. `cargo test --all --locked` | ||
| passes 2,414 tests, with 27 existing ignored tests across 175 result records; | ||
| this includes the unchanged custom-250ms main-loop HTTP contract, stream tests, | ||
| WebSocket tests and CLI compatibility. Release build, fmt, strict Clippy, | ||
| fuzz-target compilation, and documentation validation (36/36) pass. Local logs | ||
| are the `*-final*` and `wait-resources-green.log` files under the same ignored | ||
| report directory. Final source review found no remaining blocking issue. | ||
|
|
||
| The first remote run for `68c46650`, | ||
| [CI 35509162373](https://github.com/WebFirstLanguage/wfl/actions/runs/35509162373), | ||
| is not acceptance evidence: Windows integration failed in unchanged | ||
| `trusted_proxy_test::default_and_untrusted_peers_ignore_forged_forwarding` when | ||
| binding port 53684 reported Windows address-in-use error 10048. Its Linux | ||
| integration sibling was then cancelled, so neither long-boundary step ran. | ||
| Both ordinary program sweeps and the other completed jobs passed. The fixture | ||
| obtains a free port before spawning the server, which leaves a reuse window; | ||
| the observed log does not establish who occupied it. The final local full Rust | ||
| run passes that test unchanged. A new exact-head CI run must pass all gates; | ||
| the earlier failure is retained rather than treated as a waiver. | ||
|
|
||
| The final no-drop candidate also passes the actual 305-second WFL boundary | ||
| 1/1 with `--execution-timeout 330`; `green-final-long.log` records the result. | ||
|
|
||
| The final candidate also completes Scriptorium's unchanged full functional | ||
| suite at clean consumer commit `df8039cc252e2e48772ed88b9d273b98e30a67c5`: | ||
| 42 suites, 41 functional passes and the sole deliberate failure-propagation | ||
| fixture produces exit 1, using `--execution-timeout 1200`. Its local consumer | ||
| log `target/full-cli-budget-reviewed-red.log` spans about 202 seconds. This | ||
| particular rerun does not exceed 300 seconds. The earlier initial-candidate | ||
| consumer run at `2d1d6e2` plus command/docs changes spans 509 seconds with the | ||
| same 41-pass/1-deliberate-failure outcome, and the final runtime's dedicated | ||
| 305-second WFL test separately verifies the longer-boundary behavior. These | ||
| host-dependent run durations are evidence, not a performance benchmark. | ||
|
|
||
| The proxy fixture's port race is remedied without retries or weakened tests: | ||
| its WFL server binds port zero and reports the actual bound address, and the | ||
| existing Rust harness reads a complete readiness record and validates loopback | ||
| address/nonzero port before connecting. All seven existing test bodies, | ||
| assertions, security settings and cleanup remain unchanged. The updated fixture | ||
| passes 7/7 locally at the final product source in the normal debug harness; | ||
| the release candidate is unchanged. This edits an existing fixture only and | ||
| adds no Rust test scenarios or test drivers. | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,38 @@ | ||
| # A finite execution budget for long batch invocations | ||
|
|
||
| The complete Scriptorium WFL suite exceeded the CLI's historical 300-second | ||
| cap on Linux. Increasing `.wflcfg` could not help because the cap was applied | ||
| after loading configuration. Making the test runner a lifetime-exempt server | ||
| loop would hide the boundary instead of supporting legitimate batch work. | ||
|
|
||
| `wfl --execution-timeout 1200 batch.wfl` now grants one explicit finite deadline | ||
| to that invocation. Only `BudgetLimits.max_duration` changes; the interpreter | ||
| still receives its original capped configuration. This separation preserves | ||
| HTTP and streaming timeouts, main-loop process waits, explicit process-wait | ||
| deadlines, permissions and memory/operation/depth limits. Ordinary foreground | ||
| operations remain part of the same invocation budget, and executed files share | ||
| the same budget rather than resetting it. Child WFL processes retain their own | ||
| configuration unless the caller explicitly passes an override to them. | ||
|
|
||
| The option accepts whole seconds from one through one year, rejects duplicate | ||
| or malformed operands, and must precede the source filename. Normal scripts and | ||
| test-mode scripts use the same option. Configuration maintenance, editor launch | ||
| and environment dumps refuse an irrelevant override instead of ignoring it. | ||
|
|
||
| All new scenarios and drivers are WFL. Fast tests cover CLI routing, the existing | ||
| configuration baseline, longer/shorter deadlines, child cleanup and isolation, | ||
| and unchanged HTTP limits. A separate 305-second WFL case in both integration | ||
| jobs proves execution past the old ceiling with an explicit 330-second bound. | ||
| The existing outer CI bound remains finite. See the matching engineering | ||
| evidence record for Red/Green results, review and exact-head CI. | ||
|
|
||
| Review exposed two policy boundaries that needed stronger tests: a shorter | ||
| invocation override must not shorten main-loop HTTP or streamed responses, and | ||
| a duration wait must observe the deadline even when it is the final statement. | ||
| The budget now retains its original operation duration separately when the CLI | ||
| overrides invocation lifetime, preserving existing embedded custom budgets. | ||
| Duration waits check eagerly and poll passive sleeping/receiving in bounded | ||
| intervals. WebSocket handlers are awaited normally so errors run their cleanup | ||
| and unwind interpreter state; the runtime does not cancel a handler future to | ||
| interrupt a duration wait. Dump modes reject the misplaced new option while | ||
| preserving their handling of unrelated trailing arguments. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,7 @@ | ||
| # Trusted CLI conformance drivers own only synthetic child programs. | ||
| timeout_seconds = 60 | ||
| execution_logging = false | ||
| debug_report_enabled = false | ||
| allow_shell_execution = true | ||
| shell_execution_mode = sanitized | ||
| kill_on_shutdown = true |
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🔍 Exact-head CI evidence remains pending
The evidence record leaves exact-head CI pending, including both five-minute boundary runs. Confirm every required job passed before merge.
Was this helpful? React with 👍 or 👎 to provide feedback.