Improve CPU histogram execution and extend application evaluation - #24
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 5c618497b3
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if digest(input_path) != frozen["artifacts"]["worker-input"]["sha256"]: | ||
| raise ValueError("worker packet differs from frozen hash") |
There was a problem hiding this comment.
Select the Formula packet when replaying A12
When this helper receives an A12 summary emitted by openboost_worker_smoke.py, the job's input_npz points to the dedicated formula-input artifact, while this unconditional check compares it with the ordinary worker-input hash. Those packets differ because the ordinary packet retains age as a predictor, so the helper raises before replaying any A12 cell. Select the expected artifact role based on the application.
Useful? React with 👍 / 👎.
| np.savez( | ||
| out / "features.npz", | ||
| x=arrays["x_validation"], | ||
| row_ids=arrays["validation_row_ids"], | ||
| ) |
There was a problem hiding this comment.
Carry A7 exposure into replay feature packets
When replaying an A7 summary produced by openboost_worker_smoke.py, fitting succeeds from the original packet, but this reconstructed inference packet drops exposure_validation. openboost_predict.py explicitly requires an exposure field for A7, so its subprocess exits nonzero and every otherwise successful A7 replay is recorded as an error. Include the application-specific exposure alongside the features.
Useful? React with 👍 / 👎.
Summary
Full Covertype validation previously exceeded the fixed 90-second fit cap. Hoist invariant candidate row hashing and reuse selected histogram columns while preserving exact candidate identities and reduction behavior. All five frozen Covertype folds now finish within the unchanged cap and reproduce predictions exactly in fresh processes.
Extend the current CPU evaluation workers for quantiles, counts/exposure, severity, aggregate loss, matched frequency-severity composition, fixed-scale survival, structured Formula and query-aware ranking. Commit five-fold validation and replay artifacts for the real-data integrations; ranking currently has synthetic adapter checks only. Record the CPU exit audit, remaining gates and first GPU implementation scope in Sprint 062.
Validation
Scope and remaining work
These bounded validation runs do not establish baseline quality or speed parity, formal authoring/adoption acceptance, or completion of v1. CUDA remains unimplemented. Real MSLR binding, real searches, joint composition selection, installed D5 probes and formal gate reconciliation remain open. Histogram reuse increases temporary contiguous storage; the tradeoff is documented with the bounded profile evidence.