Conversation
Scores buildRenderPlan on a synthetic multi-scene take (click-triggered navigation gap, typing, scene reload, end-of-take punch) at ~39fps and ~60fps source cadence, plus the real demo-recipe take recorded on main @ 013d6b8 (events + frame index only). Asserts: z <= 1.05 at scene markers and on the first new-page frame after a nav gap; >= 700ms wide rest per scene; punches reach >= 90% by the click (or are skipped); |dz| < 1e-4 per frame over the final 300ms; < 5% of frames blend different sources. Checks that fail on main run as it.fails (PENDING table); each fix commit removes its entries. Baseline on the real take: z at scene marker 1.42, z after nav gap 1.10, 0ms wide rest on scene 2, first click arrival 0%, 13.4% blended frames (52% on the 39fps synthetic). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a9fb580e2b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| const sceneStarts = markers.map((m, i) => { | ||
| if (i === 0) return 0; | ||
| const g = gaps.find((g) => g.tA >= m - MARKER_WINDOW_BEFORE_MS && g.tA <= m + MARKER_WINDOW_AFTER_MS); | ||
| return g ? g.tB : m; |
There was a problem hiding this comment.
Include unmarked navigation gaps in wide-rest checks
When navigation is triggered by a click without a scene marker—as the synthetic fixture deliberately does between 5200 ms and 5900 ms—this list contains no start for the newly loaded page. Consequently, wideRestMs only measures the take head and marked reload, so the wide-rest test can pass even if the camera is wide for only one frame after the unmarked navigation and immediately zooms again; include navigation gaps not already attributed to a marker as scene starts.
Useful? React with 👍 / 👎.
| /** the real demo-recipe take recorded on main @ 013d6b8 (events + frame | ||
| * index only — the frames themselves are not needed to score the plan) */ | ||
| function demoTake(): { log: EventLog; index: FrameIndexEntry[] } { | ||
| const dir = join(import.meta.dirname, "fixtures", "takes", "demo-main"); |
There was a problem hiding this comment.
Avoid using import.meta.dirname at the Node 20 floor
On Node 20.0–20.10, which is allowed by the package's node >=20 engine constraint, import.meta.dirname is unavailable, so join() receives undefined and this suite fails while constructing the real-demo scenario. Resolve the directory from import.meta.url with fileURLToPath instead, or raise the declared minimum Node version to 20.11.
Useful? React with 👍 / 👎.
…60fps source - Screencast JPEG q92 at DPR 2 (resolution unchanged) instead of PNG; frames are written as frames/NNNNNN.jpg. - The renderer serves frames by extension (frameMimeType), so takes recorded as PNG before this change still render. - Repaint beacon now covers the viewport at 1-2e-4 opacity (no 8-bit channel moves). The 1px corner beacon stopped registering damage on the demo dashboard once a hovered row met a timer re-setting identical text: the source fell to that timer's 20Hz for the rest of the scene. Demo recipe source fps (capture health): 52.7 -> 57.5 avg including the scene-change reload gap; no sub-60 stretch left outside the nav gap. Tests: frameMimeType unit test; record e2e asserts .jpg frames and a new e2e holds >= 45fps over a hovered-dashboard hold (was 20.7fps). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Linear cross-blending across any 25-500ms source gap (and a 350ms late crossfade across navigation/long gaps) mixed two DIFFERENT source frames: a visible double exposure on 13% of the real demo take's frames and 52% of a 39fps take's. The plan has no pixels to prove two frames near-identical, and blending near-identical frames is a visual no-op, so every gap is now floor-held and a page change is a clean cut. The blend array stays in the plan (all -1/0) so the host-page contract is unchanged. Behaviour change: the three "source cross-blend" plan tests encoded the old dissolve; they are rewritten to assert the hold/cut (short residual gap, nav gap, irregular sub-60fps source) and the dense-60fps case is kept. Motion suite: blended share 13.4% / 52.1% -> 0% (blend checks flipped). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…, settled ending Render plan: - Page boundaries = scene markers + any source gap >= 250ms (catches click-triggered navigations the event log never marks). A marker's scene opens at the first NEW frame after its reload gap, not at the marker (emitted before the goto, while frames are dropped). - Shots never outlive a boundary: the zoom-out completes 800ms before a scene change (or as the old page freezes for a click-nav); no bridging and no GLIDE_Z park on a stale focus across a boundary; the spring state snaps to wide at a real cut, hidden by the picture change. - Establishing shot (800ms at z=1) anchored to the new page's first frame. - ZOOM_LEAD 600 -> 750ms; a punch that cannot reach 90% by its event (ARRIVE 600ms) is skipped instead of landing after the click, as is one the next navigation would cut to < 300ms after the event. - Tail >= 1000ms and the take runs >= 1700ms past the last punch so the camera is at rest at the end; plan.fade drives a picture fade in/out from black with the same lengths as the music afades (shared constants). Executor: - 1000ms pre-roll at the take head and after each new page has painted, so the establishing shot plays and the first punch arrives before the first click. - The next scene's entry goto is skipped when page.url() already equals it (and the previous scene completed): no redundant reload/freeze. Behaviour change: the establishing-shot plan test asserted the camera "glides" at z > 1.05 between scenes (6000ms); it now asserts z < 1.02 there and at the next scene marker. The motion metric's click window no longer reaches into the next beat's lead-in. Motion suite (real demo take): z at scene marker 1.42 -> 1.01, z after nav gap 1.10 -> 1.00, scene-2 wide rest 0 -> 4300ms, all PENDING flipped. Real record+render: marker 1.01, arrival min 0.96, final z 1.000 at rest. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…at16 accumulation - imageSmoothingQuality "high" on both canvases (the default low/bilinear filter aliased text when downscaling the 3840px source). - Each source frame is pre-resampled once (resizeQuality high) to 1.5x the content width, so every blur pass draws a <= 1.5x downscale. - Blur pass count from the MAX displacement of all four content corners (was top-left only: up to ~7px ghost spacing on the far side during zooms), rounded up to a power of two and capped at 32. - Accumulate in a float16 canvas when available: in 8-bit, 48 passes of white at 1/48 summed to 240/255 (probed in the bundled Chromium); float16 with power-of-two passes stays within 1 level. 8-bit fallback caps at 8. blurPassCount is exported and embedded verbatim into the host page. Demo render time 2.7s -> 3.7s for 711 frames (float16 path confirmed in the render log). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…licks - typingPlan (seeded): log-normal inter-key gaps around ~100ms, x1.8 after spaces/punctuation, 45ms floor, mean compressed no lower than 60ms on a short slot (overruns shift the schedule instead of pasting), a 250-400ms beat before the first key and 250-350ms before Enter. Replaces the uniform min(90, remaining/len) metronome. - Off-screen targets are reached by an eased (cubic in-out, 350-900ms) page scroll that centres them, filmed as motion; scrollIntoViewIfNeeded stays as a backstop for nested scroll containers. Post-scroll settle 350->150ms. - The pointer rests 100ms on the target before pressing and holds the press 70ms (was press/release back to back). Tests: typingPlan unit tests; a new /keys fixture route reports pointer, press, keystroke, Enter and scroll timings to the test server, and a record e2e asserts the eased scroll (>= 5 intermediate positions), >= 80ms settle, >= 55ms hold, a beat before the first key and Enter, varied >= 35ms gaps. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…an honest budget reserve A reasoning model (DeepSeek v4-pro on the pulse dashboard) spent all 8000 max_tokens thinking and returned empty content with finish_reason "length" on every one of the 4 attempts, failing the run. Retrying at the same size mostly truncates again. - OpenAICompatibleClient doubles max_tokens after an empty answer with finish_reason "length" (8k -> 16k -> 32k), capped at escalationCeiling() = 4x the requested size. Empty answers that were not truncated retry at the same size. - If the provider rejects an escalated size with a 400, the client falls back to the largest size it accepted and stops escalating (instead of failing the run on a provider output cap). - BudgetedLlmClient reserves escalationCeiling(maxTokens) for the call's completion, so the pre-send refusal covers the escalated attempt. Tests: escalation sequence, ceiling, no escalation on non-truncated empties, 400 fallback, budget reserve of the ceiling. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Found verifying on a real generate run: a click on a nav link navigated
with NO source gap (Chromium paint holding keeps a fast local navigation
streaming at 60fps), so the gap-based boundary never fired and the camera
carried its punch into the old page's link straight onto the new page.
- Event-Log Schema v0 gains an optional, strict `navigation` event
({t, observed_t}, no URL). Additive: older parsers drop unknown types
with their existing forward-compat warning.
- The recorder logs it when a main-frame navigation request the recipe did
not issue (entry/goto navigations are wrapped and excluded) commits.
- The plan treats it as a snap boundary at the commit time (unless its
reload also left a source gap, which already covers it): the navigating
click's punch is skipped (it would last < 300ms), and the new page opens
wide with its establishing read.
- resolveFocus finds the action's own event even if a navigation was logged
after it.
Tests: schema accepts the event (and stays strict), plan cuts wide at a
gapless logged navigation, record e2e logs exactly one navigation for a
clicked link and none for the next scene's entry goto.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The "wide on the first new-page frame" metric only looked at source gaps; it now also checks the frame after each logged action-triggered navigation (gapless fast navigations). On the real generate take whose first scene clicks a nav link: main's planner 1.42 -> 1.00 on this branch. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With action-triggered navigations now logged, a bare source gap no longer has to stand in for them. Gaps near a scene marker or a logged navigation still cut at >= 250ms; an unattributed gap must be >= 500ms (above the ~400ms stall ceiling the capture e2e documents). A 300ms hiccup mid-shot on a slow machine previously truncated the punch and snapped z ~1.18 -> 1.0 with the same page still showing. Tests: new plan test (300ms unlogged stall: no discontinuity, punch kept); the gap-bridging test now logs its navigation, as the recorder does; the motion metric applies the same attribution rule. CI robustness: the same-URL e2e now counts GET /dash hits on the fixture server (1, not 2) instead of asserting a < 250ms max frame gap, and the hovered-dashboard fps floor is 35 (CI's software compositor sustains ~47; the regression was 20). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Independent review — progress log (updated as findings land)
|
… wallpaper slab The compositor offset fx(1−z) + (cx−fx)(1−1/z) moved the focus only ~30% toward centre at z=1.42 and had no bounds: on the real gen2 take the plan shoved the window off the left edge while 307px of wallpaper showed on the right (and 131px on top at another punch). cameraTransform (plan.ts) now aims the focus at the canvas centre and clamps the offset per axis: content stays fully on canvas while z·content < canvas and covers it once z·content ≥ canvas. The legal range ramps open from the plain scale-about-centre offset (g = 0 at z=1 → 1 at coverage), so wide shots stay exactly centred and a z≈1.1 glide only nudges the window. The host page embeds the same function via toString for both the blur passes and the cursor, replacing two inline copies of the old formula. Failing-first: new framing check in the motion-quality suite (one-sided wallpaper, uncovered-while-coverable, focus-centring error vs the closest feasible position) ran as it.fails on all three scenarios plus a new edge/corner-targets take; all four now pass at 0px. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A logged click-navigation with no source gap snaps the camera from a held punch straight to wide at the commit time. That time fell inside an output frame's recorded shutter half, so the compositor's motion blur averaged the whole window from z=1.42 to z=1 into one smeared flash frame — visible at ~2.15s of the PR's own gen2 render. Snaps now apply at the start of the first output frame at/after the cut, so every frame is wholly before or after the jump. For gap boundaries this is the same frame that first shows the new page. Failing-first: new maxIntraFrameDz metric (z range across one frame's shutter samples; a moving spring spans < 0.01) on a scenario shaped like gen2's first beat (hover punch, click, logged nav 80ms later, no gap) measured a straddled snap; it now stays < 0.02 on every scenario. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The arrival metric's window was cut off at the next beat's lead-in, so on main's demo take the first click's late punch (starting ~260ms AFTER the click) read as "never punched" and scored as skipped. New maxLateRise metric: zoom gained from the click to 400ms later, stopped where the next beat's 750ms lead-in could begin, so only this click's own punch counts. Branch: 0.016 on the real demo take. Mutation with main's start rule (lead 600, late start allowed, no skip): 0.060 on that same take, and the new check fails on all five scenarios (arrival alone failed on three). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…n budget BudgetedLlmClient checked the budget once per call and reserved the 4x escalation ceiling (32k at maxTokens 8000), but one call runs up to four attempts — 8k → 16k → 32k → 32k bills ~88k plus four prompts — so a single truncating call could overshoot the advertised "hard" ceiling by ~56k. The up-front 4x reserve also refused the first call outright for any --max-tokens below ~40k, even when no escalation was ever needed. The wrapper now passes the budget left as ChatOptions.spendLimit and the OpenAI-compatible client checks it before every attempt: an escalated max_tokens shrinks to the room left, and when even the requested size no longer fits it stops with TokenBudgetExceededError. The wrapper's pre-check reserves only the first attempt (prompt + maxTokens). Failing-first: a provider that truncates every attempt and bills it in full metered ~88k against a 40k budget (now ≤ 40k, TokenBudgetExceededError); a 20k budget refused the call up front (now attempts 8k, then escalates only into the remaining room). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The recipe ceiling summed only action durations and holds, but a take adds a 1s head pre-roll, ~1.5s per later scene (entry reload + 400ms settle + 1s pre-roll on the new page, at least the 1s entry allowance) and a ~1.7s settled tail. A 4-scene recipe at 58s therefore validated yet rendered ~65s, breaking the README's ≤60s promise (real runs: pulse 11.9s of recipe → 16.65s video, gen2 13.1s → 17.6s). parseRecipe now enforces estimatedTakeMs (budgets + head + per-scene change + tail) ≤ 60s. The script stage's own target (≤ 50s over ≤ 4 scenes, ≈ 57.2s estimated) still fits; an action that overruns its slot can still add a little. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The framing checks only scored coverage and one-sided exposure, so setting the ramp to full strength at every z (gx = gy = 1) survived the whole suite: wide shots would slide sideways while the focus spring returns, and a z≈1.1 glide would jam the window against an edge. Two unit tests on cameraTransform: at z=1 the offset is exactly zero for focus points on every edge and corner; at z=1.1 the smaller margin stays ≥ half the centred margin while the window still moves toward an off-centre focus. The gx = gy = 1 mutant now fails both. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Independent review: summary (head cfb4568)I checked the PR's claims against real renders (contact sheets and full-res stills of the branch vs main), real Fixed on this branch
Verdicts on the PR's claims
Residuals (not fixed)
Gates at cfb4568: typecheck ✓ · vitest 409/409 ✓ · build ✓ · npm audit 0 vulns ✓ · check:pack ✓ · CI all green, including browser-e2e. Recommendation: ready to merge. 🤖 Generated with Claude Code |
Replaces the June 760x428 13fps GIF with an unedited `generate` take of the Pulse dashboard, rendered at 1080p60 on this branch and encoded as a 1200px 30fps animated WebP (6.6MB, about the same weight as the old GIF). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Output-quality pass on capture → plan → compositor. Each item was measured first: a failing check in the new motion-quality suite or a failing e2e test, then the fix.
Items
test/motion-quality.test.ts+test/helpers/motion-metrics.tsscorebuildRenderPlanon synthetic takes (39fps and 60fps source, a click-triggered nav gap, typing, a scene reload, a punch at the end of the take) and on the real demo take from main. Checks that failed on main ran asit.failsand each fix commit flipped its own.frameMimeTypeunit; e2e checks for.jpgframes; new e2e holds ≥45fps on a hovered dashboard (was 20.7)blendarray stays (all -1/0) so the host-page contract is unchanged.page.url()already equals it. Also: fast local click-navigations leave no source gap because of paint holding, so the recorder now logs a strictnavigationevent (no URL) for navigations it didn't issue, and the plan cuts wide on it.navigationlogged for a clicked link and none for entry gotosimageSmoothingQuality="high". Each source is pre-resampled once to 1.5× the content width. Blur pass count comes from the max displacement of all 4 corners, rounded up to a power of two and capped at 32. The accumulator is float16: in the bundled Chromium, 48 passes of white at 1/48 give 240/255 in 8-bit, while float16 with power-of-two passes stays within 1 level. The 8-bit fallback caps at 8 passes.blurPassCountunit tests (the page embeds the same function viatoString)typingPlan: log-normal gaps around ~100ms, ×1.8 after spaces and punctuation, 45ms floor, mean never below 60ms (no paste), 250–400ms before the first key, 250–350ms before Enter. Eased 350–900ms page scroll that centres the target. 100ms settle before the press, 70ms hold.typingPlanunit tests; new/keysfixture reports real pointer, key and scroll timings to an e2elength×4 and the run failedfinish_reason: "length"with empty content,max_tokensdoubles (8k→16k→32k, ceiling 4×). A 400 on an escalated size falls back to the last accepted size.BudgetedLlmClientreservesescalationCeiling(maxTokens).Metrics
Regression suite: same take files, main's planner vs this branch
* In the first commit the arrival window read the next beat's lead-in, so main's first click scored 0%. The window was tightened so it no longer does. As a result it can't see main's late punch on that click, which overlapped the next beat.
Real record + render, demo recipe (
recordthenrender --music pulse)freezedetect n=-60dB d=0.5† main's planner scored on the branch's JPEG take.
The branch's freezedetect stretches are static shots, not capture stalls. The source is continuous 60fps there. They cover the typing hold and the dashboard's wide establishing shot plus the hover hold on a full-width row, which gets no punch. The dashboard's small counter changes fall under the −60dB threshold.
Real
generateon the demo app (DeepSeek v4-pro, text-only)The run includes a click on a nav link, a same-URL next scene (no reload now), and a typed form. Results: capture 58.6fps; z at scene markers 1.014; z after the logged click-navigation 1.42 → 1.00 (main planner vs branch on the same take); wide rest ≥917ms per scene; min arrival 0.96; final z 1.0001 at rest; 0% blended. ffprobe: h264 1920×1080 60/1, 17.6s. The LLM usage for this run (~23.5k tokens) did not trigger
max_tokensescalation.Behaviour changes to be aware of
maxTokens, the escalation ceiling. Analyze and script each reserve about 32k plus the prompt, so--max-tokensvalues below roughly 40–50k now refuse the first call up front.navigationevent. Old parsers drop it with their forward-compat warning. TypeScript consumers that switch exhaustively onKnownEventwill need a new case.Not verified
max_tokens. The 400 fallback is covered by unit tests only, and no real run truncated.blur accumulator: float16).🤖 Generated with Claude Code