Skip to content

chore(api): log client-aborted requests in responseTimeLogger - #1820

Merged
ClemRz merged 3 commits into
developfrom
fix/1818-log-client-aborts
Sep 28, 2026
Merged

ClemRz merged 3 commits into
developfrom
fix/1818-log-client-aborts

Conversation

@ClemRz

@ClemRz ClemRz commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

🤔 What

Makes client-aborted requests visible in the API's own logs. responseTimeLogger gains a res.on('close') branch alongside the existing finish one.

22 added lines in config/http.js plus one integration test, no behaviour change for normal requests.

🤷‍♂️ Why

  • Aborts were completely invisible from the app side. res.on('finish') never fires when a client disconnects mid-flight. An aborted request logged a Req :: line on the way in and nothing on the way out, so the 499 existed only in the reverse proxy's access log. chore(api): log client-aborted requests (499) in responseTimeLogger #1818 had to be reconstructed from proxy logs for exactly this reason.
  • That blind spot caused a misdiagnosis. chore(api): log client-aborted requests (499) in responseTimeLogger #1818 was originally filed as perf(api) against the four full-world bbox geoloc coordinate endpoints. Curling production disproved it — all four return 200 and complete reliably (warm connection: entrancesCoordinates 308ms TTFB / 822ms total for 133 970 points, massifsCoordinates 170/324ms, organizations 417/420ms, networksCoordinates 627/628ms; all four in parallel, 3.09s wall clock, 4× 200). The recorded aborts range from 1ms to 1109ms, and nobody abandons a 2ms request.
  • The aborts are ordinary client behaviour. There is no AbortController in the map fetch path (packages/web-app/src/actions/Map.js), each session aborts all four endpoints in the same millisecond — the signature of leaving the page — and 1 of the 4 observed sessions is HeadlessChrome/152. The "27 aborts" are ~4 events × 4 endpoints.
  • It generalises. A client can abandon any request. This belongs in shared middleware, not in the endpoints that happened to surface it.

Worth stating plainly: this is a hygiene fix, not a stability or performance fix. Nothing was broken by the abort — the handler runs to completion and the late write is silently discarded. Nothing here makes anything faster.

🔍 How

res.on('close', () => {
  if (res.writableEnded) return;
  // … log 499 at info with elapsed ms
});

Four decisions worth a reviewer's attention:

res.writableEnded is the discriminator. A normal response emits close too, immediately after finish. Verified ordering: on completion close fires with writableEnded === true; on abort it fires with writableEnded === false. So the early return is the whole mechanism, and because finish never fires after an abort there is no double-logging in either direction.

Elapsed time is tracked from middleware entry, not read from X-Response-Time. An aborted response often never writes its headers, so that header is unreliable precisely in the case being logged. The elapsed value is the field that separates a benign instant disconnect from a genuinely slow response someone gave up on — the distinction that was missing when #1818 was triaged.

info, not error. An abort is not a server fault and must not count against the 5xx error budget, so it deliberately does not reuse the statusCode >= 500 branch. It also does not dump the request body the way the 5xx path does.

The literal 499. Not a real HTTP status — it is the nginx convention for "client closed request". Reusing it lets these lines be grepped against the proxy access log that was previously the only record.

🧪 Testing

test/integration/4_routes/client-abort-logging.test.js — 7 tests, all passing. It requires the real responseTimeLogger out of config/http.js and mounts it on a throwaway Express app, with a handler that responds 300ms after it is reached and a client that destroys its socket 100ms after it is reached. Same directory as traceId.test.js and security-headers.test.js, which cover config/http.js middleware the same way.

A throwaway app rather than supertest against the lifted server: the abort has to land while the handler is still running, so the handler's duration has to be controlled, and supertest offers no way to destroy the socket at a chosen moment.

Asserted on abort — one line, tagged 499, at info and not error, elapsed ms present, and no second line from the late discarded write. Asserted on normal completion — still exactly one line, tagged 200, not misreported as an abort (this is the test that fails if the res.writableEnded guard is dropped).

CASE 1 normal:  Res :: GET /normal 200 301.656ms
CASE 2 abort:   Res :: (client aborted) GET /slow 499 101ms

Both client timers are anchored on a signal from the handler, not on http.get() (thanks @Paul-AUB). The terminal handler runs strictly after responseTimeLogger, so when it signals, the close listener is registered and startedAt is recorded. Anchoring on http.get() instead meant a connect slower than the 100ms abort delay would destroy the socket before the middleware ever ran, logging nothing and failing the test for a reason unrelated to the middleware. Reproduced by setting the abort delay to 0 — 5 failures, empty capture — then 7 passing on that same configuration after the fix. The settle window is anchored the same way, so a slow connect cannot eat the margin that exists to catch a second log line. A request that never reaches the handler now rejects with an explicit message rather than hanging to mocha's timeout.

Also confirmed the underlying premise directly — after the abort, finish never fires, the handler still runs to completion, and the late write throws nothing:

[req aborted] fired
[res close] fired (writableEnded=false)
WORK STILL RAN to completion. req.aborted=true
write attempted: no throw (silently discarded)
=> finish fired? false

eslint and prettier --check clean on both changed files.

Full suite: 3483 passing, 0 failing. That reconciles exactly against the develop baseline of 3471 passing / 5 failing — 3476 tests, plus the 7 new ones. The 5 baseline failures (shard 0: 1, shard 3: 4, all in Changes/get-recent-comment-relevance-swap.test.js, which passes standalone) happened to pass this run, confirming they are pre-existing parallel-shard DB contention rather than anything related to this PR. Not addressed here.

📸 Previews

n/a — log output only, shown above.

Related, deliberately not in this PR

  • networksCoordinates costs ~630ms server-side to return 455 points / 6.9 KB — the slowest TTFB of the four while returning the least data. It and organizations are the two coordinate endpoints without a snapshot service. Reviewed and accepted as-is; no follow-up issue.
  • Cancelling in-flight work on abort. The DB query completes and the payload is fully serialised before being discarded — 3.8 MB thrown away on entrancesCoordinates. True cancellation needs pg_cancel_backend over a second connection; not worth that complexity at ~4 abort events/day.
  • The full-world bbox is by design and was left alone. Map.js:57-64 documents it: bulk coordinates fetched once at startup, feeding a supercluster kD-tree so panning needs no further API calls. Capping bbox area on these four, as chore(api): log client-aborted requests (499) in responseTimeLogger #1818 originally proposed, would break the map. (perf(db): geoloc entrances sort spills to disk on large bounding boxes (work_mem undersized) #1812's cap applies to GET /geoloc/entrances, a different bounded endpoint.)
  • Stale benchmark comment in Map.js:57-64 claims "~2.6 MB uncompressed, ~700 KB gzipped"; actual is 3.82 MB / 1.51 MB. Frontend repo.
  • Brotli yields nothing over gzip on these payloads (1 511 849 vs 1 504 117 bytes — marginally worse), suggesting a low quality level for dynamic responses. Infrastructure, not code.

@ClemRz
ClemRz requested a review from Paul-AUB September 24, 2026 22:42
@ClemRz ClemRz self-assigned this Sep 24, 2026
Paul-AUB
Paul-AUB previously approved these changes Sep 25, 2026
@ClemRz

ClemRz commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

@Paul-AUB — heads-up that the review above appears to have landed on the wrong PR, so I haven't actioned any of it.

The review header states its own target:

Review target: branch fix/document-cyclic-parent vs origin/develop (source: manual — PR 1820's head branch; upstream selection was misconfigured)

This PR is a single file — config/http.js, +22/-0 — adding a res.on('close') branch to responseTimeLogger so client aborts stop being invisible in the app logs. None of the files in the review are part of it:

  • sql/9_20_2026_08_18_document_no_self_parent_constraint.sql
  • api/controllers/v1/document/update.js
  • api/controllers/v1/document/update-with-new-entities.js
  • config/constants/document.js

Three things that may help track down the tooling issue:

  • fix/document-cyclic-parent does not exist on origin, so it looks like a stale local branch.
  • That work is already on develop, merged as fdb68604 fix(document): prevent cyclic parent relationships. The four findings also look like they were addressed at the time, in 42617c82 fix(document): address code review findings (2026-08-19) — so they are probably resolved rather than outstanding. Worth re-checking against current develop before re-filing them anywhere.
  • The note that gh was unauthenticated locally means the PR description was never fetched, which is likely why the branch mismatch went unnoticed.

Since the APPROVED state reflects that other diff rather than this one, could you re-run against fix/1818-log-client-aborts when you have a moment? I'm pushing a test commit shortly, so there will be a second commit to look at anyway.

responseTimeLogger only logs on res.on('finish'), which never fires when a
client disconnects mid-flight. An aborted request therefore logged a
'Req ::' line on the way in and nothing on the way out, leaving the 499
visible only in the reverse proxy's access log. #1818 had to be
reconstructed from proxy logs for exactly this reason, and the missing
timing is what made benign instant disconnects look like slow responses:
the aborts recorded there range from 1ms to 1109ms, and nobody abandons a
2ms request.

Nothing was broken by the abort itself. The handler runs to completion and
the late write is silently discarded, so this closes an observability gap
rather than a fault.

- Log on res.on('close') when res.writableEnded is false; normal
  completion also emits 'close', right after 'finish', so that guard is
  what separates the two
- Report the literal 499, the nginx convention for "client closed
  request", so these lines correlate with the proxy's access log
- Log at info, not error: an abort is ordinary client behaviour and must
  not count against the 5xx error budget
- Track elapsed time from middleware entry instead of reading
  X-Response-Time, which is unreliable here because an aborted response
  often never writes its headers

Applies to every route through the shared middleware, since a client can
abandon any request, not just the full-world geoloc coordinate dumps that
surfaced it.

Closes #1818
- Add integration test mounting the real responseTimeLogger middleware
- Assert an aborted request logs one line tagged 499 at info level
- Assert the elapsed wait is reported and no second line is emitted
- Assert a normal completion is still logged once and not as an abort
@ClemRz
ClemRz force-pushed the fix/1818-log-client-aborts branch from 7112fcb to 0050c50 Compare September 28, 2026 21:56
Paul-AUB
Paul-AUB previously approved these changes Sep 28, 2026

@Paul-AUB Paul-AUB left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggestions (Should Consider)

  1. [test/integration/4_routes/client-abort-logging.test.js:63] Consider starting the 100 ms abort timer after the server receives the request. It currently starts immediately after http.get(), so a delayed connection or a busy CI worker can destroy the socket before responseTimeLogger runs, leaving lines empty and making the test fail for reasons unrelated to the middleware. A signal from the handler would make the disconnect deterministic.

The close handler and its writableEnded guard behave as intended in the covered abort and normal-response cases. All seven new cases passed in an isolated Express run using the real middleware. The full integration bootstrap could not run locally because PostgreSQL refused the connection.

- Signal from the terminal handler once responseTimeLogger has run
- Start the abort and settle timers from that signal, not from http.get
- Reject with an explicit message if the handler is never reached

A slow connect could previously destroy the socket before the middleware
ran, logging nothing and failing the test for an unrelated reason.
@ClemRz

ClemRz commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

Good catch, and it was a real bug rather than a theoretical one. Fixed in 766040e.

I reproduced it before fixing it, by shrinking ABORT_AFTER from 100 to 0 so the destroy reliably lands before the server sees the request — which is the same end state a stalled connect on a loaded worker produces:

5 failing
  1) should log the aborted request:
     AssertionError: expected Array [] to have property length of 1 (got 0)
  2) should tag the abort with status 499:
     TypeError: Cannot read properties of undefined (reading 'line')
  ...

Empty lines, exactly as you predicted, and the failure message points nowhere near the actual cause — which is the worst part of it.

The fix is the handler signal you suggested. The terminal middleware runs strictly after responseTimeLogger, so by the time it fires, the close listener is registered and startedAt is recorded:

app.use((req, res) => {
  handlerReached.emit('reached');
  setTimeout(() => res.json([1, 2, 3]), HANDLER_DELAY);
});

Two details beyond the literal suggestion:

The settle timer is anchored on the signal too, not just the abort timer. It was also measured from http.get(), so a slow connect ate into the window that exists to catch a second log line from the late discarded write. Anchoring only the abort would have left should not also log a completed response able to pass for the wrong reason.

A missed signal now fails loudly. If the request never reaches the handler the promise rejects with Request never reached the handler after 5s, rather than hanging to mocha's 10s timeout.

Verification, on the same ABORT_AFTER = 0 configuration that produced the 5 failures above: 7 passing. Then restored to 100 and re-ran: 7 passing. Full suite on the rebased branch head: 3483 passing, 0 failing.

On the local bootstrap — the suite needs a PostgreSQL instance; npm run test:snapshot seeds the template database the bootstrap restores from, which is usually what's missing when the connection is refused.

@ClemRz
ClemRz merged commit bdda763 into develop Sep 28, 2026
3 checks passed
@ClemRz
ClemRz deleted the fix/1818-log-client-aborts branch September 28, 2026 22:29

@Paul-AUB Paul-AUB left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the current PR head (766040ee) against develop, including the change since the previous review.

The handler signal now starts both the abort and settle timers after responseTimeLogger has registered its listeners. The explicit timeout also makes a request that never reaches the handler fail clearly. This resolves my previous test timing suggestion; I found no remaining issues.

All seven focused tests passed locally against the real middleware in the isolated Express harness. The repository's full integration bootstrap was unavailable locally because PostgreSQL refused the connection.

This branch was successfully deployed

1 active deployment
build — 766040ee Deployed Sep 28, 2026 by ClemRz via build-test #3883
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

chore(api): log client-aborted requests (499) in responseTimeLogger

2 participants