Skip to content

Make proxy upstream response-header timeout configurable - #83

Open
mkalkere wants to merge 1 commit into
NVIDIA:developfrom
mkalkere:fix/configurable-proxy-header-timeout
Open

mkalkere wants to merge 1 commit into
NVIDIA:developfrom
mkalkere:fix/configurable-proxy-header-timeout

Conversation

@mkalkere

@mkalkere mkalkere commented Sep 14, 2026 •

Copy link
Copy Markdown

Changelog title

Proxy upstream response-header timeout is now configurable

Changelog body

  • nvpair-proxy accepts --response-header-timeout (or the NVPAIR_PROXY_RESPONSE_HEADER_TIMEOUT environment variable) to raise the 120 s default when engines queue requests or load models slowly. Invalid values log a warning and keep the default.

Bumps

  • services: patch
  • nvpair-cluster-manager: none
  • nvpair-engine-manager: none
  • nvpair-errors: none
  • nvpair-job-scheduler: none
  • nvpair-manual-nodes: none
  • nvpair-node-info: none
  • nvpair-node-scanner: none
  • nvpair-node-settings: none
  • nvpair-proxy: patch
  • nvpair-tui: none
  • nvpair-ui-broker: none
  • nvpair-workload-manager: none

Summary

Rebased onto the unified services/nvpair-proxy. The 120 s
ResponseHeaderTimeout on the upstream transport is now configurable:
--response-header-timeout flag first, then
NVPAIR_PROXY_RESPONSE_HEADER_TIMEOUT, then the 120 s default. A missing,
unparseable, or non-positive value logs a warning and keeps the default, so
a bad setting can never silently disable the timeout.

firstBodyTimeout tracks the same resolved value, since both bound the same
wait — how long an engine may take to start its work. Documented in
docs/proxy-response-header-timeout.mdx (new) and
services/nvpair-proxy/spec.md; troubleshooting entry added for the
502-after-120-seconds symptom.

Test plan

  • go build ./..., go vet ./... — clean
  • gofmt -l — clean
  • go test -race -count=1 . — full nvpair-proxy suite green
  • header_timeout_test.go: TestResolveResponseHeaderTimeout (flag >
    env > default precedence; invalid, zero, and negative values fall back),
    TestProxyTransportUsesConfiguredHeaderTimeout (the configured value
    reaches transports built after startup)

@Noah-Tervalon-Nvidia

Copy link
Copy Markdown
Collaborator

Heads up — develop has just taken a change that merges the Ollama and LM Studio
proxies into a single nvpair-proxy process, hosting one facade per enabled
engine. services/ollama-proxy/ and services/lmstudio-proxy/ are gone.

This PR edits files in those directories, so it will need updating before it can
merge. The good news is that the client-facing ollama-proxy: and
lmstudio-proxy: namespaces are unchanged — only the process behind them moved —
so the change should port across to services/nvpair-proxy/ fairly directly.

Apologies for the churn, and thanks for the contribution. Happy to help work out
where it lands if it isn't obvious.

@Noah-Tervalon-Nvidia
Noah-Tervalon-Nvidia changed the base branch from main to develop September 21, 2026 21:58
The 120 s ResponseHeaderTimeout on the upstream transport is now
overridable: --response-header-timeout flag first, then the
NVPAIR_PROXY_RESPONSE_HEADER_TIMEOUT environment variable, then the
120 s default. A missing, unparseable, or non-positive value logs a
warning and keeps the default, so a bad setting can never silently
disable the timeout. firstBodyTimeout tracks the same resolved value
because the two bound the same wait: how long an engine may take to
start its work.

Ported onto the unified nvpair-proxy after the engine proxy
unification; the per-proxy copies are gone, so the flag, the
resolver, and the test exist exactly once.

Tests: TestResolveResponseHeaderTimeout (flag > env > default
precedence, invalid/zero/negative fall back), and
TestProxyTransportUsesConfiguredHeaderTimeout (the configured value
reaches transports built after startup).

Signed-off-by: Mallikh Kaula <mallikh@users.noreply.github.com>
@mkalkere
mkalkere force-pushed the fix/configurable-proxy-header-timeout branch from ff0650b to 014b7b1 Compare September 21, 2026 22:27
@mkalkere

Copy link
Copy Markdown
Author

Rebased onto the unified layout and pushed. The configurable header timeout is ported across to services/nvpair-proxy/:

  • --response-header-timeout flag, then NVPAIR_PROXY_RESPONSE_HEADER_TIMEOUT, then the 120 s default; invalid values warn and keep the default.
  • firstBodyTimeout tracks the same resolved value, since both bound the same wait.

Details in docs/proxy-response-header-timeout.mdx.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants