Skip to content

fix: Show the offline page when a backend or tunnel is down - #503

Draft
scotwells wants to merge 3 commits into
mainfrom
fix/offline-page-empty-backend
Draft

scotwells wants to merge 3 commits into
mainfrom
fix/offline-page-empty-backend

Conversation

@scotwells

@scotwells scotwells commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Summary

A user reaching a proxy gets a generic "temporarily unavailable" error in two cases where the platform is working fine. Their service has no running endpoints, or the tunnel behind it is down.

The branded offline page that says nothing is running behind the address exists and is configured, but the edge answers both cases in a way that skips the stage its selection rule depends on.

Sending that traffic to an empty backend restores the signal, so the existing rule matches and the right page is served.

Test plan

  • A proxy with no running endpoints serves the offline page rather than the generic error
  • A proxy whose tunnel is down serves it too, while the tunnel agent still gets its own terse reply
  • Idle services and downed tunnels each use one shared sink, however many exist
  • Build, lint, tests and the config parity gate pass

Related to #502

A user reaching an HTTPProxy whose NetworkService has no ready endpoints
gets the generic "this service is temporarily unavailable" page, which
reads as a platform fault. The offline page that says nothing is running
behind the address exists and is configured, but cannot be reached.

Envoy Gateway collapses such a route to a bodiless 503 direct_response.
That short-circuits before the router filter, so the response carries no
UH flag, and UH is what the offline mapper selects on. Neither response
side offers a way out: local_reply_config overrides a direct_response
body, and route-level request_headers_to_add never runs. Both were
measured against the edge's Envoy build rather than assumed.

Pointing those routes at an endpoint-less cluster restores the UH flag and
the existing mapper matches unchanged. One shared cluster is used rather
than one per backend, because an endpoint-less cluster carries ~108
resident data-plane stats and a per-backend cluster would scale that by
the number of idle services.

Key changes:
- Rewrite EG's bodiless 503 direct_response routes to forward to a single
  shared endpoint-less cluster, preserving match, metadata and per-filter
  config so a governed route keeps its WAF settings
- Discriminate on status 503 plus an absent body, which EG uses only for
  no-ready-endpoints and which leaves connector-offline routes untouched
- Gate on the branded page being configured, so behaviour is unchanged
  when branding is off
- Count rewritten routes in nso_extension_empty_backend_routes_total

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@scotwells

Copy link
Copy Markdown
Contributor Author

Both response side fixes were measured against the edge's Envoy build before settling on this one, because each looked plausible on paper and both fail.

Setting the body on the short circuited response does not work. A route carrying an explicit body still returned the generic page, and the response length showed the route's own body was discarded rather than preferred. Marking the route with a header does not work either, at route level or virtual host level, because header changes are applied by the stage that never runs for a short circuited response. A header rule does match what the client sends, so anyone shipping that approach would let a user summon the offline page on any error by sending the header themselves.

That leaves the empty backend, which restores the signal the existing rule already matches, so no change to the page selection logic was needed.

On cost, an empty backend adds 8 counters per request over the short circuit, all of them cluster scoped, and every retry counter stays at zero because there is nothing to retry onto and no connection is attempted. The real cost is 108 resident stats per empty backend, which is why there is one shared backend rather than one per service.

The discriminator rests on the upstream project using 503 with no body only for this case. I read that at the version production runs and at the current version. Both have two sites emitting 503 and both are the no endpoints case, one for HTTP routes and one for gRPC routes. Every other collapse uses 500, and no body is ever set on any of them.

Verified end to end on the two cluster environment. A proxy with no endpoints now serves the offline page, a second idle proxy shares the same backend without duplicating it, a proxy with real endpoints still returns 200, and the tunnel offline path is untouched.

A user reaching a proxy whose connector tunnel is down gets the generic
"temporarily unavailable" page. The terse body the offline path writes has
never reached a user, because the branded page overrides a direct_response
body, so the string only ever appeared in a config dump.

The cause is the same one behind the empty-backend case: a direct_response
short-circuits before the router filter, so the response carries no UH flag
for the offline page to select on. Sending user traffic to an endpoint-less
cluster restores it.

The comment this reverses argued an endpoint-less cluster would bring retry
and connect noise. Measurement does not bear that out. With no hosts Envoy
fails at host selection, so every retry counter stays at zero and no
connection is attempted. The cost is resident stats, which one shared
cluster holds constant.

Key changes:
- Route an offline tunnel's user traffic to a shared endpoint-less cluster
  when the branded page is configured, keeping the deterministic 503 when
  it is not
- Leave the CONNECT route alone, since it answers the connector agent
  rather than a browser
- Give the tunnel and empty-backend cases separate sinks so data-plane
  stats and the parity scanner can tell them apart
- Teach both offline-route scanners the forwarding shape, so the parity
  gate keeps counting these routes

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@scotwells scotwells changed the title fix: Serve the offline page when a backend has no endpoints fix: Show the offline page when a backend or tunnel is down Sep 25, 2026
@scotwells

Copy link
Copy Markdown
Contributor Author

The tunnel half reverses a comment in the offline path which argued that an endpoint-less cluster would bring retry and connect noise, so that claim was measured rather than taken on trust. With no hosts Envoy fails at host selection, so every retry counter stays at zero and no connection is attempted. An endpoint-less cluster adds 8 counters per request over the short circuit, all of them cluster scoped. The real cost is 108 resident stats per cluster, which is why there are two shared sinks rather than one per service and one per connector.

The two sinks are deliberately separate. Sharing one made the parity gate count an idle service as a downed tunnel, since that scanner identifies offline connector routes. Keeping them apart also lets data-plane stats say which of the two a request hit. A test asserts the two names never converge.

The tunnel agent keeps its own terse reply. It is a machine client rather than a browser, so the branded page would be the wrong answer there.

Both offline route scanners needed teaching the forwarding shape. Without that the parity gate would have stopped counting these routes and passed anyway, which is the quiet failure worth watching for in review. Verified on the two cluster environment with a tunnel and two idle services live at once: the gate counts 4 offline routes, both sinks exist once each, and the wider suite passes.

Two things remain unverified. The empty backend case is told apart by a 503 with no body, which I read at the version production runs and at the current version, but I did not exercise the upstream project's other collapse paths to watch them emit 500. And nothing here has run on a real edge, so the propagation path is untested.

Comment blocks added with the offline-page fix explained mechanism at
length where the code already carries it. Keep the reasoning a reader
cannot recover from the code and drop the rest.

Also revert the operator image tags three kustomizations picked up from
a local test environment.

Key changes:
- Reduce doc comments in the mutate, metrics and scan paths
- Shorten both new chainsaw scenario headers
- Restore config image tags to their committed values

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant