Conversation
A user reaching an HTTPProxy whose NetworkService has no ready endpoints gets the generic "this service is temporarily unavailable" page, which reads as a platform fault. The offline page that says nothing is running behind the address exists and is configured, but cannot be reached. Envoy Gateway collapses such a route to a bodiless 503 direct_response. That short-circuits before the router filter, so the response carries no UH flag, and UH is what the offline mapper selects on. Neither response side offers a way out: local_reply_config overrides a direct_response body, and route-level request_headers_to_add never runs. Both were measured against the edge's Envoy build rather than assumed. Pointing those routes at an endpoint-less cluster restores the UH flag and the existing mapper matches unchanged. One shared cluster is used rather than one per backend, because an endpoint-less cluster carries ~108 resident data-plane stats and a per-backend cluster would scale that by the number of idle services. Key changes: - Rewrite EG's bodiless 503 direct_response routes to forward to a single shared endpoint-less cluster, preserving match, metadata and per-filter config so a governed route keeps its WAF settings - Discriminate on status 503 plus an absent body, which EG uses only for no-ready-endpoints and which leaves connector-offline routes untouched - Gate on the branded page being configured, so behaviour is unchanged when branding is off - Count rewritten routes in nso_extension_empty_backend_routes_total Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Both response side fixes were measured against the edge's Envoy build before settling on this one, because each looked plausible on paper and both fail. Setting the body on the short circuited response does not work. A route carrying an explicit body still returned the generic page, and the response length showed the route's own body was discarded rather than preferred. Marking the route with a header does not work either, at route level or virtual host level, because header changes are applied by the stage that never runs for a short circuited response. A header rule does match what the client sends, so anyone shipping that approach would let a user summon the offline page on any error by sending the header themselves. That leaves the empty backend, which restores the signal the existing rule already matches, so no change to the page selection logic was needed. On cost, an empty backend adds 8 counters per request over the short circuit, all of them cluster scoped, and every retry counter stays at zero because there is nothing to retry onto and no connection is attempted. The real cost is 108 resident stats per empty backend, which is why there is one shared backend rather than one per service. The discriminator rests on the upstream project using 503 with no body only for this case. I read that at the version production runs and at the current version. Both have two sites emitting 503 and both are the no endpoints case, one for HTTP routes and one for gRPC routes. Every other collapse uses 500, and no body is ever set on any of them. Verified end to end on the two cluster environment. A proxy with no endpoints now serves the offline page, a second idle proxy shares the same backend without duplicating it, a proxy with real endpoints still returns 200, and the tunnel offline path is untouched. |
A user reaching a proxy whose connector tunnel is down gets the generic "temporarily unavailable" page. The terse body the offline path writes has never reached a user, because the branded page overrides a direct_response body, so the string only ever appeared in a config dump. The cause is the same one behind the empty-backend case: a direct_response short-circuits before the router filter, so the response carries no UH flag for the offline page to select on. Sending user traffic to an endpoint-less cluster restores it. The comment this reverses argued an endpoint-less cluster would bring retry and connect noise. Measurement does not bear that out. With no hosts Envoy fails at host selection, so every retry counter stays at zero and no connection is attempted. The cost is resident stats, which one shared cluster holds constant. Key changes: - Route an offline tunnel's user traffic to a shared endpoint-less cluster when the branded page is configured, keeping the deterministic 503 when it is not - Leave the CONNECT route alone, since it answers the connector agent rather than a browser - Give the tunnel and empty-backend cases separate sinks so data-plane stats and the parity scanner can tell them apart - Teach both offline-route scanners the forwarding shape, so the parity gate keeps counting these routes Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
The tunnel half reverses a comment in the offline path which argued that an endpoint-less cluster would bring retry and connect noise, so that claim was measured rather than taken on trust. With no hosts Envoy fails at host selection, so every retry counter stays at zero and no connection is attempted. An endpoint-less cluster adds 8 counters per request over the short circuit, all of them cluster scoped. The real cost is 108 resident stats per cluster, which is why there are two shared sinks rather than one per service and one per connector. The two sinks are deliberately separate. Sharing one made the parity gate count an idle service as a downed tunnel, since that scanner identifies offline connector routes. Keeping them apart also lets data-plane stats say which of the two a request hit. A test asserts the two names never converge. The tunnel agent keeps its own terse reply. It is a machine client rather than a browser, so the branded page would be the wrong answer there. Both offline route scanners needed teaching the forwarding shape. Without that the parity gate would have stopped counting these routes and passed anyway, which is the quiet failure worth watching for in review. Verified on the two cluster environment with a tunnel and two idle services live at once: the gate counts 4 offline routes, both sinks exist once each, and the wider suite passes. Two things remain unverified. The empty backend case is told apart by a 503 with no body, which I read at the version production runs and at the current version, but I did not exercise the upstream project's other collapse paths to watch them emit 500. And nothing here has run on a real edge, so the propagation path is untested. |
Comment blocks added with the offline-page fix explained mechanism at length where the code already carries it. Keep the reasoning a reader cannot recover from the code and drop the rest. Also revert the operator image tags three kustomizations picked up from a local test environment. Key changes: - Reduce doc comments in the mutate, metrics and scan paths - Shorten both new chainsaw scenario headers - Restore config image tags to their committed values Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Summary
A user reaching a proxy gets a generic "temporarily unavailable" error in two cases where the platform is working fine. Their service has no running endpoints, or the tunnel behind it is down.
The branded offline page that says nothing is running behind the address exists and is configured, but the edge answers both cases in a way that skips the stage its selection rule depends on.
Sending that traffic to an empty backend restores the signal, so the existing rule matches and the right page is served.
Test plan
Related to #502