us.proxy is flapping operational → degraded → outage on the public status page while the proxy answers in 3.5 ms. The proxy is not the problem; the probe route is.
What is actually being measured
The status node probes the US proxy at proxy1.ciris-services-1.ai. That record is Cloudflare-proxied:
Cloudflare: A proxy1.ciris-services-1.ai -> 108.61.242.236 proxied=True
resolves to: 104.21.42.160, 172.67.163.135 (CF edge)
So a probe originating on the origin host leaves the machine, reaches Cloudflare's edge, and comes back to the same machine. Measured from that host, three ways:
direct to the container 3.4 – 4.3 ms stable
via proxy1.ciris-services-1.ai 136 – 1502 ms ~10x spread <- what you measure
via proxy1.ciris-services-2.ai 515 – 594 ms stable (EU)
Any threshold inside that 136–1502 ms band will flap, and the recorded transitions show exactly that — operational -> outage at 15:04:21, outage -> degraded 62 seconds later, with upstream_status: operational throughout and every LLM provider sourced from cirisproxy.us reading healthy in the same payload.
Confirmed from the status container's live sockets that this is the route in use — connections to 172.67.163.135:443 and 104.21.42.160:443, i.e. CF edge, not an origin or a docker name.
EU is only healthy by accident
proxy1.ciris-services-2.ai is grey-clouded — it resolves straight to 46.224.81.217. So the EU probe already does the right thing: it reaches the service. US and EU latency figures are therefore not comparable and never have been; one includes a CDN round trip and the other does not.
A direct route now exists
Bridge-side, ciris-proxy has been added to the lens_cirislens network, which ciris-status is already on. Verified after the change, from inside the status node's network namespace:
docker DNS: ciris-proxy -> 172.20.0.5 (also resolves as `proxy`)
probe: http=200 3.2 ms / 3.8 ms / 5.0 ms
So http://ciris-proxy:4000/health/liveliness is available and is ~400x faster and far more stable than the current target. This did not change what the adapter probes — that is the ask here.
The ask
Make the probe target operator-configurable, so the deployment can point a component at the service rather than at whatever public route happens to exist.
On mechanism, one thing worth stating plainly rather than assuming: the operator's preference is environment variables set by the bridge. That conflicts with DEPLOY.md §2's zero-env contract, which we have been careful to keep — it is why --diagnostics was built as a CLI flag rather than accepting CIRIS_DIAGNOSTICS. So the requirement is operator control of probe targets; the mechanism is yours. A CLI flag or a config:* CEG object both satisfy it without breaking the contract, and if env vars are genuinely the right answer for this class of setting, that is a deliberate change to §2 rather than something to slip in.
Worth noting the general principle too: a status page that probes its own service through a CDN is measuring the CDN, and from the origin host it is measuring a hairpin no real user ever traverses. Whatever the mechanism, targeting the service is the correct measurement.
Bridge-side note
Adding the network required recreating the proxy container. Caddy's active health checker marked the upstream down for ~30 s during that window and returned no upstreams available (503) to a handful of requests before logging host is up. Public path is back to 200. Flagging it because it was a real, if brief, user-visible cost of creating this route.
us.proxyis flapping operational → degraded → outage on the public status page while the proxy answers in 3.5 ms. The proxy is not the problem; the probe route is.What is actually being measured
The status node probes the US proxy at
proxy1.ciris-services-1.ai. That record is Cloudflare-proxied:So a probe originating on the origin host leaves the machine, reaches Cloudflare's edge, and comes back to the same machine. Measured from that host, three ways:
Any threshold inside that 136–1502 ms band will flap, and the recorded transitions show exactly that —
operational -> outageat 15:04:21,outage -> degraded62 seconds later, withupstream_status: operationalthroughout and every LLM provider sourced fromcirisproxy.usreading healthy in the same payload.Confirmed from the status container's live sockets that this is the route in use — connections to
172.67.163.135:443and104.21.42.160:443, i.e. CF edge, not an origin or a docker name.EU is only healthy by accident
proxy1.ciris-services-2.aiis grey-clouded — it resolves straight to46.224.81.217. So the EU probe already does the right thing: it reaches the service. US and EU latency figures are therefore not comparable and never have been; one includes a CDN round trip and the other does not.A direct route now exists
Bridge-side,
ciris-proxyhas been added to thelens_cirislensnetwork, whichciris-statusis already on. Verified after the change, from inside the status node's network namespace:So
http://ciris-proxy:4000/health/livelinessis available and is ~400x faster and far more stable than the current target. This did not change what the adapter probes — that is the ask here.The ask
Make the probe target operator-configurable, so the deployment can point a component at the service rather than at whatever public route happens to exist.
On mechanism, one thing worth stating plainly rather than assuming: the operator's preference is environment variables set by the bridge. That conflicts with DEPLOY.md §2's zero-env contract, which we have been careful to keep — it is why
--diagnosticswas built as a CLI flag rather than acceptingCIRIS_DIAGNOSTICS. So the requirement is operator control of probe targets; the mechanism is yours. A CLI flag or aconfig:*CEG object both satisfy it without breaking the contract, and if env vars are genuinely the right answer for this class of setting, that is a deliberate change to §2 rather than something to slip in.Worth noting the general principle too: a status page that probes its own service through a CDN is measuring the CDN, and from the origin host it is measuring a hairpin no real user ever traverses. Whatever the mechanism, targeting the service is the correct measurement.
Bridge-side note
Adding the network required recreating the proxy container. Caddy's active health checker marked the upstream down for ~30 s during that window and returned
no upstreams available(503) to a handful of requests before logginghost is up. Public path is back to 200. Flagging it because it was a real, if brief, user-visible cost of creating this route.