Skip to content

feat: Issue wildcard certificates through the certificate service - #526

Draft
scotwells wants to merge 7 commits into
mainfrom
feat/certificate-service-consumer
Draft

scotwells wants to merge 7 commits into
mainfrom
feat/certificate-service-consumer

Conversation

@scotwells

@scotwells scotwells commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Customers who want a wildcard hostname cannot get a certificate for it today, because cert-manager on the edge validates over HTTP and a wildcard can only be validated over DNS.

With the certificate service setting on, only wildcard hostnames get their certificate from the Milo certificate service, over DNS-01, and exact hostnames stay on cert-manager exactly as they are now.

A wildcard keeps serving its current certificate while a replacement is pending or failing, and any failed, stalled or overdue renewal shows on the hostname, raises a listener condition, counts in a metric and, after two hours, fires an alert.

Moving exact hostnames to the service is deferred until the wildcard path is proven, and that earlier migration work is kept on a backup branch.

API

Operator config, everything off by default:

gateway:
  certificateService:
    enabled: true
    secretNamespace: certificates-system
    verifyChain: true
    trustedRootsFile: /etc/certificate-service/roots.pem
  • enabled is refused when hostname verification is disabled.
  • verifyChain refuses issued material whose chain does not build to trustedRootsFile, or to the system roots when that is empty.

Request the operator makes for each wildcard listener:

apiVersion: certificates.miloapis.com/v1alpha1
kind: TLSCertificate
spec:
  dnsNames:
    - "*.shop.example.com"
  issuance: DNS01
  secretName: shop-https-hostname-0

Listener conditions on the Gateway:

conditions:
  - type: CertificateRenewalBlocked
    status: "True"
    reason: RenewalFailing
    message: "The TLS certificate for *.shop.example.com has not been issued: DNS problem: NXDOMAIN looking up TXT"
  - type: CertificateIssuanceBlocked
    status: "True"
    reason: IssuanceFailing
    message: "A TLS certificate cannot be issued for *.shop.example.com: names under datum.net are denied"

Hostname status on the HTTPProxy:

hostnameStatuses:
  - hostname: "*.shop.example.com"
    conditions:
      - type: CertificateReady
        status: "False"
        reason: ChallengeInProgress
        message: "We're provisioning and applying a certificate to this hostname - it may take a few minutes. Publish these DNS records to continue: _acme-challenge.shop.example.com CNAME 3f9c1a7b2e4d8a60.acme-dns.staging.env.datum.net"

CertificateReady reasons for a wildcard:

- status: "True"
  reason: CertificateIssued      # serving on the edge
- status: "False"
  reason: Pending                # requested, or issued and being applied
- status: "False"
  reason: ChallengeInProgress    # waiting on DNS, records in the message
- status: "False"
  reason: ProvisioningFailed     # refused, rejected, or failing with nothing serving
- status: "True"
  reason: RenewalFailing         # serving, but the replacement is blocked

Metrics and alert:

metrics:
  - nso_certificate_service_failures_total{namespace,name,listener,reason}
  - nso_certificate_service_listener_failing{namespace,name,listener,reason}
reasons:
  - Rejected
  - Refused
  - IssuanceFailed
  - NotReady
  - RenewalOverdue
  - MaterialRefused
  - UntrustedChain
  - NamespaceRefused
  - StepFailed
  - NotOwned
alert:
  name: CertificateServiceIssuanceFailing
  severity: warning
  for: 2h

Test plan

  • Exact hostnames are never handed to the service, and turning the setting off and on again requests and replaces nothing
  • A failed order, a request stuck short of Ready, an overdue renewal and an untrusted chain each raise a condition and the metric, while one failing wildcard leaves others on the gateway untouched
  • Build, vet, lint, the full suite and the alert rule tests pass locally and in CI
  • On staging with the setting on, a wildcard issues over DNS and an exact hostname keeps its cert-manager certificate

Related to datum-cloud/enhancements#913

Custom hostname certificates can now come from the Milo certificate
service instead of cert-manager on the downstream hub. The path is off
by default; with gateway.certificateService.enabled the gateway
controller requests a TLSCertificate per claimed hostname in the project
control plane, answers its HTTP-01 challenges on the downstream gateway,
and mirrors the issued Secret to the edge under the name the listener
already references.

Key changes:
- Add gateway.certificateService {enabled, kubeconfigPath} with
  validation that hostname verification stays on, since the service
  issues for whatever name it is handed
- Vendor certificates.miloapis.com/v1alpha1 TLSCertificate types under
  internal/certificates until the service module is importable
- Extract HTTP-01 solver route creation so the Challenge path and the
  TLSCertificate path share it; serve only challenges whose dnsName is
  the hostname the listener claimed, and drop routes when the challenge
  leaves status
- Mirror the edge Secret from the service-side Secret named in
  status.serviceSecretRef, read with the operator's own credentials,
  never from the project copy
- Keep a valid cert-manager Certificate until its renewal time, then
  retire it and request a TLSCertificate, so turning the flag on issues
  nothing up front and never races a cert-manager renewal
- Derive listener health and HTTPProxy CertificateReady from
  TLSCertificate conditions, with the listener kept serving on any valid
  downstream Secret
- Watch TLSCertificates across project control planes when enabled

Flag off is unchanged: the existing certificate tests pass untouched.

Related to datum-cloud/enhancements#913
Review of the certificate service consumer found that a failure in the
new issuance step could stop a gateway's routes, DNS and status from
reconciling, that issued material was mirrored without being checked,
and that a flag turned back off would reissue every migrated hostname.

Key changes:
- Record a failed certificate step in the listener's status and retry
  it with backoff instead of ending the gateway reconcile; routes, DNS
  and status continue whatever the service does
- Refuse a service-side Secret reference outside the configured service
  namespace, and verify key pair, hostname coverage and validity before
  any material may replace what is serving
- Never delete a TLSCertificate this gateway does not control, and name
  TLSCertificates from a hash of gateway and listener so two listeners
  can never share one
- Scope HTTP-01 solver routes to the claimed hostname and skip tokens
  that are not ACME tokens without failing the step
- Hand a hostname to the service ahead of cert-manager's renewal time,
  never while cert-manager has a renewal in flight
- With the flag off, keep serving a Secret the service handed over or
  issued until it is due for renewal, then return the hostname to
  cert-manager; remove the gateway's own TLSCertificates and solver
  routes so the service stops renewing it
- Report a hostname as ready while its previous certificate still serves,
  and name the DNS records the service is waiting on

Related to datum-cloud/enhancements#913
A hostname can serve for weeks on its previous certificate while every
attempt to replace it fails; until now that only reached the log. The
customer and operators now hear about it before the expiry does, and the
hand-over to the service no longer waits forever on a stuck cert-manager
renewal.

Key changes:
- Report a blocked renewal on the listener as CertificateRenewalBlocked
  and on the HTTPProxy hostname as CertificateReady=True with reason
  RenewalFailing, whether the failure is a rejected request, a refused
  service-side Secret, or a failed call; count each in
  nso_certificate_service_failures_total by reason
- Hand over anyway when cert-manager's renewal has failed, has been in
  flight for over a day, or under a week of lifetime remains
- Repeat the last message instead of the failing calls while a gateway
  is in backoff, and forget the gateway when it is deleted
- Bound each read of the service cluster to ten seconds
- Drop the hand-over marker once cert-manager has the Secret back
- Pin the TLSCertificate naming scheme with golden values

Related to datum-cloud/enhancements#913
The blocked-renewal condition on a listener was never removed, so one
transient failure read as a permanent one on the listener and on the
HTTPProxy hostname. It now goes away on the next pass that does not
block, and the hostname status no longer claims an issued certificate
that never reached the edge.

Key changes:
- Remove CertificateRenewalBlocked from a listener whose certificate
  step no longer reports a problem
- Report a hostname whose issued certificate could not be taken, and
  which has nothing serving, as ProvisioningFailed rather than issued;
  an issued certificate not yet mirrored is Pending
- Keep one cached message per listener while a gateway is in backoff
- Count a service rejection once per transition and drop a deleted
  gateway's failure series

Related to datum-cloud/enhancements#913
A new hostname whose issued certificate could not be read was reported
as "being applied" for as long as the failure lasted, because the
gateway only flagged blocked renewals on listeners that already served.
It now flags blocked issuance on the others as well, and HTTPProxy reads
both.

Key changes:
- Add the CertificateIssuanceBlocked listener condition for a listener
  with nothing serving whose certificate step failed; HTTPProxy maps it
  to ProvisioningFailed
- Drop a removed listener's failure series and tracker entries while
  the gateway lives on
- Compare counter deltas in tests so repeated runs pass

Related to datum-cloud/enhancements#913
0xmc
0xmc previously approved these changes Oct 2, 2026

@0xmc 0xmc left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving to unblock you, but two comments. Ping me if you make changes you want reviewed.

Add alert as proposed.

All staging hostnames will gradually switch to untrusted certs, so perhaps limit scope for just wildcard certs for this initial deployment.

Only single-label wildcard hostnames are issued by the Milo certificate
service, always over DNS-01. Exact hostnames stay on cert-manager whether
the flag is on or off, so the service never takes over a certificate that
already serves. Migrating exact hostnames waits until the wildcard path is
proven; the full migration is kept on backup/certificate-service-full-
migration.

Removed, since only the exact-hostname migration needed them:
- the hand-over of cert-manager Certificates 48h before renewal, with its
  hold, stuck-renewal and expiry-floor timing
- rollback cleanup when the flag goes off, and the cert-manager path's
  check for a handed-over Secret before reissuing
- retire and delete of legacy Certificates on the service path
- the certificate-service Secret label used to track hand-overs
- HTTP-01 answering for TLSCertificates and the solver refactor that
  shared route code with the cert-manager solver

Key changes:
- surface every failing or stalled renewal, not only rejections: a failed
  ACME order, a request short of Ready past an hour, and a served
  certificate past its renewal point raise CertificateRenewalBlocked or
  CertificateIssuanceBlocked and count once in the failure metric
- add nso_certificate_service_listener_failing so a lasting failure stays
  visible after its single count
- refuse issued chains the configured roots do not trust, behind
  verifyChain and an optional trustedRootsFile
- back off per listener, so one failing hostname no longer delays the
  others on its gateway
- tell a hostname the API refused why, instead of promising to retry
- recheck wildcard listeners hourly so a quiet stall is noticed
Add CertificateServiceIssuanceFailing, a warning raised when a wildcard
listener's certificate has failed to issue or renew for two hours. It
reads the failing-listener gauge, so a failure counted once still alerts,
and the failure counter, so a step that keeps failing and retrying does
too.

Key changes:
- add the rule in its own group so existing alert state is untouched
- add promtool tests for lasting, retrying, cleared and young failures
- add a runbook section mapping each reason to where it failed
@scotwells scotwells changed the title feat: consume certificate service behind flag feat: Issue wildcard certificates through the certificate service Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants