Skip to content

RFC 0001: runtime/platform integrations needed to realize full potential #9

Description

@adiled

Runtime / platform integrations needed to realize RFC 0001's full potential

RFC 0001 (the node-unit-spec extensions on the feat/orchfile-rfc-0001 branch: time / space / trust axes) outgrows the current two-axis model (Runtime × Platform). The v1.0-rc grammar was served by "how to run" × "how to supervise." RFC 0001's three axes each land on a different gap, and full realization needs more than integrations — it needs a third axis plus several cross-cutting subsystems.

Below is the map, grounded in the current code (src/runtime/mod.rs, src/platform/mod.rs, src/orchdi.rs, docs/init/03-platform-interface.md).

1. Time axis → orchdi grows from babysitter to supervisor (the keystone)

orchdi.rs today is a dependency poller + executer: waits on DepSpec readiness, runs pre_start, spawns start, SIGTERM/SIGKILLs on stop. That's a babysitter, not a supervisor. RFC 0001's entire time axis is supervision semantics that launchd/systemd can't express, so orchdi is the platform that must carry it:

  • Probe engine — STARTUP / LIVENESS / READINESS as distinct pollers with distinct outcomes. The spec is explicit: readiness loss is not a failure and never triggers ON_FAILURE. Today orchdi only polls dependencies; it needs a three-tier probe loop (liveness fail → restart, readiness fail → withhold traffic only, startup gates the others).
  • Notify socket — READY notify = systemd's Type=notify pattern: service pushes readiness over a unix socket instead of being polled. orchdi needs its own notify socket + a READY event loop.
  • WATCHDOG timer — keep-alive window, a miss = hang = restart (deadlock/livelock detection). New timer in orchdi.
  • Managed lifecycle — LIFECYCLE managed exposes configure/activate/deactivate (Active→Inactive); a readiness failure deactivates, not restarts. That's a real state machine in orchdi.
  • ON_FAILURE — restart | state:<name> | unit:<service> escalation: orchdi must trigger machine-state transitions and start other units.

2. Space axis → wants Linux + containerd/podman; systemd is mostly already native

The space axis is node-local requirements (ARCH arm64, DEVICE /dev/ttyS1, REQUIRES_CAP hugepages,sriov) — Linux-only. macOS/launchd has no device nodes, no hugepages, no sriov. Full realization needs the Linux + containerd/podman runtime/platform pair the justfile scaffolds but doesn't ship yet.

systemd gets ~70-80% of RFC 0001 for free — the work is mapping new Service fields onto existing directives:

RFC 0001 systemd native
READY notify Type=notify
WATCHDOG WatchdogSec=
STARTUP ExecStartPre=
LIVENESS Restart= + probe
CONDITION path-exists ConditionPathExists=
ASSERT path-exists AssertPathExists= (hard-fail variant exists)
SLICE Slice= (cgroup)
INSTANCES + %i template units foo@.service + %i
CAPABILITY CapabilityBoundingSet=
READONLY_ROOT / NO_NEW_PRIVILEGES / PRIVATE_TMP / SECCOMP ReadOnlyDirectories= / NoNewPrivileges= / PrivateTmp= / SystemCallFilter=
MEMORY/CPUS/LIMIT_* already mapped

Space-axis gaps in systemd: the flat machine-state model (systemd has runlevels/targets, not a flat MACHINE_STATES set), WINDOW (recurring), and managed lifecycle — no init system does these.

3. Trust axis → runtime security hardening + external backends

Splits down the middle:

Runtime side (sandbox directives are container-runtime concerns): CAPABILITY, READONLY_ROOT, NO_NEW_PRIVILEGES, PRIVATE_TMP, SECCOMP, EPHEMERAL. Today orchd-osx honors user / nofile / nproc / cgroup caps / memory / cpus / publish / volumes — but not seccomp, not capability bounding sets, not readonly-root, not private-tmp. Integration: extend the apple-osx guest init + --spec blob to apply seccomp filters and capability sets in the guest (kernel is already pinned), remount root ro, private-tmp as tmpfs, and EPHEMERAL = volatile root / no persistent volume. The bare runtime can't seccomp on macOS (Linux/systemd concern). containerd/podman get it nearly free via OCI securityContext.

Backend side (neither runtime nor platform — new subsystems): SECRET needs a secrets resolver per scheme (vault://, tpm://, file://), resolved at tend/run, injected into env, never in orch parse output. IDENTITY spiffe://... scope=session needs a SPIRE/SPIFFE integration minting an SVID per service / per-session. ON_TAMPER needs TPM/tamper-signal handling; AUDIT needs a tamper-evident log. These are the "platform tooling" the spec deliberately defers to — external backends orchd wires, not grammar.

4. The honest architectural conclusion

Realizing RFC 0001 fully is not "add a cluster orchestrator as a runtime" — it's four things the current model doesn't have:

  1. A third axis: the machine-state controller. MACHINE_STATES/DEFAULT_STATE/STATE answer when services run — orthogonal to how-to-run (runtime) and how-to-supervise (platform). Needs a state manager above the platform: boot into DEFAULT_STATE, activate services whose STATE includes it; on ON_FAILURE state:<name> or WINDOW transition, deactivate/activate the set.
  2. The provider registry. CONDITION/ASSERT/WINDOW predicates (network-online, time-synced, power-state, tariff, contact-pass, reverse-DNS custom) are a third plugin axis at the semantic layer — same registry pattern as runtime/platform, but for admission predicates. systemd covers path-exists; the rest need provider implementations. Per the spec: "the provider set is a registry, not grammar."
  3. The change engine. UPDATE (ab|inplace, rollback-on-health-fail, signature-required) and ROLLOUT (rolling/blue-green/canary across instances, gated on readiness) are a higher-order orchestration over sow/plant/tend — a new walk staging template instances through health-gated rollouts. Not a runtime, not a platform: a change-management layer.
  4. Profile expander + observability. PROFILE/ASSURANCE/LABEL expand at graft/compose time (orch-side, not runtime/platform): a profile-document registry that expands a PROFILE into concrete directives, pinned in output. METRICS/TRACES/LOG_FORMAT need a scrape server + OTLP exporter + log-format contract at the platform layer.

Integrations, ranked by what actually unlocks the RFC

  1. orchdi → full supervisor (probe engine, notify socket, watchdog, managed lifecycle, ON_FAILURE) — unlocks the time axis, the biggest share.
  2. Linux + containerd/podman pair — unlocks the space axis (device nodes, cgroups, capabilities) macOS/launchd structurally can't.
  3. apple-osx guest security — seccomp / caps / sandbox / ephemeral in the VZ guest init.
  4. Machine-state controller — the missing third axis, on top of the platform.
  5. Provider registry, secrets resolver, SPIFFE identity, audit — the semantic/trust backends.
  6. Change engine + observability + profile expander — the higher-order and graft-time layers.

Deliberately not needed

The RFC scopes out PLACEMENT/AFFINITY/ANTI_AFFINITY/REDUNDANCY as "a separate fleet-composition layer." The space axis is node-local admission, not fleet scheduling. Realizing RFC 0001 means building a better node — not wrapping a cluster. Fleet orchestrators belong as a supervised service in the Orchfile, not as a runtime.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions