Skip to content

feat: hot-reload skills in long-running agents - #2

Merged
wi-kshitij-nagvekar merged 5 commits into
mainfrom
feat/skill-hot-reload-on-sync
Sep 22, 2026
Merged

wi-kshitij-nagvekar merged 5 commits into
mainfrom
feat/skill-hot-reload-on-sync

Conversation

@wi-kshitij-nagvekar

Copy link
Copy Markdown
Collaborator

Summary

Makes long-running Agents pick up skill changes from Git without restarting the Pod.

OpenCode discovers skills once, when its server instance starts, and caches them. So editing a skill repository does not reach a running Agent even after the cloned files on disk are updated. Git contexts already had a sync option; skill sources had none, so the only way to apply a skill change was to bounce the Agent's Deployment.

This adds sync to skill sources and implements the re-scan in the existing git-sync sidecar. Because a git-sync sidecar is already created for every mount with sync.enabled, and skills already flow through the same gitMount pipeline as Git contexts, the plumbing is largely shared — the new part is the server re-scan.

How it works

  1. A skill source with sync.enabled gets a git-sync sidecar, exactly as synced Git contexts do.
  2. The sidecar polls the remote and updates the clone in place (existing behavior, including GitHub App token refresh).
  3. When the commit changes, it asks the OpenCode server to re-scan, so the change is visible without a restart.

The re-scan uses POST /instance/dispose, which rebuilds the server instance on the next request. Verified behavior: skills are cached at instance start, dispose makes newly added and edited skills visible, the server stays healthy, and sessions survive it.

skills:
- name: org-skills
  git:
    repository: https://github.com/my-org/standards-skills.git
    path: .claude/skills
    sync:
      enabled: true
      interval: 15m   # default for skills

Two things this gets right

A reload never aborts work in progress. Disposal during an active turn terminates it (the assistant message ends with MessageAbortedError). The sidecar therefore checks GET /session/status and only disposes when no session is busy; a busy agent defers to the next cycle. The file update still happens immediately — only the re-scan waits.

An owed reload survives a sidecar restart. My first cut tracked "reload pending" in memory. Live testing showed the failure: a sidecar restarting while a reload was deferred would see local HEAD == remote HEAD, conclude there was nothing to do, and leave the server stale until the repository changed again — potentially days. The hash the server last scanned is now persisted on the shared volume (world-writable, for random-UID/SCC environments), so a restarted sidecar still knows the server is behind.

API

  • sync added to GitSkillSource, reusing the existing GitSync type so the field shape matches Git contexts.
  • reload added to GitSync so Git contexts can opt into the same re-scan. Contexts keep today's behavior by default; skill sources always reload when syncing, since re-scanning is the point of syncing skills.
  • Rollout is rejected for skills via CEL. It would need the controller to compare remote refs, which is not implemented for authenticated repositories, and would otherwise be a silently ineffective option.
  • Skill sources default to 15m (vs 5m for contexts), because each applied change costs a reload.

CRD manifests and deepcopy regenerated (make update-scripts update-crds).

Measured cost

Measurement Result
POST /instance/dispose ~3ms
First turn after dispose vs. warm ~+3.5–4.5s
Idle gap without dispose ~+0.5–1s (so the cost is the dispose, not idleness)
Memory over 30 consecutive reloads oscillates, then settles below baseline — churn, not a leak

At a 15m interval the penalty is roughly 0.4% of one turn. Pre-warming after disposal was measured as ineffective. There is no lighter re-scan endpoint — disposal is the only mechanism the server exposes.

Scope

Reload applies to Agent servers only. Task pods are ephemeral, always start fresh, and never build git-sync sidecars.

Plugins are out of scope: plugin packages are installed into an emptyDir by plugin-init, so a reload cannot make a new plugin's node_modules appear. That needs a separate install mechanism.

Known limitation

A mount using names maps fixed subpaths, so edits to mounted skills hot-reload but a newly added skill directory is not mounted until the Pod restarts. Omitting names mounts the whole directory and makes additions dynamic. Documented in the API and docs rather than worked around, since fixing it needs a different mount strategy.

Tests

  • Idle disposes; busy refuses to dispose; disposal and status-endpoint failures propagate; malformed status payloads fail closed (treated as busy) so an API change cannot abort someone's work.
  • processSkills sync propagation, including the 15m default and implicit reload for skills.
  • Sidecar reload env gating, and server wiring that points the sidecar at the agent's own port (test uses a non-default port so a hardcoded URL would fail).
  • make lint (0 issues), make test (891 passed), make verify all pass.
  • Live end-to-end against a real opencode serve with the compiled git-sync binary: a pushed skill change reached the running server without a restart; a busy session caused deferral and left the in-flight turn untouched; the restart scenario above was reproduced and then verified fixed.

E2E (make e2e-setup) was not run; it needs the Kind environment.

Docs

  • docs/adr/0043-hot-reload-skills-in-agents.md — mechanism, measured cost, constraints, alternatives.
  • website/docs/features/skills.md, git-auto-sync.md, context-system.md — usage and field reference.

Long-running Agents cache discovered skills at instance start, so updating
the cloned files on disk is invisible to a running OpenCode server. Skills
could not be kept in sync at all: Git contexts had a sync option, skill
sources did not.

Add sync to GitSkillSource, reusing the existing GitSync type so the field
shape matches Git contexts. Skills default to a 15m interval rather than the
5m used for contexts, because each applied change triggers a server instance
reload and skill catalogs change less often.

Add reload to GitSync for contexts to opt into the same re-scan behavior.
Skill sources always reload when syncing, since re-scanning is the point.

Rollout is rejected for skills: it needs the controller to compare remote
refs, which is not implemented for authenticated repositories, and would
otherwise be a silently ineffective option.
Populate the sync fields for skill sources in processSkills, mirroring the
existing Git context handling, and mark skill mounts as needing a server
re-scan after an update.

Pass OPENCODE_RELOAD_URL to git-sync sidecars whose mount requests a reload,
pointing at the agent's own server over loopback. Only Agent servers get it:
Task pods are ephemeral and never build git-sync sidecars, so the reload
concept does not apply there.
When a reload is requested, the sidecar asks the server to re-scan its
configuration so updated skills take effect without a Pod restart.

Disposal releases the whole server instance, so it must not run during an
active turn: doing so aborts the user's work (verified: the assistant message
ends with MessageAbortedError). The sidecar therefore checks session activity
and defers the reload while any session is busy, retrying every cycle.

The "server is behind" condition is persisted on the shared volume as the
commit hash the server last scanned. In-memory tracking was not enough: a
restarted sidecar would see an unchanged remote, conclude there was nothing
to do, and leave the server stale until the repository changed again.
Unit tests for the reload path against a fake OpenCode server: idle disposes,
busy refuses to dispose, disposal and status failures propagate, and malformed
status payloads fail closed rather than being read as idle.

Cover processSkills sync propagation (including the 15m default and implicit
reload for skills), the sidecar reload env gating, and the server wiring that
points the sidecar at the agent's own port. The server test uses a non-default
port so a hardcoded URL would fail.
Document why file sync alone cannot update skills in a running Agent (OpenCode
caches discovery at instance start) and how to enable the reload.

Add ADR 0043 recording the mechanism, the measured reload cost, and the
constraints: the idle gate that prevents aborting an active turn, the persisted
state that survives a sidecar restart, the names-filtering limitation, and why
plugins are out of scope.
@wi-kshitij-nagvekar
wi-kshitij-nagvekar merged commit 4996fba into main Sep 22, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants