Repository navigation
refactor(compute): define driver shutdown and restart lifecycle contract #2417
Description
Activity
- addedarea:gatewayGateway server and control-plane workGateway server and control-plane work
on Jul 22, 2026 - added a parent issue
on Jul 22, 2026 This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.
- addedstate:staleInactive item at risk of automatic closure.Inactive item at risk of automatic closure.
on Aug 6, 2026 - addedstate:acceptedA maintainer decided OpenShell should pursue this issueA maintainer decided OpenShell should pursue this issueand removedstate:staleInactive item at risk of automatic closure.Inactive item at risk of automatic closure.
on Aug 14, 2026 🏗️ build-plan
Implementation Plan
Issue type:
refactor
Complexity: Medium
Confidence: High — clear pathSummary
Unify local compute lifecycle around the public
ComputeDriverRPC contract. On graceful gateway shutdown, preserve the persisted logical running intent but invokeStopSandboxfor each Docker, Podman, and VM sandbox that should be running. On gateway startup, invoke the idempotentStartSandboxRPC for that same intent set. A stacked follow-up removes the remaining server-side behavioral checks for concrete driver names by negotiating gateway-managed lifecycle and process-identity behavior as public driver capabilities.Scope
- Keep refactor(compute): unify gateway restart reconciliation #2743 focused on shared graceful-shutdown stop and startup reconciliation for Docker, Podman, and VM.
- Do not transition persisted running-intent rows to
Stoppedduring gateway shutdown; shutdown is an infrastructure lifecycle event, not an explicit user stop. - Keep Kubernetes compute cluster-owned and running across gateway shutdown.
- Extend
compute_driver.protoin refactor(compute): support external driver parity #2744 with additive, forward-compatible capability negotiation for gateway-managed lifecycle and unspecified process-identity handling. - Replace runtime behavior gates based on
docker,podman,kubernetes, orvmnames with negotiated capabilities or generic structural validation. - Allow
--drivers <canonical-name> --compute-driver-socket <path>to select a remote driver; the endpoint override takes precedence over built-in construction. - Do not change default packaging, spawn additional driver subprocesses, or remove in-process built-in constructors.
Implementation Steps
- In refactor(compute): unify gateway restart reconciliation #2743, add a best-effort shutdown sweep that calls the public
StopSandboxRPC for every persisted running-intent sandbox on Docker, Podman, and VM before terminating any gateway-managed driver process. - Preserve sandbox phases during the shutdown sweep so startup can distinguish running intent from an explicit user stop.
- Reconcile that persisted intent through
StartSandboxon startup before watch processing begins. - In refactor(compute): support external driver parity #2744, expose a gateway-managed lifecycle capability covering both graceful-shutdown stop and startup start, advertise it from Docker, Podman, and VM, and keep Kubernetes opted out.
- Keep process-identity default handling as a separate capability because it controls request normalization before a driver RPC is sent.
- Remove the gateway-listener driver-name allowlist while retaining callback authorization and existing address, multicast, wildcard, and port validation.
- Permit canonical built-in names with
--compute-driver-socket; without an endpoint, retain existing built-in construction.
Test Plan
- Unit-test shutdown
StopSandboxtargeting for running-intent phases, exclusion of stable stopped/deleting/error phases, phase preservation, failure continuation, and Docker/Podman/VM versus Kubernetes selection. - Unit-test startup reconciliation and process-identity behavior by capability, including a canonical name over a fake UDS driver and conservative behavior with no capability.
- Update Docker, Podman, and VM gateway-restart E2E coverage to assert running compute stops while the gateway is down, restarts when the gateway returns, retains state, and leaves explicitly stopped sandboxes stopped.
- Run focused crate tests,
mise run pre-commit,mise run test, andmise run ci, plus available local-runtime E2E lanes.
Risks & Open Questions
- Shutdown RPC failures must not prevent attempts to stop the remaining sandboxes or termination of a gateway-managed driver process; failures are surfaced after cleanup.
- A gateway crash or forced kill cannot perform graceful RPC cleanup. Startup reconciliation remains idempotent for the retained intent.
- Legacy external drivers advertise no gateway-managed lifecycle capability and remain operator-owned.
- The driver process and sandbox compute lifecycle remain distinct: the gateway uses sandbox RPCs first, then terminates only a process it owns.
Documentation Impact
- Update architecture, local-driver READMEs, and the compute-driver reference in refactor(compute): unify gateway restart reconciliation #2743 for the stop-on-shutdown/start-on-startup contract.
- Keep refactor(compute): support external driver parity #2744 free of public-facing documentation changes as requested; update internal driver/debug guidance for capability-based external-driver diagnostics.
- LSM compatibility is unchanged: this changes lifecycle orchestration, not process visibility or execution controls.
Revision 3 — graceful gateway shutdown stops Docker, Podman, and VM compute through
StopSandboxwhile preserving restart intentImplementation is ready in #2743. Graceful gateway shutdown now invokes the public
StopSandboxRPC for running-intent Docker, Podman, and VM sandboxes while preserving their persisted logical intent; gateway startup invokesStartSandboxto restart them. Explicitly stopped sandboxes remain stopped, and Kubernetes remains cluster-owned. Docker restart E2E passes with an assertion that the container is stopped while the gateway is down.Implementation is ready in stacked PR #2744, based on #2743. The follow-up replaces concrete driver-name lifecycle gates with the additive
GATEWAY_MANAGED_LIFECYCLEcapability, covering both graceful-shutdownStopSandboxand startupStartSandbox. Docker, Podman, VM, and an external UDS driver can share the same public contract; Kubernetes and legacy external drivers opt out by omitting the capability. It also retains separate process-identity negotiation and canonical-name UDS parity. Merge order: #2743, then #2744.- added 3 commits that reference this issue
on Aug 19, 2026 - added a commit that references this issue
on Aug 20, 2026 - added a commit that references this issue
on Aug 20, 2026 Closure evidence: #2743 unified graceful shutdown and startup reconciliation through the public
StopSandboxandStartSandboxRPCs. #2744 extended the same lifecycle contract to external drivers. #2786, #2822, and #2823 established generic driver registration, standalone first-party drivers, and explicit gateway-versus-operator lifecycle ownership.
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsDone
Problem Statement
The OpenShell Drivers roadmap (#1051) requires in-tree and out-of-tree compute drivers to interact with the gateway through the same gRPC contract, but driver shutdown and restart semantics are not defined by
compute_driver.proto.The gateway currently has different lifecycle ownership for in-process drivers, the gateway-managed VM driver process, and operator-managed external driver endpoints. Without a shared contract, graceful gateway shutdown, driver restart, sandbox recovery, and reconnection behavior remain driver-specific and cannot be covered consistently by conformance tests.
Proposed Design
Define a common driver lifecycle contract before adding implementation-specific hooks. The design should specify:
Keep process supervision separate from the protocol. Package managers or operators may own an external process even when the gateway uses a common lifecycle RPC.
Alternatives Considered
Agent Investigation
Definition of Done