Skip to content

refactor(compute): define driver shutdown and restart lifecycle contract #2417

Description

@elezar

Problem Statement

The OpenShell Drivers roadmap (#1051) requires in-tree and out-of-tree compute drivers to interact with the gateway through the same gRPC contract, but driver shutdown and restart semantics are not defined by compute_driver.proto.

The gateway currently has different lifecycle ownership for in-process drivers, the gateway-managed VM driver process, and operator-managed external driver endpoints. Without a shared contract, graceful gateway shutdown, driver restart, sandbox recovery, and reconnection behavior remain driver-specific and cannot be covered consistently by conformance tests.

Proposed Design

Define a common driver lifecycle contract before adding implementation-specific hooks. The design should specify:

  • which component owns starting, stopping, and restarting each driver process;
  • whether the protocol needs an explicit graceful-shutdown RPC, a drain or quiesce operation, capability negotiation, or a combination;
  • what a successful shutdown response guarantees for running sandboxes and watch streams;
  • how the gateway detects driver loss and reconnects after a driver restart;
  • how persisted sandboxes are reconciled after gateway or driver restart;
  • how unmanaged out-of-tree drivers express unsupported lifecycle operations;
  • timeout, idempotency, and error semantics; and
  • which behaviors belong in the driver conformance suite (spike: driver conformance suite for compute drivers #2183).

Keep process supervision separate from the protocol. Package managers or operators may own an external process even when the gateway uses a common lifecycle RPC.

Alternatives Considered

  • Continue relying on process termination and reconnect behavior only. This leaves graceful drain and recovery semantics implicit and driver-specific.
  • Make the gateway supervise every driver process. This does not fit operator-managed out-of-tree endpoints.
  • Add a shutdown RPC without first defining its guarantees. This risks encoding incompatible behavior across drivers.

Agent Investigation

Definition of Done

  • Document lifecycle ownership for in-process, gateway-managed, and operator-managed drivers.
  • Decide whether new protocol methods or capabilities are required.
  • Define shutdown, disconnect, restart, reconnect, and reconciliation semantics.
  • Identify migration work for all in-tree drivers and external drivers.
  • Add or update implementation issues and conformance scenarios.
  • Update the relevant architecture documentation.

Activity

  1. github-actions commented on Aug 6, 2026

    @github-actions

    This issue has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. Comment or remove the state:stale label to keep it open.

  2. added
    state:acceptedA maintainer decided OpenShell should pursue this issue
    and removed
    state:staleInactive item at risk of automatic closure.
    on Aug 14, 2026
  3. drew commented on Aug 14, 2026

    @drew
    Collaborator

    🏗️ build-plan

    Implementation Plan

    Issue type: refactor
    Complexity: Medium
    Confidence: High — clear path

    Summary

    Unify local compute lifecycle around the public ComputeDriver RPC contract. On graceful gateway shutdown, preserve the persisted logical running intent but invoke StopSandbox for each Docker, Podman, and VM sandbox that should be running. On gateway startup, invoke the idempotent StartSandbox RPC for that same intent set. A stacked follow-up removes the remaining server-side behavioral checks for concrete driver names by negotiating gateway-managed lifecycle and process-identity behavior as public driver capabilities.

    Scope

    • Keep refactor(compute): unify gateway restart reconciliation #2743 focused on shared graceful-shutdown stop and startup reconciliation for Docker, Podman, and VM.
    • Do not transition persisted running-intent rows to Stopped during gateway shutdown; shutdown is an infrastructure lifecycle event, not an explicit user stop.
    • Keep Kubernetes compute cluster-owned and running across gateway shutdown.
    • Extend compute_driver.proto in refactor(compute): support external driver parity #2744 with additive, forward-compatible capability negotiation for gateway-managed lifecycle and unspecified process-identity handling.
    • Replace runtime behavior gates based on docker, podman, kubernetes, or vm names with negotiated capabilities or generic structural validation.
    • Allow --drivers <canonical-name> --compute-driver-socket <path> to select a remote driver; the endpoint override takes precedence over built-in construction.
    • Do not change default packaging, spawn additional driver subprocesses, or remove in-process built-in constructors.

    Implementation Steps

    1. In refactor(compute): unify gateway restart reconciliation #2743, add a best-effort shutdown sweep that calls the public StopSandbox RPC for every persisted running-intent sandbox on Docker, Podman, and VM before terminating any gateway-managed driver process.
    2. Preserve sandbox phases during the shutdown sweep so startup can distinguish running intent from an explicit user stop.
    3. Reconcile that persisted intent through StartSandbox on startup before watch processing begins.
    4. In refactor(compute): support external driver parity #2744, expose a gateway-managed lifecycle capability covering both graceful-shutdown stop and startup start, advertise it from Docker, Podman, and VM, and keep Kubernetes opted out.
    5. Keep process-identity default handling as a separate capability because it controls request normalization before a driver RPC is sent.
    6. Remove the gateway-listener driver-name allowlist while retaining callback authorization and existing address, multicast, wildcard, and port validation.
    7. Permit canonical built-in names with --compute-driver-socket; without an endpoint, retain existing built-in construction.

    Test Plan

    • Unit-test shutdown StopSandbox targeting for running-intent phases, exclusion of stable stopped/deleting/error phases, phase preservation, failure continuation, and Docker/Podman/VM versus Kubernetes selection.
    • Unit-test startup reconciliation and process-identity behavior by capability, including a canonical name over a fake UDS driver and conservative behavior with no capability.
    • Update Docker, Podman, and VM gateway-restart E2E coverage to assert running compute stops while the gateway is down, restarts when the gateway returns, retains state, and leaves explicitly stopped sandboxes stopped.
    • Run focused crate tests, mise run pre-commit, mise run test, and mise run ci, plus available local-runtime E2E lanes.

    Risks & Open Questions

    • Shutdown RPC failures must not prevent attempts to stop the remaining sandboxes or termination of a gateway-managed driver process; failures are surfaced after cleanup.
    • A gateway crash or forced kill cannot perform graceful RPC cleanup. Startup reconciliation remains idempotent for the retained intent.
    • Legacy external drivers advertise no gateway-managed lifecycle capability and remain operator-owned.
    • The driver process and sandbox compute lifecycle remain distinct: the gateway uses sandbox RPCs first, then terminates only a process it owns.

    Documentation Impact


    Revision 3 — graceful gateway shutdown stops Docker, Podman, and VM compute through StopSandbox while preserving restart intent

  4. moved this from Todo to In progress in OpenShell Roadmapon Aug 14, 2026
  5. drew commented on Aug 14, 2026

    @drew
    Collaborator

    Implementation is ready in #2743. Graceful gateway shutdown now invokes the public StopSandbox RPC for running-intent Docker, Podman, and VM sandboxes while preserving their persisted logical intent; gateway startup invokes StartSandbox to restart them. Explicitly stopped sandboxes remain stopped, and Kubernetes remains cluster-owned. Docker restart E2E passes with an assertion that the container is stopped while the gateway is down.

  6. drew commented on Aug 14, 2026

    @drew
    Collaborator

    Implementation is ready in stacked PR #2744, based on #2743. The follow-up replaces concrete driver-name lifecycle gates with the additive GATEWAY_MANAGED_LIFECYCLE capability, covering both graceful-shutdown StopSandbox and startup StartSandbox. Docker, Podman, VM, and an external UDS driver can share the same public contract; Kubernetes and legacy external drivers opt out by omitting the capability. It also retains separate process-identity negotiation and canonical-name UDS parity. Merge order: #2743, then #2744.

  7. added 3 commits that reference this issue on Aug 19, 2026
    6702450
    b341901
    90efe6d
  8. moved this from In progress to Done in OpenShell Roadmapon Sep 19, 2026
  9. drew commented on Sep 19, 2026

    @drew
    Collaborator

    Closure evidence: #2743 unified graceful shutdown and startup reconciliation through the public StopSandbox and StartSandbox RPCs. #2744 extended the same lifecycle contract to external drivers. #2786, #2822, and #2823 established generic driver registration, standalone first-party drivers, and explicit gateway-versus-operator lifecycle ownership.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:computearea:gatewayGateway server and control-plane workstate:acceptedA maintainer decided OpenShell should pursue this issue

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions