Skip to content

Commit 7909fb5

Browse files
authored
refactor(compute): unify gateway restart reconciliation (#2743)
* refactor(compute): unify gateway restart reconciliation Remove the Docker-specific gateway shutdown cleanup and reconcile persisted running intent through ComputeDriver::StartSandbox for Docker, Podman, and VM drivers. Explicitly stopped sandboxes remain stopped. Refs #2417 Signed-off-by: Drew Newberry <anewberry@nvidia.com> * refactor(compute): stop local sandboxes on shutdown Signed-off-by: Drew Newberry <anewberry@nvidia.com> * test(e2e): match managed Podman containers Signed-off-by: Drew Newberry <anewberry@nvidia.com> * fix(compute): synchronize lifecycle sweeps Signed-off-by: Drew Newberry <anewberry@nvidia.com> --------- Signed-off-by: Drew Newberry <anewberry@nvidia.com>
1 parent 701382d commit 7909fb5

14 files changed

Lines changed: 896 additions & 316 deletions

File tree

‎.agents/skills/debug-openshell-cluster/SKILL.md‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -194,6 +194,13 @@ For source checkout development, restart the local gateway with:
194194
mise run gateway:docker
195195
```
196196

197+
During a graceful gateway restart, Docker, Podman, and VM sandboxes with
198+
running intent should stop before the gateway exits and restart after it
199+
returns. Check for `Stopped sandbox during gateway shutdown` and `Started
200+
sandbox during gateway startup` in gateway logs. A sandbox explicitly stopped
201+
through the CLI remains stopped. Kubernetes sandboxes are cluster-owned and do
202+
not follow this local gateway lifecycle.
203+
197204
### Step 5: Check Podman-Backed Gateways
198205

199206
```bash

‎architecture/compute-runtimes.md‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -110,6 +110,15 @@ compute to zero, and VM retains its launch request and writable overlay beside
110110
a stop marker. Delete remains a separate operation that removes these
111111
resources.
112112

113+
On graceful gateway shutdown, persisted running intent for Docker, Podman, and
114+
VM is stopped through the shared `StopSandbox` RPC before any gateway-managed
115+
driver process exits. The gateway does not persist `Stopped` for this
116+
infrastructure event. On startup, it reconciles the retained intent through the
117+
shared idempotent `StartSandbox` RPC before watch processing begins. Explicitly
118+
`Stopped` sandboxes are excluded from both sweeps. Kubernetes workloads are
119+
cluster-owned and continue running without gateway shutdown or startup
120+
lifecycle calls.
121+
113122
## Deletion Lifecycle
114123

115124
Lifecycle requests use per-sandbox gates to serialize stop, start, and

‎crates/openshell-driver-docker/README.md‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -26,6 +26,11 @@ policy. Start starts that same container, so files in the resolved OCI
2626
workspace remain available. A durably stopped sandbox is excluded from
2727
gateway startup recovery and stays stopped across gateway restarts. Delete
2828
continues to force-remove the container and clean up driver-owned material.
29+
Graceful gateway shutdown sends `StopSandbox` for each sandbox whose persisted
30+
phase requires running compute without changing that persisted intent. On
31+
startup, the gateway sends an idempotent `StartSandbox` request for the same
32+
sandboxes, restarting their retained containers. Explicitly stopped sandboxes
33+
remain excluded.
2934

3035
Before creating the container, the driver inspects the final sandbox image and
3136
captures its immutable image ID, raw OCI `Config.User`, and OCI

‎crates/openshell-driver-docker/src/lib.rs‎

Lines changed: 0 additions & 72 deletions
Original file line numberDiff line numberDiff line change
@@ -1055,69 +1055,6 @@ impl DockerComputeDriver {
10551055
}
10561056
}
10571057

1058-
pub async fn stop_managed_containers_on_shutdown(&self) -> Result<usize, Status> {
1059-
let containers = self.list_managed_container_summaries().await?;
1060-
let targets = containers
1061-
.into_iter()
1062-
.filter_map(|container| {
1063-
let state = container.state.unwrap_or(ContainerSummaryStateEnum::EMPTY);
1064-
if container_state_needs_shutdown_stop(state) {
1065-
summary_container_target(&container)
1066-
} else {
1067-
None
1068-
}
1069-
})
1070-
.collect::<Vec<_>>();
1071-
let target_count = targets.len();
1072-
let mut stopped = 0usize;
1073-
let mut failures = Vec::new();
1074-
let stop_timeout_secs = self.config.stop_timeout_secs;
1075-
1076-
let mut stop_results = futures::stream::iter(targets.into_iter().map(|target| {
1077-
let docker = self.docker.clone();
1078-
async move {
1079-
let result = docker
1080-
.stop_container(
1081-
&target,
1082-
Some(
1083-
StopContainerOptionsBuilder::default()
1084-
.t(docker_stop_timeout_secs(stop_timeout_secs))
1085-
.build(),
1086-
),
1087-
)
1088-
.await;
1089-
(target, result)
1090-
}
1091-
}))
1092-
.buffer_unordered(16);
1093-
1094-
while let Some((target, result)) = stop_results.next().await {
1095-
match result {
1096-
Ok(()) => {
1097-
stopped += 1;
1098-
}
1099-
Err(err) if is_not_found_error(&err) || is_not_modified_error(&err) => {}
1100-
Err(err) => {
1101-
warn!(
1102-
container = %target,
1103-
error = %err,
1104-
"Failed to stop Docker sandbox container during shutdown"
1105-
);
1106-
failures.push(target);
1107-
}
1108-
}
1109-
}
1110-
1111-
if !failures.is_empty() {
1112-
return Err(Status::internal(format!(
1113-
"failed to stop {} of {target_count} Docker sandbox containers during shutdown",
1114-
failures.len()
1115-
)));
1116-
}
1117-
1118-
Ok(stopped)
1119-
}
1120-
11211058
async fn reserve_pending_sandbox(&self, sandbox: &DriverSandbox) -> Result<(), Status> {
11221059
let mut pending = self.pending.lock().await;
11231060
if pending
@@ -3200,15 +3137,6 @@ fn summary_container_target(summary: &ContainerSummary) -> Option<String> {
32003137
.or_else(|| summary_container_name(summary))
32013138
}
32023139

3203-
fn container_state_needs_shutdown_stop(state: ContainerSummaryStateEnum) -> bool {
3204-
matches!(
3205-
state,
3206-
ContainerSummaryStateEnum::RUNNING
3207-
| ContainerSummaryStateEnum::RESTARTING
3208-
| ContainerSummaryStateEnum::PAUSED
3209-
)
3210-
}
3211-
32123140
/// States from which a managed container can be brought back to running by
32133141
/// `start_container`. Skip `Restarting` (already coming up), `Removing`,
32143142
/// `Dead` (terminal), `Paused` (needs `unpause`, not `start`), and

‎crates/openshell-driver-podman/README.md‎

Lines changed: 6 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -32,6 +32,12 @@ The stop call waits until Podman reports the container as stopped or exited.
3232
This keeps an immediate start from racing a rootless Podman stop that is still
3333
finishing after its API request returns.
3434

35+
Graceful gateway shutdown sends `StopSandbox` for each sandbox whose persisted
36+
phase requires running compute without changing that persisted intent. On
37+
startup, the gateway sends an idempotent `StartSandbox` request for the same
38+
sandboxes, restarting their retained containers. Explicitly stopped sandboxes
39+
remain excluded.
40+
3541
## Architecture
3642

3743
The Podman driver communicates with the Podman daemon over a Unix socket and

‎crates/openshell-driver-vm/README.md‎

Lines changed: 5 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -194,10 +194,11 @@ during the first prepare.
194194
195195
The driver also writes the accepted `DriverSandbox` launch request to
196196
`<state-dir>/sandboxes/<id>/sandbox.pb`. If the gateway restarts, it starts a
197-
new VM driver process; that process scans the sandbox state directories,
198-
restarts each persisted VM launcher, and preserves any existing `overlay.ext4`
199-
instead of cloning a fresh overlay template. If a restart happened before the
200-
overlay was created, the driver creates it during the start attempt.
197+
new VM driver process. During graceful shutdown, the gateway first sends the
198+
shared `StopSandbox` request for each persisted running-intent sandbox, which
199+
stops its launcher while retaining the launch request and `overlay.ext4`.
200+
After driver initialization, the gateway sends the idempotent `StartSandbox`
201+
request for that retained intent. Explicitly stopped sandboxes remain excluded.
201202
202203
Stop writes a marker in the sandbox state directory before terminating
203204
the launcher and releasing host GPU and network allocations. It retains

0 commit comments

Comments
 (0)