Repository navigation
feat: fail-closed capability negotiation for gateway↔supervisor signals under version skew #2949
Description
Activity
- addedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Aug 26, 2026 📋 triage-agent
Triage Assessment
Classification: validated-feature
Summary
The reported gap is real and confirmed against
main. The gateway carries per-request behavior signals to the supervisor as SSH env requests, and the supervisor acknowledges every env request withchannel_successregardless of whether it understands the variable. A newer gateway paired with an older, baked-in supervisor therefore cannot distinguish an honored signal from a silently dropped one. High confidence in the evidence.Investigation
Verified on
main:env_requestreplieschannel_successunconditionally for any variable (crates/openshell-supervisor-process/src/ssh.rs:854). This is the exact silent-ack mechanism the report describes and rules out flippingwant_replyon the env request as a detection path.- The
OPENSHELL_MAIN_READ_ONLYsignal already rides this path onmain(ssh.rs:848), so the skew surface spans at least the main-attach pipeline today, independent of the exec case. - The second cited signal (
no_login_shell) is not yet onmain; it originates from the still-open PR feat(sandbox): add --no-login-shell to skip shell startup files on exec #2852 (gator:in-review). The issue's line numbers reference that PR branch, so exec-site line refs (sandbox.rs:2170/:2359) and theno_login_shellhandling will not resolve againstmainuntil feat(sandbox): add --no-login-shell to skip shell startup files on exec #2852 lands. This does not affect validity — the underlying transport mechanism and themain_read_onlysignal are already present. - Duplicate scan (all states): no existing issue proposes capability negotiation for this SSH-signal pipeline. Push gateway-owned desired state to supervisors #1731 (
state:review-ready) explicitly excludes negotiation and is scoped to the config-push pipeline; feat: support in-place supervisor upgrade via re-exec #2589 (re-exec upgrade) is complementary. The report's Relationship and Alternatives analysis is accurate. - Feasibility of the two design directions is plausible: the
SupervisorHello/SessionAcceptedcapability-set approach is additive and composes across pipelines; the SSH identification-banner alternative is feasible in russh 0.62 as described, with the noted plumbing caveat thatremote_sshid()lives on the clientSessionand must be captured in aHandlercallback.
Report note
Filed with
topic:securityby the reporter as a security-relevant opt-out gate, not a vulnerability disclosure (no live exploit) — consistent withSECURITY.mdhandling as a feature. Treated as a feature request here, not routed as a security report.Impact Signals
- Affected users/scope: Operators who upgrade the gateway in place while long-lived sandboxes created before the upgrade remain running. Occurs within a single normal deploy, not only across releases. Live surface today is the
main_read_onlysignal; grows tono_login_shellwhen feat(sandbox): add --no-login-shell to skip shell startup files on exec #2852 lands, and to any future skew-sensitive signal. - Regression: No — pre-existing architectural gap, not a recent behavior change.
- Workaround: Unavailable. The gateway cannot tell an honored signal from a dropped one; recreating sandboxes after an upgrade would avoid skew but there is no signal indicating it is necessary.
- Evidence quality: High — the central silent-ack mechanism was verified directly in
mainsource.
Human Decision Required
Decide whether OpenShell should address this issue. If yes, apply
state:accepted, associate it with a roadmap item, or do both, and decide
whether the work remains human-owned. Either action records acceptance;
roadmap placement additionally records sequencing. Sequencing note: the
design's preferred direction (Alternative 4) couples to #1731's bootstrap;
the standaloneSupervisorHello(Alternative 1) does not.
To queue investigation or planning for an unattended agent, also apply
agent:plan-requested. You can instead directly ask an agent to use
create-spikeorbuild-from-issueon this issue; the agent will warn about
missing expected workflow labels and continue without changing them. If no,
close it as not planned and record the rationale.I think there may be another option worth considering alongside Alternatives 1/4: negotiate one protocol-baseline capability per session, then let each SSH control request fail closed.
The shape would be:
- A supervisor that implements strict signal handling advertises a baseline such as
strict-openshell-env-v1inSupervisorHello(or the Push gateway-owned desired state to supervisors #1731 bootstrap). - If that baseline is absent, the gateway treats any skew-sensitive behavior as unsupported. This handles existing supervisors, whose unconditional ACKs cannot be trusted.
- If the baseline is present, the gateway sends reserved
OPENSHELL_*signals withwant_reply=trueand waits for the corresponding channel result before continuing with exec/subsystem/etc. - The supervisor ACKs only after a single handler has recognized, validated, and applied the signal. Unknown or invalid reserved signals receive
channel_failure. Non-OpenShell env requests retain the current compatibility behavior.
I would make the handler itself the source of truth rather than maintaining a separate allowlist, so adding support and deciding whether to ACK cannot drift:
- recognized + valid + applied → ACK
- unknown/invalid
OPENSHELL_*→ reject - ordinary SSH env request → existing behavior
This would avoid a growing per-feature capability set: after the strict-reply baseline is established, the request/reply becomes the feature check for future signals. It also catches a gateway that accidentally omits a preflight gate.
The tradeoff is that this differs from the issue's current preferred design and acceptance criterion: it adds one SSH request/reply round trip for every signaled operation. Also, in russh,
set_env(true).awaitonly sends the request; the gateway must consume the subsequentChannelMsg::SuccessorFailurebefore sending the operation request.So I see this as a choice between:
- per-feature capabilities negotiated once, with no per-command round trip; or
- one strict-protocol baseline plus per-signal fail-closed replies, with a smaller capability registry but an added round trip.
It boils down to would the latter tradeoff be worth considering, or is the zero steady-state round-trip requirement decisive here.
- A supervisor that implements strict signal handling advertises a baseline such as
hey @pimlock, since #2967 is adding the
SessionAcceptedbootstrap, that exchange now looks like a good fit for a small capability set (Alternative 4), rather than a separateSupervisorHello.The main consideration is timing. If that's the direction, a field would need reserving before the bootstrap shape settles, so raising it now.
@varshaprasad96's strict-baseline idea would also slot into the same field, the open question there is whether the per-signal round-trip is acceptable or the zero-round-trip requirement is decisive.
I don't have a strong preference on how it's carried, the key outcome either way is that an unsupported signal fails loud (
FailedPrecondition) instead of silently no-opping. Does folding it intoSessionAcceptedseem reasonable?IMO, using #2967’s initialization exchange seems like the right lifecycle boundary. One nuance is direction (we can discuss more): SessionAccempted is Gateway -> supervisor, while these capabilities describe what the supervisor implements. I'd put supported_capabilities in
ConfigBootstrapResult(orSupervisorHello?) and cache it on the live session before making bootstrap complete and enabling relays.Reacted by Artem Lytvyn
References: originates from PR #2852 (
--no-login-shell, which Closes #2668) and review finding GATOR-7d79a247-01 on that PR.Related issues: #1731 (push gateway-owned desired state — deliberately excludes capability negotiation for the config-push pipeline; see Relationship below), #2589 (in-place supervisor upgrade via re-exec — complementary root-cause fix; see Alternatives).
Suggested labels:
area:gateway,area:supervisor,topic:security(this gates a security-relevant opt-out; not a vulnerability disclosure — no live exploit — so it is filed as a feature per SECURITY.md).User Story
As an operator running long-lived sandboxes who upgrades the gateway in place, I want the gateway to reject a requested sandbox behavior when the sandbox's supervisor doesn't support it, so that a security- or correctness-relevant option never silently degrades to the old behavior without any signal.
Problem Statement
The supervisor is baked into each sandbox at creation and never updates for that sandbox's lifetime, while the gateway is upgraded independently. When a newer gateway sends a behavior signal to an older supervisor that predates it, the supervisor accepts the request but ignores the signal — there is no mechanism for the gateway to learn which behaviors the live supervisor actually understands. Today these signals are carried as SSH environment requests, which an older supervisor acknowledges with success even when it does not act on them.
Impact / Why This Matters
Without this, a gateway-issued option can silently no-op against an already-running sandbox, and the caller receives a normal success as if it had taken effect.
--no-login-shell(PR feat(sandbox): add --no-login-shell to skip shell startup files on exec #2852, addressing sandbox exec always runs commands through a login shell, so sandbox-user startup files run before the requested command #2668). Against a pre-change supervisor, the command still runs through a login shell (bash -lc), the user's profile files are sourced, and the caller sees a clean success. For a security opt-out on a trusted exec boundary, silent failure is worse than an explicit error.no_login_shell(exec) andmain_read_only(main attachment) — so the skew spans more than one pipeline (exec, interactive shell/PTY, direct-tcpip tunnels, subsystems, main attach all share the same gateway↔supervisor SSH connection).Current workaround: none. The gateway cannot tell an honored signal from a silently dropped one.
Proposed Design
From the operator's and caller's perspective:
FailedPrecondition: "supervisor predates this capability; recreate the sandbox") instead of silently running the old behavior.Internal wire format is left to implementers; directions are in Alternatives.
Relationship to #1731 (why this is not the rejected negotiation)
#1731 ("Push gateway-owned desired state to supervisors") resolved to not add capability negotiation, a protocol-version field, or a rollout mode, using a required bootstrap exchange plus bounded timeout as its compatibility gate. That decision is scoped to the config/policy/provider/inference push pipeline over the reverse
ConnectSupervisorsession, where the gateway controls both ends of a single required exchange and can fail the whole init if the bootstrap does not complete.This issue is a different pipeline: per-request behavior signals carried as SSH env requests (exec, interactive/PTY, tunnels, subsystems, main attach). Two properties make #1731's gate insufficient here:
channel_successunconditionally (crates/openshell-supervisor-process/src/ssh.rs:726) — so there is no bootstrap-style all-or-nothing gate to hang the result on.So the fail-closed check must be consulted at each affected call site, keyed on something learned once per session. This does not reintroduce config-push negotiation; it reuses #1731's spirit (learn compatibility once, fail clearly) on the SSH transport. If maintainers prefer to fold this into #1731's bootstrap result (advertise SSH-behavior capabilities in
SessionAcceptedrather than a newSupervisorHello), that satisfies this issue's acceptance criteria — see Alternative 4.Acceptance Criteria
FailedPrecondition-style error rather than silently falling back.no_login_shellfails closed: with a profile marker seeded, requesting a non-login exec against a pre-change supervisor returns an unsupported-feature error and does not source the profile.Alternatives Considered
SupervisorHellocapability set (preferred direction). Supervisor advertises a set of supported capability strings once per session; gateway caches it on the live session and gates each behavior against it. Additive (new behavior = new string), composes across pipelines, no version-ordering tables. Cost: a new negotiation message/contract to define and maintain, and one exchange at session open.SSH-2.0-<software>), already exchanged at connection start; gateway parses the remote id and gates on a version threshold. No new message, no proto field, no extra round-trip. Narrower: carries a version, not a capability list, so each future behavior needs its own "added in version ≥ X" mapping on the gateway side. (Verified feasible in russh 0.62: supervisor can setserver::Config.server_id; client exposesSession::remote_sshid()— though it lives on the clientSession, not theHandle, so the gateway must capture it in aHandlercallback and thread it out to the exec path.)SessionAcceptedbootstrap. Rather than a newSupervisorHello, the supervisor advertises its SSH-behavior capability set as one field of the existing bootstrap result Push gateway-owned desired state to supervisors #1731 already requires. The gateway caches it on the live session and gates behaviors identically. Reuses an exchange that must exist anyway (no new message, no extra round-trip) and stays consistent with Push gateway-owned desired state to supervisors #1731's "learn compatibility once" gate. Cost: couples this fail-closed path to Push gateway-owned desired state to supervisors #1731 landing first, and mixes SSH-transport concerns into a config-push message. Viable if Push gateway-owned desired state to supervisors #1731 ships before or with this work.Direction (1)/(4) is preferred because the skew already spans multiple pipelines and is expected to grow; a session-scoped capability set turns "re-solve skew per feature" into "add one string," at zero steady-state cost. Choice between a standalone
SupervisorHello(1) and folding into #1731's bootstrap (4) depends on sequencing of #1731.Agent Investigation
crates/openshell-server/src/grpc/sandbox.rs:2170(interactive) and:2359(unary) — both viaset_env(false, NO_LOGIN_SHELL_ENV.0, ...)(want_reply=false). These are the sites gator flagged.crates/openshell-supervisor-process/src/ssh.rs:env_request(fn at:701) —no_login_shellvariable check at:715/set at:718,main_read_onlyat:720; the flag maps to the shell argument inlogin_shell_flagat:1058(-cvs-lc), covered by testlogin_shell_flag_controls_profile_sourcingat:1778.env_requestunconditionally replieschannel_successat:726for any variable, so an older supervisor acks an unknown signal with success — there is no reply-based signal to detect a dropped capability (this rules out flippingwant_replyon the env request).OPENSHELL_NO_LOGIN_SHELL(exec) andOPENSHELL_MAIN_READ_ONLY(main attachment), confirming the pattern spans pipelines.server::Config.server_id(set) andSession::remote_sshid()(read, client side), with the plumbing caveat noted above.state:accepted; origin PR feat(sandbox): add --no-login-shell to skip shell startup files on exec #2852 is still open.