Skip to content

methods decorated with compute and launched in separate process can't be interrupted #1

Description

@krafczyk

Encountered when running dryml code in a jupyter notebook.

I decorated a training method with @dryml.compute and realized I had included debug print statements in a forward method of a pytorch model (which is called very frequently) and the cell output was being flooded with messages. I interrupted the cell, but realized that the subprocess doesn't receive the interrupt signal, and in fact won't stop because it waits for the closing signal from one of the queues. I had to terminate the process.

Could be more reason to go to ray for this type of execution?

Activity

  1. added this to the DRYML 0.3 milestone on Jun 27, 2026
  2. krafczyk commented on Oct 6, 2026

    @krafczyk
    CollaboratorAuthor

    Current-API review (2026-10-06): keep this issue open. The legacy @dryml.compute API is superseded, and Execute-backed dispatch.run now handles notebook KeyboardInterrupt: it tries confirmed pre-GO cancellation, otherwise requests post-GO cancellation, then re-raises without waiting in the shell (merged PR #53, 8161e8a84b4e90d13a7172c0ecaaa1f40b3fb5ce; src/dryml/dispatch/api.py:299-317). A real-kernel test in tests/dispatch/test_notebook.py verifies the shell/task remain usable after interruption, but allows either interrupt delivery or normal completion; a request is not proof that a worker stopped.

    Remaining gap: generic dryml.execute.Executor.run() and Core Executor.run() still call submit(...).result() without requesting cancellation on Ctrl-C (src/dryml/execute/executor.py:312-372, src/dryml/core/execute.py:2592-2652). Their one-off run() helpers attempt cleanup after a caught interrupt without first cancelling nonterminal work (src/dryml/execute/executor.py:1408-1431, src/dryml/core/execute.py:3012-3035). Thus the old observable symptom is still reachable through these supported direct Execute APIs; close(cancel=False) can wait for ongoing work. Post-GO stop must remain best-effort and report cancellation only after owned-process/group termination is confirmed, as existing backend tests require.

    Suggested fix: bring the Dispatch pre-GO-cancel/post-GO-request sequence to both blocking generic/Core run surfaces and leave cleanup with the owned future until terminality. Test blocking Ctrl-C, late worker termination, retained recovery handle, and subprocess/Windows semantics. This does not require claiming signal delivery directly to the worker, killing escaped descendants, or native Windows Ray support.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions