Repository navigation
fix(runner): prevent panic when timeout races with process completion - #288
Open
Souravrajvi0 wants to merge 1 commit into
Open
Souravrajvi0 wants to merge 1 commit into
Souravrajvi0 wants to merge 1 commit into
Conversation
When execution finishes near WORKER_TIMEOUT, the timer callback could still send on output channels after the completion path had already finished, causing "send on closed channel" and crashing the daemon. Wait for an already-fired timer callback before signaling done, guard the timeout error with sync.Once, and make channel sends recover from closed-channel panics as a safety net. Fixes langgenius#284
3 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes #284
Under high concurrency, when code execution finishes near
WORKER_TIMEOUT, thetime.AfterFunctimeout callback could still send on output channels after the normal completion path had finished. This causedpanic: send on closed channeland crashed the entire daemon process.Root cause
timer.Stop()does not prevent an already-scheduled timer callback from executing. When the python process exits right around the timeout deadline, the completion goroutine and the timer callback race — the callback may send onexecErrorafter channels are drained/closed.Changes
timer.Stop()returns false (timer already fired), wait for the timer callback to finish viasync.WaitGroupbefore signalingdone.sync.Onceso the timeout error is written at most once.recover()to prevent panics if a channel is already closed (defensive safety net, as suggested in the issue).Test plan
TestCaptureOutputTimeoutRaceDoesNotPanic— 200 iterations with execution time straddling the timeout deadlineTestSendBytesRecoversFromClosedChannel— verifies closed-channel sends do not panicTestCaptureOutput*tests pass