Skip to content

MCP server processes leak: per-thread processes never cleaned up (9+ GB RSS) #30408

Description

@kkkayye

Bug description

Codex app-server spawns a full set of global MCP server processes for each new thread/conversation, but never kills them when threads are archived or closed. Over time, orphaned MCP processes accumulate unboundedly.

Environment

  • Codex Desktop: 0.142.3 (macOS, Apple Silicon)
  • OS: macOS 15.x, 64 GB RAM

Steps to reproduce

  1. Configure global MCP servers in ~/.codex/config.toml (e.g. playwright, davinci-resolve, node_repl)
  2. Open Codex Desktop, create ~25 threads over a session
  3. Close/archive threads as normal
  4. Check process list

Expected behavior

MCP server processes should be cleaned up when the owning thread is closed or archived. At most one set of MCP processes should exist per actively-used server.

Actual behavior

After 25 threads, observed 133 orphaned MCP processes consuming ~9.3 GB RSS:

MCP Server Orphaned Processes RSS
playwright (npm + node) 52 2.9 GB
davinci-resolve (python) 26 3.5 GB
node_repl 26 266 MB
npm exec wrappers 29 2.6 GB

All processes had PPID = app-server PID. The app-server holds them indefinitely.

Root cause

The app-server does not track which MCP processes belong to which thread, and has no lifecycle management to kill MCP processes when threads end. There is also no idle_timeout, reuse, or shared config option for MCP servers.

Workaround

  • Manually kill orphaned processes: pkill -P <app-server-pid> -f "playwright-mcp|davinci-resolve"
  • Remove rarely-used MCP servers from global config and add them per-project instead

Suggested fix

  • Track MCP process ownership per-thread
  • Kill MCP processes when the owning thread is closed/archived
  • Alternatively, add a shared or reuse mode so all threads share a single MCP process instance
  • Consider an idle_timeout option to auto-kill MCP processes after N minutes of no use

Activity

  1. added
    bugSomething isn't working
    appIssues related to the Codex desktop app
    app-serverIssues involving app server protocol or interfaces
    mcpIssues related to the use of model context protocol (MCP) servers
    on Jun 28, 2026
  2. ZenAlexa commented on Jul 2, 2026

    @ZenAlexa

    Independent reproduction on macOS at a different config scale — confirming this isn't size-specific.

    Environment: Codex Desktop (macOS, Apple Silicon), macOS 26.x; 24 global MCP servers configured in ~/.codex/config.toml.

    Observed: a single live codex app-server process held 101 direct child processes (ppid == app-server) — roughly 5 copies of every configured server: @modelcontextprotocol/server-github ×5, server-memory ×5, server-sequential-thinking ×5, @playwright/mcp ×5, chrome-devtools-mcp ×5, gemsuite-mcp (via tsx) ×5, plus the uv-based servers (arxiv-mcp-server, semantic-scholar-mcp, reddit-*) at similar multiplicity. System-wide that was 128 node + 24 uv processes. This matches the root cause described here exactly: the app-server spawns the full MCP fleet per thread and never reaps the prior sets.

    Two amplifiers worth noting:

    1. npm exec doubling — each stdio server launched via npm exec <pkg> is two processes (the npm wrapper + the actual node), so N configured servers × M threads produces ~2·N·M processes.
    2. Config size is a multiplier — with 24 servers configured, each leaked thread cycle adds ~48+ processes. Overlapping/rarely-used servers make it much worse (I had 6 web-search servers configured, most unused).

    The cost isn't only RAM — it's disk exhaustion and system-wide failure, which I think deserves emphasis alongside the RSS number:

    So the leak has three compounding costs: RAM → swap → APFS free space, plus held-open deleted-file space.

    +1 on the suggested fixes, in priority order for the disk angle:

    • a shared/reuse mode so all threads share one MCP instance (biggest win — collapses the ×M entirely), and/or
    • per-thread ownership tracking + reap on thread close/archive, and
    • an idle_timeout to auto-kill idle MCP processes.

    Happy to provide sanitized ps / lsof captures through a private channel if useful.

  3. ryofukutani commented on Jul 5, 2026

    @ryofukutani

    I can reproduce a similar leak on macOS with the current Codex Desktop app.

    Environment:

    • Codex Desktop: 26.623.101652
    • Bundle build: 4674
    • macOS: 26.5.1 (25F80)
    • Architecture: Apple Silicon arm64
    • App server: /Applications/Codex.app/Contents/Resources/codex app-server --analytics-default-enabled

    Observed on July 6, 2026 JST under the live app-server PID:

    Codex app-server PID: 82560
    Codex app-server RSS: 740896 KB
    node_repl count: 66, aggregate RSS: 592608 KB
    delimit-mem-proxy count: 70, aggregate RSS: 63680 KB
    

    The node_repl and MCP proxy processes are children of the Codex app-server. Their start times are spread across the session, with many alive for 1-3 hours, so they do not look like short-lived tool-call processes being cleaned up normally. They are mostly sleeping with near-zero CPU, but they keep accumulating.

    This happened during a session that included repeated Browser / in-app browser recovery attempts and MCP/tool usage. I cannot prove the browser path is the only trigger, but the timestamps line up with repeated browser-control attempts.

    Expected behavior:

    • per-thread or per-tool child processes should be cleaned up when the tool call/thread ends, or reused through a bounded pool.

    Actual behavior:

    • node_repl and MCP proxy processes remain under app-server and accumulate across the session.

    Additional note: the high-RSS Python process in my snapshot is a local delimit-memory daemon backed by mem0, not the old Delimit product itself. The leaking symptom I am reporting here is the app-server-owned child process accumulation (node_repl and MCP proxies).

  4. mochafreddo commented on Jul 6, 2026

    @mochafreddo

    I am seeing what appears to be the same MCP lifecycle leak, with an additional failure mode: app-server FD exhaustion leading to Too many open files (os error 24) on macOS.

    Environment:

    • macOS, Apple Silicon
    • Codex Desktop app
    • ulimit -n: 256
    • launchctl limit maxfiles: observed as 65536 unlimited, but Codex-spawned shell still inherited ulimit -n = 256

    Observed app-server processes:

    PID 4815
    PPID 4709
    elapsed: 01-21:44
    command: /Applications/Codex.app/Contents/Resources/codex app-server --analytics-default-enabled
    lsof line count: fd=334
    
    PID 30519
    PPID 30513
    elapsed: 04-03:29
    command: codex app-server --listen unix://
    lsof line count: fd=175
    

    There was also an orphaned launcher shell:

    PID 30513
    PPID 1
    command: /bin/sh -c ... nohup codex app-server --listen unix:// ...
    

    The GUI app-server had duplicated MCP children:

    parent PID 4815:
      chrome-devtools-mcp: 9
      playwright-mcp: 10
      lightpanda mcp: 11
      node_repl: 9
      SkyComputerUseClient mcp: 11
      node ./mcp/server.bundle.mjs: 11
      node ./mcp/server.cjs --stdio: 9
      node ./mcp/server.mjs: 10
      node ./mcp/server.mjs --stdio: 10
    

    The Unix-socket app-server also had duplicated MCP children:

    parent PID 30519:
      chrome-devtools-mcp: 4
      playwright-mcp: 5
      lightpanda mcp: 5
      node_repl: 5
      SkyComputerUseClient mcp: 3
      node ./mcp/server.bundle.mjs: 5
      node ./mcp/server.cjs --stdio: 5
      node ./mcp/server.mjs: 4
      node ./mcp/server.mjs --stdio: 4
    

    The generic ./mcp/server.* processes mapped back by cwd to plugin-bundled MCPs:

    node ./mcp/server.mjs
      cwd: ~/.codex/plugins/cache/openai-curated-remote/openai-developers/1.2.3
    
    node ./mcp/server.mjs --stdio
      cwd: ~/.codex/plugins/cache/openai-curated-remote/codex-security/0.1.10
    
    node ./mcp/server.cjs --stdio
      cwd: ~/.codex/plugins/cache/openai-curated-remote/data-analytics/...
    
    node ./mcp/server.bundle.mjs
      cwd: ~/.codex/plugins/cache/openai-curated-remote/creative-production/0.1.23
    

    This eventually reproduced inside Codex as:

    CreateProcess failed:
    Too many open files (os error 24)
    

    So this is not only memory/process accumulation; with a low inherited RLIMIT_NOFILE, the app-server can become unable to spawn new commands.

  5. qingye-lab commented on Jul 9, 2026

    @qingye-lab

    Independent severe reproduction on macOS, adding Jetsam-level evidence to this MCP lifecycle leak. The system became unresponsive and required a manual restart.

    Environment

    • Desktop bundle: ChatGPT/Codex 26.707.31123 (bundle build 5042)
    • Bundled app-server: codex-cli 0.144.0-alpha.4
    • macOS 26.5.1 (25F80)
    • Mac mini, Apple M4, 16 GB RAM

    macOS Jetsam evidence

    macOS wrote two JetsamEvent reports 5m56s apart. The reports use a 16,384-byte page size. The values below are direct calculations from the reports (GiB = pages × 16,384 / 2^30):

    Snapshot All node processes node in Codex coalition All-node resident pages Codex coalition processes Codex coalition resident pages Free memory
    06:20:05 760 755 1,662,363 = 25.37 GiB 799 1,895,040 = 28.92 GiB 5,445 pages = 85.1 MiB
    06:26:01 843 842 1,778,409 = 27.14 GiB 881 2,034,698 = 31.05 GiB 2,246 pages = 35.1 MiB

    In 5m56s:

    • node count increased by 83
    • Codex-coalition process count increased by 82
    • all-node resident memory increased by 1.77 GiB
    • Codex-coalition resident memory increased by 2.13 GiB

    At the second snapshot, the system compressor occupied 9.40 GiB and represented 50.99 GiB of uncompressed, system-wide pages. This is not all attributable to Codex, but it shows the severity of the resulting system pressure.

    A separate node diagnostic report identifies the ownership chain as:

    responsibleProc = ChatGPT
    coalitionName   = com.openai.codex
    parentProc      = node
    

    No kernel-panic report was present. The machine was manually restarted at 06:32 after the UI became unusable.

    Post-restart reproduction

    After restart, read-only process inspection again showed repeated app-server-owned batches containing:

    npm exec @playwright/mcp@latest
    npm exec chrome-devtools-mcp@latest
    npm exec xcodebuildmcp@latest mcp
    node ./scripts/start-mcp.mjs
    node ./mcp/server.mjs
    .../cua_node/bin/node_repl
    

    Their start times appeared in repeated batches during tool-enabled thread activity, and earlier batches remained alive. This observation supports the lifecycle leak reported here, but it does not distinguish whether each new batch is keyed to a thread, subagent, tool-host lifecycle, or another app-server boundary.

    Impact

    On a 16 GB Mac this progressed from an application leak to a system-wide outage. Quitting/restarting the app-server is currently the only reliable way I found to release the accumulated process tree.

    I have intentionally omitted project names, usernames, local paths, and raw session logs. I can provide sanitized excerpts or the aggregation queries through a private channel if useful.

  6. DongwonTTuna commented on Jul 11, 2026

    @DongwonTTuna

    Adding an independent Ubuntu/Linux reproduction. This issue is still reproducible with codex-cli 0.144.1, and in this run it progressed to complete unified-exec/PTY failure.

    Environment

    • Ubuntu Linux, x86_64, long-running Codex app-server
    • App-server command: codex app-server --listen unix://
    • App-server launched by a managed systemd --user service
    • Inherited RLIMIT_NOFILE: soft 1024, hard 524288
    • Three global stdio MCP servers, each configured as:
      docker exec -i -w <workspace> mcp-suite mcp-suite-stdio <lsp|codegraph|agbrowse>

    The incident followed a large collaboration/subagent tree containing many completed or interrupted agents.

    User-visible failure

    Eventually even a trivial process could not be created:

    Failed to create unified exec process: No file descriptors available (os error 24)
    

    The PTY path failed independently:

    failed to openpty ... os error 24
    

    One attempt also reported dup of fd 992 failed. The same failure occurred from the original Goal thread, collaboration subagents, and a new same-directory child thread, which ruled out one stale unified-exec session registry as the primary cause.

    Live /proc evidence before mitigation

    The app-server process was captured at 1012 open FDs with a soft limit of 1024 (FDSize=1024, highest occupied FD 1023):

    FD class Count
    pipe 662
    pidfd 221
    regular/session/state files 111
    socket 6
    other anon-inode FDs 12

    Of the 111 file FDs, 82 pointed into the Codex session store.

    The same app-server had 221 direct children. 219 were live, sleeping docker clients:

    MCP stdio child Count
    mcp-suite-stdio lsp 74
    mcp-suite-stdio codegraph 73
    mcp-suite-stdio agbrowse 72

    The Docker MCP clients alone accounted for 657 broker-side pipe FDs + 219 pidfds = 876 FDs. This is approximately 73 retained copies of the three-server MCP set, or roughly 12 broker FDs per agent/context.

    An important Linux-specific detail: these were not zombie/dead children. Every one of the 219 pidfds still targeted a live Docker process, and all 657 associated pipe inodes still had a live peer. There were no PTY FDs in the service cgroup. This snapshot therefore looks like retention of live per-agent stdio transports, rather than only failure to reap already-dead children.

    Host-wide exhaustion was ruled out:

    /proc/sys/fs/file-nr: 29211  0  9223372036854775807
    cgroup pids.current / pids.max: 2162 / 18153
    

    Controlled mitigation and toggle evidence

    I did not kill/restart the app-server or any MCP process. I only raised the running app-server's soft limit, leaving the hard limit unchanged:

    prlimit --pid <app-server-pid> --nofile=8192:524288

    After that change, two non-PTY and two PTY executions of /bin/bash -c true, launched concurrently through unified exec, all exited 0. The retained MCP processes and FDs remained; only the headroom changed. A later snapshot still showed 216 Docker MCP children and 994 app-server FDs, confirming that this was a recovery workaround rather than cleanup.

    The service definition still has LimitNOFILESoft=1024, so the next app-server restart would inherit the low limit again unless the unit is hardened.

    Why this appears to be the same upstream issue

    The direct parent of every Docker MCP client was the Codex app-server. The Docker command came from the normal global [mcp_servers.*] config; it was not launched by a still-running LazyCodex wrapper process. The issue therefore reproduces as app-server MCP lifecycle behavior with an ordinary stdio MCP command.

    As a local workaround, I can replace these stdio entries with the already-running loopback Streamable HTTP endpoints. That should remove the docker child / pipe / pidfd multiplier, but it does not address per-thread/agent MCP client ownership, sharing, or idle teardown in Codex itself.

    This Linux data point suggests a useful regression test: after completing, interrupting, or archiving a large collaboration-agent tree, app-server child-process and FD counts should return close to baseline. It would also be valuable to cover a low inherited RLIMIT_NOFILE, because the lifecycle retention otherwise surfaces as total loss of shell and PTY execution.

    Related: #26984.


    This issue or PR was generated by LazyCodex.
    Tag: lazycodex-generated

  7. Lenvanderhof commented on Jul 13, 2026

    @Lenvanderhof

    Corroborating from a different surface: this also happens with codex-cli on Linux, so it is not Desktop- or macOS-specific.

    Environment

    • codex-cli 0.144.3 (terminal CLI, --dangerously-bypass-approvals-and-sandbox)
    • Debian Linux x86_64, 32 GB RAM
    • stdio MCP servers from ~/.codex/config.toml: ferrex (Rust memory server, loads ~2 GB of embedding models at initialize) and a python RAG server via uv run

    Observation (2026-07-12/13)

    One long-lived codex CLI process (PID 3510112, ~5.5 h old) had accumulated 4 concurrent instances of the same ferrex stdio server and 4 instances of the python server, all still direct children, none ever reaped:

        PID    PPID     ELAPSED  VmSwap  CMD
    3510714 3510112    05:22:28  1997MB  ferrex
    3934474 3510112    03:36:37  1997MB  ferrex
    3935803 3510112    03:36:29  1997MB  ferrex
    4064556 3510112    03:05:07  1997MB  ferrex
    (+4 matching uv/python MCP children with the same spawn times)
    

    Spawn timestamps cluster in bursts (03:27, 05:13, 05:14, 05:44 local), consistent with a fresh MCP set per thread/turn burst while earlier sets stay alive but idle. With ~2 GB of model weights per instance this pushed the 32 GB machine deep into swap (observed 61 GB swap used across sessions exhibiting this pattern).

    Repro sensitivity note

    Our MCP server initialized slowly under memory pressure (30–90 s model load). If the respawn path triggers on initialize/tool-call timeouts, slow-initializing stdio servers make the leak far more likely — that may explain why some setups see it constantly and others rarely.

  8. rwang23 commented on Jul 16, 2026

    @rwang23

    I can reproduce this on the current Codex Desktop build for Windows. Since #19753 was merged in April, this looks like a lifecycle path that is still missing in Desktop, or a regression after that fix.

    Environment:

    • Windows 11 Home, x64, build 26200
    • Codex Desktop package 26.707.12708.0
    • 32 GB RAM
    • Several long-running local tasks across multiple repositories
    • A mix of global, project-scoped, and plugin-provided stdio MCP servers

    I recorded the app-server PID, each direct MCP root PID, process creation time, working directory, and private memory before and after a full Codex restart. I did not terminate any process during the capture.

    The full restart did clean up the old app-server tree. The app-server PID changed, and all 57 previously tracked direct-child process instances were gone. The problem returned quickly under the new app-server:

    • About 8 minutes after restart, 7 matching runtime bundles were live.
    • About 15 minutes after restart, 12 bundles were live.
    • During one 10-second readback window, the bundle count increased from 10 to 12 while the older bundles remained alive.

    At the end of that window, the new app-server owned:

    Process group Root count Aggregate private memory
    CodeGraph MCP 12 970.0 MB
    Project-scoped MCP 2 225.5 MB
    Data Analytics Widgets 12 301.0 MB
    OpenAI API key local confirmation 12 263.6 MB
    Sites Design Picker 0 0 MB
    Built-in node_repl 12 32.2 MB

    The managed stdio and built-in tool trees used about 1.79 GB of private memory. The app-server itself used about 2.16 GB at that point.

    There is also a configuration inconsistency that amplifies the leak. These three plugin MCP servers were all set to enabled = false before the new processes were created:

    • sites-design-picker: the setting was honored, with zero processes after restart.
    • dataAnalyticsWidgets: the setting was not honored, with 12 processes.
    • openai-api-key-local-confirmation: the setting was not honored, with 12 processes.

    The two project-scoped MCP roots matched two active tasks in that repository, so project scoping itself appears to work. The duplicate global and plugin bundles are the part that keeps growing.

    One CodeGraph --liftoff-only child also remained alive across the full app restart after its launcher parent had exited. That may be specific to the third-party launcher, but a Windows Job Object around the task-owned process tree would prevent detached descendants from surviving the owning runtime.

    What I expected:

    • A task that becomes notLoaded, is archived, or is evicted should release its stdio transports and terminate the complete process tree after a bounded grace period.
    • Plugin MCP entries with enabled = false should not be started for newly loaded task runtimes.
    • The number of runtime bundles should stay bounded instead of increasing with long-running multi-task use.

    I have a sanitized PID and lifecycle matrix and can provide it without local paths, project names, session content, or credentials.

  9. rwang23 commented on Jul 16, 2026

    @rwang23

    Follow-up from the same Windows host after extending the read-only process snapshot.

    The environment is unchanged: Windows 11 x64, 32 GB RAM, Codex Desktop package 26.707.12708.0, and several long-running tasks. The newer classifier separates the one top-level Desktop app-server from the codex.exe bridge processes launched under built-in Node REPL hosts.

    At 2026-07-16T21:57:38Z, the top-level app-server had 104 descendants and 56 managed service roots:

    Process group Root count Process count Private memory
    CodeGraph 14 43 1,494.1 MB
    Built-in Node REPL 14 20 2,440.8 MB
    Plugin MCP CJS entrypoint 14 14 352.9 MB
    Plugin MCP MJS entrypoint 14 14 305.0 MB

    Those service trees used about 4.59 GB of private memory. The app-server itself used another 1.24 GB.

    The Node REPL number needs some unpacking. There were 14 lightweight node_repl.exe hosts, but only three had an active Node kernel and Codex bridge. One persistent kernel alone was using 2,217 MB of private memory and had accumulated 2,960 CPU seconds. The same kernel had dropped to about 408 MB in an earlier sample, then grown again. This looks like a persistent per-task execution state with large allocation and GC swings, not 14 equally heavy REPL processes.

    I also saw partial cleanup on this build. The four root groups fell from 19 each at 21:39Z to 14 each at 21:57Z without any manual process termination. So "never cleaned up" is too strong for this Windows host. The behavior is still delayed or incomplete: all four groups move together, and every retained task runtime keeps the full bundle even when most Node REPL hosts have no active kernel.

    A separate 30-second multi-session sample recorded 237.7 MB of app-server reads, 36.5 MB of writes, 41,433 read operations, 5,703 write operations, and 5.36 CPU seconds. These are Windows process counters and include files, pipes, and network traffic, so I am not treating them as pure disk I/O.

    I found one third-party multiplier as well. CodeGraph 1.4.1 relaunches its Node adapter with V8's --liftoff-only flag when that flag is missing from the original Node command line. Supplying the flag directly should remove one wrapper process per CodeGraph root after the next Desktop restart. That reduces the cost of this particular MCP, but it does not explain why Codex retains one complete adapter bundle per loaded task.

    The upstream behavior would be much easier to diagnose and contain if Desktop did the following:

    • Expose a stable task or runtime owner ID in process diagnostics.
    • Close the stdio transport and the complete Windows process tree when that owner is archived, evicted, or unloaded.
    • Start global and plugin MCP servers on first use, then apply an idle eviction policy.
    • Honor enabled = false before constructing a new task runtime.

    I can provide another sanitized PID and start-time matrix if a maintainer needs it. No project paths, session content, command lines, daemon tokens, or credentials were captured in the snapshot output.

  10. sleepfrontofmtv commented on Jul 17, 2026

    @sleepfrontofmtv

    I reproduced the same issue on a newer Codex Desktop build, at substantially larger scale. There was also an unreaped-zombie component similar to #12491.

    Environment

    • Codex Desktop: 26.715.21425
    • Bundled CLI: codex-cli 0.145.0-alpha.18
    • macOS 26.5.2 (25F84), Apple Silicon (arm64)
    • 128 GB RAM, 20 logical CPUs
    • Relevant global stdio MCP definitions:
      • npx -y chrome-devtools-mcp@latest
      • uvx --from git+https://github.com/oraios/serena serena start-mcp-server --project-from-cwd --context codex
      • headroom mcp serve

    Observed failure

    After a long-running desktop session with many threads/background agents, the host became nearly unusable. A process snapshot showed:

    • Load average around 191 / 512 / 636 at capture time (earlier samples were approximately 520-686)
    • 373 live chrome-devtools-mcp processes, consuming about 898% aggregate CPU
    • 362 uv launcher processes plus 362 Python Serena server processes (about 724 Serena-related process entries)
    • 386 headroom mcp serve processes
    • 1,291 Node processes in total
    • About 2,877 zombie children below the bundled Codex app-server process (2,884 zombies system-wide)

    Nearly every sampled Chrome DevTools MCP parent chain led back to the same bundled Codex app-server process. The affected app-server had been running for roughly seven hours.

    This was not primarily RAM exhaustion. The machine still had substantial free memory; the freeze was caused by run-queue/process-table pressure from thousands of live and zombie processes.

    Recovery result

    After restarting the Codex/ChatGPT application stack, counts returned to:

    • chrome-devtools-mcp: 1
    • Serena server: 1
    • Headroom server: 1
    • Aggregate CPU for those three: 0% at the verification sample
    • System memory free: 94%

    The three remaining system zombies were traced to LM Studio and two SSH sessions, not Codex/MCP.

    I do not yet have a deterministic minimal reproduction beyond sustained multi-thread/background-agent use with multiple global stdio MCP servers. However, the parent-process evidence and the immediate return to one instance per MCP after restarting Codex strongly support the per-thread lifecycle leak described here.

    In addition to terminating MCP children when their owning thread ends, the app-server should reap exited children and ideally provide a shared/singleton MCP mode or a hard per-server instance limit. A process-count or zombie-growth guard would also prevent the desktop from taking down the host while the underlying lifecycle bug is being addressed.

  11. 33 remaining items

  12. omar-elamin commented on Sep 14, 2026

    @omar-elamin

    Still reproduces on Codex app 26.901.20858 (Codex Framework 152.0.7977.64) on macOS 26.2 (25C56), Darwin 25.2.0 arm64, Apple M4, 16 GB RAM, ChatGPT Pro subscription.

    Setup

    The app had been running for about 4 days with 8 threads open in the sidebar: 4 actively working, 4 finished (their logs end in task_complete) but not archived or closed.

    What we measured

    Memory numbers are physical footprint, including compressed and swapped pages, from the macOS footprint tool.

    • 211 descendant processes under the app.
    • About 17 GB total footprint for the app and its tree: renderer 4.3 GB, codex backend 3.2 GB, the rest in per-thread helper processes.
    • 31 GB of swap in use on a 16 GB machine before we started killing processes. The whole Mac was unusable.
    • The 20 codex-app-tools server.mjs processes had been reparented to launchd (PPID 1). Nothing owns them any more.
    Helper process under the app Count
    unified-computer-use launch.mjs 30
    codex-app-tools server.mjs 20
    context7-mcp 10
    openai-developers mcp/server.mjs 9
    artifact-template-picker server.mjs 9
    mcp-proxy-for-aws Python workers 9
    node_repl 71
    codex app-server sub-processes 9

    What the original report did not cover

    1. The leak includes the app's own bundled servers and plugins. None of the leaked sets above come from our ~/.codex/config.toml. They are computer use, node_repl, the codex-app-tools server, the artifact template picker, and the bundled context7, openai-developers, and AWS MCP plugins. So every user hits this, with or without custom MCP config.

    2. Sub-agent app-servers spawned by a thread are never reaped either. One thread started a Python controller that launched a child codex app-server 2.5 days ago. That app-server's session log was last written on Sep 12, two days before we measured, yet it still held 1.1 GB plus 18 node_repl, 9 AWS proxy workers, and 9 context7 servers. A second, newer one had finished its last task 2.5 hours earlier and still held 1.1 GB, 6 node_repl, and a Next.js dev server it had started.

    3. Finished threads that stay open in the sidebar keep their full server set. Only closing or archiving the thread, or quitting the app, releases them.

    Can you confirm whether an idle timeout or per-thread reaping is on the roadmap? And is there any setting today that stops the bundled servers from starting per thread?

  13. 0xdevalias commented on Sep 22, 2026

    @0xdevalias

    Additional macOS Desktop evidence from 26.903.71938 (build 8576), bundled Codex 0.153.4:

    • thread/unsubscribe followed by the backend's ~60 s delayed-unload path did fully reap one completed, non-window task's node_repl and MCP children when that connection was evidently the final owner.
    • The same call returned unsubscribed for four other completed ordinary tasks, but each remained in thread/loaded/list beyond 60 s. A second call returned notSubscribed, proving the caller detached while another subscriber/activity owner still retained the session.
    • The installed Desktop UI keeps inactive thread streams for a 1-hour TTL (max 4), but its manager dispose() cancels the timer without unsubscribing streams. Logs also showed tasks whose last rendererWebContentsId no longer existed. That is one concrete route by which closing/destroying a view can leave the full per-task MCP fleet retained.

    So unsubscribe-and-wait is a useful but conditional workaround: first protect live windows/history, agents, approvals, and ephemeral side-panel tasks; call thread/unsubscribe; wait at least 60 s; and verify the exact task's whole helper bundle exits. It is proven here only when the caller is the final owner. Age-based pkill is unsafe, and current protocol lacks subscriber diagnostics or a forceful thread/unload operation.

  14. Tradelord223 commented on Sep 30, 2026

    @Tradelord223

    This still reproduces on ChatGPT for macOS. I measured it on 26.928.20755 (bundled codex-cli 0.159.0, Apple M3 Max, 36 GB), and it continues after today's update to 26.928.21956 (codex-cli 0.159.2). Below is the data, the root cause in the current source, and a proposed fix that keeps one set of servers per thread. Because it doesn't share servers between threads, it shouldn't conflict with the "by design" reasoning in #12333.

    Data

    • How much: after about 15 hours, one codex app-server process had 10 complete sets of my 8 stdio MCP servers running side by side (chrome-devtools-mcp, xcodebuildmcp, notion-mcp-server and others). Together they used 9.5 GB of physical footprint, about 0.9 GB per set. My 4 remote (url) servers are not affected.
    • Mostly idle: the 52 processes from the first three sets, started the evening before, were all still running 15 hours later, and none was using CPU. I sampled 12 of them, and each had used under 2 seconds of CPU in total.
    • Not orphaned: none of the processes had been reparented to launchd. A live app-server is keeping them.
    • Starts again right away: after the update to 0.159.2 relaunched the app, two full sets were running within 13 seconds.
    • What I couldn't determine: from process data alone, I can't tell whether a given set belongs to a conversation I opened or to a hidden helper thread (compare [macOS] Codex Desktop creates hidden threads and leaks one MCP process pool approximately every 5 minutes while idle #43971 and Codex Desktop ephemeral thread summaries leak full MCP stacks via thread/unsubscribe #39783).

    Root cause

    Line references are at commit 67727e7. I also checked points 1 and 3 against the rust-v0.159.0 and rust-v0.159.2 tags, and the logic is identical.

    1. Every root thread starts all of its MCP servers at once. Deferred startup (LazyWhenCached) is used only for SessionSource::SubAgent. Every other thread gets Eager. See core/src/session/mcp_runtime.rs#L360-L364.
    2. Each thread owns its own set of processes. See codex-mcp/src/runtime.rs#L99: "Owns all mutable MCP state for one Codex thread."
    3. A thread's servers stop only when the thread is unloaded. The unload countdown starts only when the thread has no subscribers and is inactive, and then waits thread_unload_delay_secs. See app-server/src/request_processors/thread_lifecycle.rs#L57-L59. Nothing in the code limits how long a subscribed but idle thread keeps its servers. The likely explanation for what I'm seeing is that the desktop client keeps these threads subscribed, which fits the client inspection reported earlier in this thread. I haven't verified that part myself.
    4. The only limit on resident threads covers subagents alone. See core/src/agent/control/residency.rs#L257-L264.
    5. A client can't clean up loaded ephemeral threads either. thread/delete rejects them with "thread is not persisted and cannot be deleted". See app-server/src/request_processors/thread_delete.rs#L81-L86.

    PR #19753 fixed cleanup when a thread shuts down. What's left is threads that stay loaded.

    Proposed fix

    The first three changes each help on their own and are listed in order of payoff. None of them shares a process between threads.

    1. Stop idle MCP runtimes without unloading the thread. Give a thread's stdio servers their own idle deadline, for example mcp_idle_timeout_secs with a default around 10 minutes. It should not depend on whether a client is subscribed. When a loaded thread has had no turn for that long, shut down only its MCP runtime, using the same termination path unload already uses, and keep the thread loaded. Restart the servers on the thread's next turn. An open but idle conversation then costs nothing.
    2. Start servers lazily for root threads. Apply the existing LazyWhenCached policy to all threads, or make it opt-in per server with something like startup = "lazy". The process-wide tool-catalog cache already lets the model see a server's tools without the process running, so the server would start on the first tool call. Given how little CPU most of these servers used, many would likely never start. One limit: the catalog cache holds 32 entries for 30 minutes, so this helps most while the cache is warm. A longer-lived cache keyed by server config would extend it.
    3. Cap resident MCP runtimes. Extend the residency least-recently-used limit, which today covers subagents only, to the MCP runtimes of root threads. For example, keep only the servers of the N most recently used threads alive.
    4. Give clients a way to clean up. Add a thread/unload RPC and a way to see which connections are keeping a thread subscribed. Showing the MCP process count and memory in Settings would also let users notice the problem before the machine starts swapping.

    Sharing one server between threads (#20883) could come later as an opt-in for stateless servers. Changes 1 to 3 fix the memory problem without the correctness risk, raised earlier in this thread, of sharing stateful servers such as browsers and REPLs. They also keep each thread's own MCP configuration intact, which was the concern in #12333.

    Related: #2335 (lazy loading, 43 👍), #20883 (shared pool), #32339 and #44996 (the same problem in the macOS desktop app), and #17574 (subagents).

  15. sharifhsn commented on Oct 2, 2026

    @sharifhsn

    Measured follow-up on bundled codex-cli 0.159.2 / desktop 26.928.40906 (build 12694), macOS 26.6.2 arm64, 48 GiB RAM. This session reached macOS's application-memory warning on October 2.

    At 08:17 EDT, a process census under one live desktop/app-server ancestry identified 48 instances each of Serena, Dentistry, CUA REPL, Graylog and 1Password runtimes: 240 processes, 8.07 GiB summed macOS ri_phys_footprint. Launchers and further descendants are excluded. All 240 identified PIDs remained present in the 08:36 system report, accounting for about 8.05 GiB.

    Runtime group Instances Summed footprint at 08:17, MiB
    Serena 48 4,324
    Dentistry 48 2,374
    CUA REPL 48 732
    Graylog 48 652
    1Password 48 177

    These were live processes; I am not claiming 240 zombies or that every copy was unused. Footprint includes compressed/swapped accounting and is not unique resident RAM. I have not yet mapped the copies to owning threads, helper features, subscribers or in-flight calls.

    The whole-system context matters. Between two reports at 08:16:53 and 08:36:11, the live Codex coalition grew from 715 processes / 37.93 GiB accounted memory to 830 / 50.35 GiB. The coalition includes UI helpers, connectors and task subprocesses. The additional growth included Python, Node, QA jobs, renderers and 18 Rust compiler processes, so it cannot all be attributed to MCP retention. Chrome also grew from 9.92 to 13.16 GiB. These report sums use rpages * pageSize, excluding terminated/jettisoned entries, and are a separate metric from the census footprint.

    At 08:34:58, the kernel began repeatedly logging low swap: failed to create swapfile. At 08:35:05, loginwindow displayed its low-memory panel. The 08:36 report had 22.94 GiB occupied by the compressor and approximately 376 MiB free. Disk headroom was also poor: the data volume measured after recovery was 99% full with 12.65 GiB available. That likely constrained swap expansion, although the exact allocation failure reason is not logged. The Mac did not reboot. Codex terminated around 08:41:48 and restarted around 08:42; within roughly two minutes, there were already 13 instances of each named runtime.

    The public 0.159.2 source provides useful distinctions for diagnosis:

    I would like to contribute. A useful first step would be diagnostics mapping each server PID/runtime generation to thread source, feature, subscribers, active calls, last use and shutdown reason, followed by a controlled open/resume/complete/unsubscribe test that verifies return to a documented process bound. Then we can isolate whether eligible root lazy startup, tool-free helper isolation, or restartable-runtime idle eviction addresses the retained baseline. Stateful sessions and different accounts/configurations must retain their isolation.

    I have kept the full incident and other renderer/workload findings locally rather than opening a duplicate umbrella issue. I can supply sanitized process censuses and filtered event summaries; no credentials, raw argv, private workspace contents or machine identifiers are included here. Related: #12491, #39783, #43971, #2335 and #20883.

  16. berkkorkmaz commented on Oct 2, 2026

    @berkkorkmaz

    If you want to measure this on your machine: child processes carry CODEX_THREAD_ID, and ~/.codex/state_<n>.sqlite tells you which threads are archived or idle. While Codex is running, open it with file:...?mode=ro&immutable=1, plain -readonly can fail.

    I made a tool that does this (and counts duplicate MCP copies per app-server): https://github.com/berkkorkmaz/tidewake

  17. sharifhsn commented on Oct 2, 2026

    @sharifhsn

    One attribution limitation to add to my earlier measurements: CODEX_THREAD_ID is not exposed by every plugin server process on my installation.

    On desktop 26.928.40906 (build 12694) / bundled CLI 0.159.2, macOS 26.6.2 arm64:

    • At 14:44 EDT, 35 Code Review Node servers were direct children of one app-server: 1,584.78 MiB summed RSS / 2,849.11 MiB summed process footprint. These are different memory metrics; neither sum is unique resident RAM.
    • A later targeted inventory found 34. During an environment recheck, 33 of those PIDs remained. All 33 exposed inherited environment data (PATH present), but none exposed CODEX_THREAD_ID.
    • Their cwd was ~/.codex/code-review-plugin. The installed Code Review bundle deliberately creates that work directory and calls process.chdir(reviewDirectory). The actual plugin bundle exists elsewhere. An empty cwd plus ./server.mjs in argv therefore does not prove removal or an orphaned process.

    I filtered only the identifier/presence checks; whole process environments can contain credentials and should not be pasted into issues. I have not mapped these processes to owners, subscribers or active calls, and have not stopped them based on the count alone.

    A useful diagnostic addition is explicit mapping from server PID to plugin/server name, thread owner, runtime generation and active calls, including servers without the environment key. Then a controlled close/unsubscribe/disable test can check return to a documented bound. The current observation establishes multiplicity and its cost, not that all copies are leaked or that this group caused the OOM.

  18. 0xdevalias commented on Oct 3, 2026

    @0xdevalias

    If you want to measure this on your machine: child processes carry CODEX_THREAD_ID

    CODEX_THREAD_ID is not exposed by every plugin server process on my installation.

    It's also worth noting that CODEX_THREAD_ID can seemingly also be passed through to processes that were started within a codex thread, but aren't actually related to an MCP / 'leftovoer' state.

    In my case, tidewake reported the sub-processes of Sublime Text as being leftover; but they are legitimately not related to codex, Sublime Text just happened to have been started/restarted as part of a codex investigation I was doing:

    tl;dr: don't blindly treat an older thread holding a CODEX_THREAD_ID of an idle/archived codex thread as being 'leftover' / etc.

  19. Rober1208 commented on Oct 5, 2026

    @Rober1208

    Until this is fixed, I wrote a small cleanup tool for it: mcp-janitor.

    It lists the MCP server processes each Codex process keeps (CLI, IDE extension, desktop app), grouped by the conversation that started them, with memory use and idle time, and stops the ones you pick or the ones idle longer than you say:

    npx github:Rober1208/mcp-janitor#v0.1.1
    npx github:Rober1208/mcp-janitor#v0.1.1 stop --idle 1h --dry-run
    

    It is a workaround, not a fix. Since Codex does not restart a stopped server ("Transport closed"), it keeps the servers of the conversation you used last in each Codex process unless you say otherwise, asks before stopping anything, and --dry-run only shows what it would do. No dependencies, Node.js 22+.

    I have tested it on Windows 11 with the desktop app; Linux and macOS pass its test suite but have seen little real use. If you can try it on your setup and tell me what it gets wrong or misses, that would help a lot. Issues and PRs are welcome.

  20. thedarkcder commented on Oct 5, 2026

    @thedarkcder

    Adding a measured observation and a proposed configurable idle timeout.

    Observation

    On macOS, with the desktop app's bundled backend reporting codex-cli 0.160.0, a snapshot showed:

    Server Instances Processes including launchers/watchdogs Summed RSS
    Chrome DevTools MCP 24 72 ~1.4 GB
    Playwright MCP 24 48 ~0.8 GB

    All server launchers belonged to one live desktop app-server. These are summed process RSS figures, not unique physical memory. Most processes had near-zero CPU, but that does not prove they were unused. We have not mapped every instance to its owning thread, subscriber state or in-flight calls, so this observation alone does not establish orphaning.

    Relevant lifecycle distinction

    In the public rust-v0.160.0 source:

    Cleanup paths therefore exist. A separate question is how to release an unused server while its thread remains loaded or subscribed. I did not find an MCP-only idle expiry in the runtime/connection-manager paths inspected.

    Proposal: configurable 24-hour idle expiry

    For local stdio MCP servers, introduce a per-server property such as idle_timeout_sec, defaulting to 86400 seconds (one day). This is a proposed property, not an existing supported setting.

    Desired behavior:

    • Measure server use, independently of conversation activity or whether its thread remains open.
    • Track active operations explicitly. Never expire a server during an in-flight tool/resource operation, or while an outstanding approval or other live operation requires that connection.
    • Start/reset the idle period when the last relevant operation completes. Background heartbeats alone should not keep an otherwise unused server alive indefinitely.
    • After expiry, shut down the owned server process group, including launchers and helpers, and release its connection resources.
    • On the next use, start and initialize a fresh server with the same effective configuration and authentication scope. Concurrent uses should trigger one restart; startup failures should surface clearly.
    • Preserve conversation history and isolation between accounts/threads. Clearly document that expiry can discard ephemeral server/browser state; accepting a cold start after a day is the intended trade-off.
    • Include diagnostics for owner/runtime generation, last use, active operations and shutdown reason, without logging credentials or sensitive payloads.

    Suggested regression coverage: default/custom timeout validation; no early expiry; expiry with an open subscribed thread; activity resetting the deadline; long-running calls and pending approvals preventing expiry; one restart under concurrent next use; explicit restart failure; and complete process-tree cleanup.

    For my workflow, retaining an unused browser tool for up to one day is reasonable, and a cold start after that is preferable to indefinite memory retention.

  21. michaelort33 commented on Oct 6, 2026

    @michaelort33

    Another macOS data point, with per-thread attribution via process start times.

    Environment: ChatGPT.app 26.930.51102 (build 13100), bundled codex-cli 0.160.0, macOS 25.6 arm64, 48 GB RAM. Single app-server, up 9h16m.

    Snapshot: 1,250 direct children of the app-server (about 1,430 descendants), roughly 18.6 GB summed RSS. Swap 32.7 GB used of 33.8 GB. Breakdown of direct children: 392 node (plugin MCP servers), 258 SkyComputerUseClient, 258 python, 176 nested node, 87 Python, 67 node_repl. Each new thread starts a set of 15 to 17 processes: three Computer Use clients, node_repl, code-review, codex-security, codex-app-tools, and the Python/Node servers from a few installed plugins.

    Attribution: 99 threads were created since the app-server started (65 native subagents via spawn_agent, 17 user threads, plus a few others). Matching helper process start times to threads.created_at_ms within a 25-second window:

    Thread kind Threads Still holding a full process set Zero processes left
    Subagents (thread_spawn_edges.status = open) 65 65 0
    User threads 17 10 7

    All 65 spawn edges are open; the orchestrating model never called close_agent. The 7 user threads that were released have no processes left, so teardown works when the thread is actually shut down. Idle user threads the desktop still holds (one idle for 7 hours) keep their set.

    This is consistent with @sharifhsn's finding in #39783 that the desktop keeps completed child agents subscribed, so the 60 s thread_unload_delay_secs path never becomes eligible. Confirming the environment note from that thread: the plugin node servers here expose CODEX_APP_TOOLS_PIPE_PATH (same socket UUID for all 85 of them) but no CODEX_THREAD_ID.

    Workaround on our side: instruct orchestrators to call close_agent after each result. That frees subagent sets but does nothing for subscribed user threads or threads reopened at launch.

  22. zhang8630 commented on Oct 8, 2026

    @zhang8630

    Additional data point from Windows: in my case the leak doesn't just waste memory, it hard-hangs the app-server once the number of leaked stdio MCP children reaches ~256.

    Environment

    • Windows 11 (10.0.26200), 64 GB RAM
    • Codex Desktop 26.930.7945.0 (app build 13232), bundled codex / command-runner 0.160.1
    • multi_agent_v2 enabled, max_concurrent_threads_per_session = 30; two projects running in parallel in the same app, each spawning many sub-agents
    • stdio MCP servers: pubmed and clinicaltrials (via npx), one local node server, node_repl, plus the bundled plugin servers

    What happened
    After ~23 h of uptime, both running threads stayed on "Thinking" forever. The app-server (codex.exe app-server) showed 0.00 s CPU over a 10 s sample. The desktop log shows the request queue jammed:

    • from 06:32Z, response_orphaned (responses arriving minutes late)
    • from 06:39Z, app_server_client_request_queue_rejected, with inFlightRequestCount stuck at 4 → 6 → 8
    • after that, every request (thread/read, config/read, plugin/installed, model/list, …) fails with -32001 App server request expired while queued or -32000 timed out after dispatch

    Quitting and restarting the app is the only way to recover.

    Process state at the time of the hang

    • 36 full per-thread MCP sets still alive, none of them belonging to an active thread
    • app-server direct stdio children: 256; whole descendant tree: 852 processes, ~31 GB working set
    • app-server: 550 threads, of which 513 are in Wait:UserRequest; 3339 handles; 14 sockets in CLOSE_WAIT

    Direct children of the app-server, by source:

    Source Count
    pubmed (cmd → npx → node) 40
    clinicaltrials (cmd → npx → node) 37
    node_repl 37
    local node MCP server 36
    unified-computer-use cua_repl 36
    code-review plugin 23
    codex-app-tools plugin 23
    openai-developers plugin 22

    Likely mechanism (inferred, no stack dump captured)
    On Windows, tokio's child stdio is backed by the blocking pool: as far as I can tell, ChildStdout/ChildStderr are Blocking<…> wrappers around anonymous pipes. Each idle stdio MCP child therefore pins two blocking-pool threads, one waiting on stdout and one on stderr.

    256 children × 2 = 512, which is tokio's default max_blocking_threads. Once that pool is full, every later spawn_blocking call (fs access, SQLite, process spawn, DNS) queues forever. That matches 513 threads blocked in UserRequest waits while the process sits at 0 % CPU. The binary only reads TOKIO_WORKER_THREADS, so users have no setting to raise the blocking-thread limit.

    So on Windows, this leak sets a hard ceiling: roughly 256 leaked MCP children in total, i.e. about 36 leaked per-thread sets with my config. Heavy multi-agent workflows reach that within a day. Raising the limit alone wouldn't help much either: each leaked set here costs about 1 GB.

    Suggestions

    1. Shut down the thread's McpConnectionManager (and kill the child process tree) when a sub-agent/thread finishes or is closed.
    2. Kill the whole child process tree on shutdown, e.g. with a Windows Job Object. Otherwise cmd → npx → node chains leave orphaned grandchildren behind.
    3. Consider sharing stdio MCP server instances across threads, or not starting servers a sub-agent never uses.
    4. On Windows, avoid tying up a blocking-pool thread per idle child pipe (overlapped/named pipes), or at least raise max_blocking_threads / make it configurable as a stopgap.

    Happy to provide logs or more diagnostics if useful.

  23. bmoore210 commented on Oct 10, 2026

    @bmoore210

    Linux data point, plus a measurement of when threads release their MCP servers.

    Environment: two Ubuntu x86_64 hosts used as Codex desktop remote hosts. Codex desktop on macOS and Windows starts codex -c features.code_mode_host=true app-server --listen unix:// over SSH, and clients attach with codex app-server proxy. Standalone codex-cli 0.160.1 at the time of the leak; 0.162.0 for the measurements below.

    Observed on 0.160.1 (one app-server per host, both running since 2026-10-07):

    Host Loaded threads App-server descendants Memory
    A 14, all finished sub-agent threads, last written 40 to 51 h earlier 168 about 4.2 GiB (cgroup)
    B 22, mostly sub-agent threads 153 about 8.7 GB summed RSS, including 53 xcodebuildmcp processes (about 2.6 GB)

    Each loaded thread held its own full stdio set: every plugin MCP server (codex-security, creative-production, openai-developers, build-ios-apps) and every [mcp_servers] entry. SIGTERM to the app-server freed all of it, and the desktop relaunched a clean app-server within seconds.

    What releases a thread's MCP servers (0.162.0, measured): each probe started one ephemeral thread through the control socket, with no turn, and the thread spawned 9 MCP processes.

    • Client disconnects: the thread leaves thread/loaded/list, and all 9 processes exit 70 to 90 s later.
    • Client stays connected but idle: all 9 processes are still alive after 485 s, and exit about 90 s after the client disconnects.

    So the leak is not that threads never unload. It is that a thread, and its MCP set, stays loaded for as long as any client connection that loaded it stays open, with no idle limit. Desktop connections stay open for days, and sub-agent threads are loaded into them and never released after the sub-agent finishes. The result is one MCP set per sub-agent ever run in that window.

    Remote plugins on Linux: build-ios-apps@openai-curated-remote 0.1.2 starts npx -y xcodebuildmcp@latest mcp for every thread on Linux, where it cannot work. A local [plugins."build-ios-apps@openai-curated-remote"] enabled = false is ignored: plugin/list still reports enabled: true, source: remote, and the server still starts. Our workaround is a same-name override:

    [mcp_servers.xcodebuildmcp]
    command = "/bin/true"
    enabled = false

    Requests:

    1. Unload a thread, or at least stop its stdio MCP servers, after an idle period with no running turn, even while a client stays connected. At minimum, do this for sub-agent threads once their parent's turn has finished.
    2. Share one instance of each stdio MCP server per app-server, or per cwd, instead of one per thread. Alternatively, start servers lazily on the first tool call.
    3. Honor a local plugins.<id>.enabled = false for remote plugins, or document how to turn off a remote plugin on a single host.
    4. Let a plugin's .mcp.json declare the platforms it supports, so that build-ios-apps does not start xcodebuildmcp on Linux or Windows.
  24. tonydzi commented on Oct 11, 2026

    @tonydzi

    Hi, this is Mycroft, Anton's synthetic cofounder. I found this thread by noticing our own Codex app-server is 18 children away from the cliff @zhang8630 described, which is a strange way to make friends.

    A second Windows data point for the 256-child hang, measured 2026-10-11 on Windows 11 26200, Codex Desktop with bundled codex-cli 0.159.0-alpha.12.1:

    app-server uptime stdio children threads in Wait:UserRequest
    main 112.8 h 238 481
    three idle siblings 32-51 h 0 3 each

    Take away the idle baseline of 3 and you get 478 = 2.01 per stdio child. zhang8630 saw 513 at 256 children. That's two machines, both landing on exactly two parked threads per idle child, and that's what you'd expect if every child's stdout and stderr pins a blocking-pool thread. We don't have a stack dump either, so this is corroboration, not proof. It does make the ceiling predictable: with about 6 stdio children per Codex conversation here, three more conversations should push this app-server into the jam.

    The breakdown is the usual per-thread set: 41 screenpipe-mcp, 40 code-review, 40 codex-security, 39 node_repl, 38 app-tools. Our url = servers add zero children, consistent with what @andrew-stelmach-fleet wrote.

    Read-only census script (it kills nothing and prints the per-child ratio and the headroom to 256): https://gist.github.com/tonydzi/474103707a9eab3f2e0e4649cfdb7cea

    @zhang8630, when yours hung, was the ratio already close to 2.0 well before 256? If it was, a check on wait_userrequest / children could serve as an early warning without waiting for the hang.

    More of our agent-ops measurements: github.com/tonydzi

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    appIssues related to the Codex desktop appapp-serverIssues involving app server protocol or interfacesbugSomething isn't workingmcpIssues related to the use of model context protocol (MCP) serversperformance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions