Skip to content

[Linux Desktop 26.1007] Background memory extraction of inactive 7 GiB chat OOMs native backend; no chat opening required #53039

Description

@SpinGiantCRM

What version of the Codex App are you using (From “About Codex” dialog)?

ChatGPT-branded Linux desktop app 26.1007.21434 (package chatgpt-desktop-bin 26.1007.21434-1).
Bundled backend: codex-cli 0.162.0-alpha.17.2. Both verified locally. The bundled backend is byte-for-byte identical to the official release binary; its source commit is 740e5af33c71225640e0c1c1555c514c2c93ab74.

What subscription do you have?

Not collected.

What platform is your computer?

CachyOS Linux x86_64, kernel 7.2.9-2-cachyos-bore, KDE Plasma 6.7.5 / Wayland, RTX 4080 SUPER / NVIDIA 615.78.08. Approximately 30.5 GiB usable RAM and 30.5 GiB swap.

What issue are you seeing?

The second shutdown was triggered by background memory extraction of an inactive, oversized chat. The user did not open that chat. The same underlying full-history loader also reproduces runaway allocation in an isolated backend.

The rollout is 7,507,994,138 bytes (6.99 GiB) with 106,635 JSONL records; its largest record is approximately 17.08 MiB. Large records include persisted compaction replacement histories. A content-free scan found 87.60% of the file's serialized bytes are embedded image URLs, including 1.54 GB of exact repeated URLs. The detailed breakdown is in the synthetic-reproduction follow-up. Contents and identifiers are withheld.

In both actual app failures, backend logs recorded Resuming rollout from "<same oversized rollout>" shortly before the OOM kill:

October 11, Australia/Adelaide (UTC+10:30) Rollout load begins App killed Gap
First incident 14:38:32 14:39:16 44 seconds
Second incident 15:56:40 15:57:33 53 seconds

The second app instance had been running only 3 minutes 52 seconds. The user had two ordinary chats and one Work chat active and explicitly confirms they did not open the oversized old chat.

Background trigger evidence: a read-only query of the local memories_1.sqlite job table found the oversized chat's job with:

  • kind = memory_stage1
  • status = running
  • started_at = 1791696400 (2026-10-11 05:26:40 UTC / 15:56:40 Adelaide)
  • finished_at = NULL
  • lease_until = 1791700000 (one hour later)
  • The job's worker was the different, currently active Work thread.

The job start matches the oversized rollout's backend load log to the second. Another small chat's memory job was claimed and loaded concurrently. This establishes background memory extraction as the second incident's trigger; it does not require navigation to the oversized chat. The first incident loaded the same oversized history, but its earlier job claim has been overwritten, so the exact first caller remains unverified.

The log text Resuming rollout from ... is emitted by the generic load_rollout_items loader and does not prove a user-facing thread/resume request. Earlier framing of this as opening/resuming the old chat was too narrow.

Local config has features.memories = true, memories.generate_memories = true, and memories.use_memories = true.

What steps can reproduce the bug?

A self-contained synthetic reproduction is now available in this follow-up. It needs no private history or database. With a constant 1 MiB final replacement history, increasing saved snapshots from 1 to 256 raised sampled backend RSS from 342 MiB to 1,058 MiB. A ~512 MiB synthetic archive exhausted an isolated 2 GiB diagnostic cgroup. The follow-up includes the full script, successful parse checks, kernel evidence, exact release verification, and fix/test targets.

The original private-history reproduction follows:

A controlled local reproduction used the installed bundled backend, a separate CODEX_HOME, an independent copy of the rollout and a filtered copy of the local thread metadata. The live chat/database was not changed. No turn was started and no authentication files were copied.

  1. Launch codex app-server --stdio inside a separate systemd scope with MemoryMax=6G and MemorySwapMax=0.
  2. Initialize the protocol connection.
  3. Request:
    {"id":2,"method":"thread/resume","params":{"threadId":"<fixture-thread-id>","excludeTurns":true,"cwd":"<isolated-fixture-directory>","config":{"mcp_servers":{}}}}
  4. Measure backend RSS while waiting for the response.

Results with the 6.99 GiB rollout:

  • At 2 seconds: 3,193,164 KiB RSS.
  • At 4 seconds: 6,188,436 KiB RSS.
  • At approximately 4.43 seconds, the diagnostic scope hit its 6 GiB hard limit.
  • Kernel event: CONSTRAINT_MEMCG; victim: codex; anonymous RSS at termination: 6,269,660 KiB.
  • The isolated scope reported Result=oom-kill; the normal desktop app remained running.

Control: the same metadata-only resume procedure with a different 28,905,963-byte rollout completed in 0.42 seconds, with sampled peak RSS 405,472 KiB and exit code 0.

Thus excludeTurns=true does not protect this installed backend from the oversized rollout load.

What is the expected behavior?

Background memory extraction and chat resume must not materialize an unbounded archive into backend memory. Extraction needs streaming/filtering and a bounded input budget before retaining records. Oversized inputs should be skipped with a recoverable diagnostic. A metadata-only resume should also avoid system-wide memory exhaustion.

Additional information

Actual watchdog evidence:

  • First killed app-owned scope: 19.90 GiB RAM peak, 17.75 GiB swap peak; system RAM/swap at the decision: 96.4% / 90.3%.
  • Second killed app-owned scope: MemoryPeak=16026378240 and MemorySwapPeak=15437815808 (approximately 14.93 / 14.38 GiB).
  • In both incidents the ChatGPT main process received SIGKILL.
  • Scope ownership was verified from journal metadata: the Chromium-named scope contained /usr/lib/chatgpt/ChatGPT; after relaunch the equivalent scope also contains the app-server/helpers.

During subsequent bounded sampling, the normal app-server stayed approximately 368–463 MiB RSS, and process counts fell 32 → 20 as helpers exited. The isolated failing test directly identifies the native backend as an allocation owner.

The verified release's source is consistent with the measured behavior: RolloutRecorder::load_rollout_items parses rollout records and retains them in a vector; get_rollout_history uses that loader. The installed binary hash matches the official release built from this tagged source. This is source inspection, not a captured allocation backtrace.

The verified release's source also identifies the background path: memory startup asynchronously launches phase 1 for eligible root sessions; phase1::sample first calls RolloutRecorder::load_rollout_items and only afterwards filters or applies the memory input token budget. The pipeline documentation confirms that eligible idle chats are processed in the background. These references are pinned to the official release commit matching the installed backend.

This closely matches #29510, which already describes huge rollout histories causing app-server growth to tens of GB. Please investigate the continuing failure on this newer backend, particularly automatic memory extraction of inactive chats and cold resume with excludeTurns=true.

The original report also noted repeated access-denied hydration requests and IAB errors. Those observations remain valid, but the new reproduction supplies a much stronger cause; their causal role has not been established.

The saved chat and its original history remain intact. Public evidence contains only metrics, record counts and sanitized log signatures.

Activity

  1. github-actions commented on Oct 11, 2026

    @github-actions
    Contributor

    Potential duplicates detected. Please review them and close your issue if it is a duplicate.

    Powered by Codex Action

  2. changed the title [-][Linux Desktop 26.1007] Work session OOM-killed after 6h: app-owned scope peaks at 19.9 GiB RAM + 17.75 GiB swap[/-] [+][Linux Desktop 26.1007] Oversized saved rollout OOMs native backend: 7 GiB history, 6 GiB RSS within 4s on metadata-only resume[/+] on Oct 11, 2026
  3. changed the title [-][Linux Desktop 26.1007] Oversized saved rollout OOMs native backend: 7 GiB history, 6 GiB RSS within 4s on metadata-only resume[/-] [+][Linux Desktop 26.1007] Background memory extraction of inactive 7 GiB chat OOMs native backend; no chat opening required[/+] on Oct 11, 2026
  4. SpinGiantCRM commented on Oct 11, 2026

    @SpinGiantCRM
    Author

    Further investigation adds an exact release match, a credential-free synthetic reproduction, and a content-free breakdown of the real oversized history.

    Exact binary/source match

    The installed bundled backend is byte-for-byte identical to the official rust-v0.162.0-alpha.17.2 Linux x86_64 musl release. I downloaded the official release archive, verified its published SHA-256, decompressed it and compared the binary hash:

    Installed binary SHA-256:
    73b73b048546993298234a295de01ef2424e7a8b29def9d0ecb1ca7cb77500d4
    Official release binary SHA-256:
    73b73b048546993298234a295de01ef2424e7a8b29def9d0ecb1ca7cb77500d4
    Release tag commit:
    740e5af33c71225640e0c1c1555c514c2c93ab74
    

    The installed loader logs also identify rollout/src/recorder.rs:1108 (load starts) and :1161 (load completes), matching that release's source.

    What made the original history large

    A streaming scan of the original 7,507,994,138-byte file, without exporting conversation contents, measured:

    Record category Count Serialized bytes
    Custom tool outputs 11,933 5,165,329,265
    Compaction snapshots 171 1,656,273,161
    Other records 94,531 686,391,712

    Embedded data:image/... URLs account for 6,577,033,437 bytes (87.60%) across 4,868 occurrences. That includes 5,105,045,283 bytes in custom tool outputs and 1,453,847,178 bytes in compaction snapshots. Local hash equality checks found 2,066 unique image URLs and 1,543,511,104 bytes of exact repeated image URLs. Only counts and byte sizes are disclosed; no image data or hashes are posted.

    This is serialized image storage, not an estimate of decoded pixel memory. It does not establish a leak in screenshot capture. It explains why loading this persisted history is unusually costly.

    Synthetic reproduction: no private history or database required

    The script below creates a fresh temporary CODEX_HOME, starts a new thread without starting a model turn, and writes repeated valid paginated compaction records with ordinals. Every compaction's replacement history contains the same 1 MiB text; the final replacement history stays constant while the stored archive grows. It then cold-resumes the thread with excludeTurns=true.

    Memories are disabled in the synthetic fixture. No auth files, original chats or original databases are used. Thus the synthetic test exercises the shared history/resume path; the separately recorded memory job establishes the real second shutdown's background caller.

    Observed on the identical official release binary:

    Saved snapshots Approximate file size Sampled peak backend RSS Outcome
    1 1.00 MiB 342.42 MiB Resume completes, 0.025 s
    32 32.01 MiB 494.03 MiB Resume completes, 0.175 s
    128 128.02 MiB 791.89 MiB Resume completes, 0.599 s
    256 256.05 MiB 1,058.05 MiB Resume completes, 1.167 s
    512 ~512.10 MiB Kernel recorded 1,542,640 KiB anonymous RSS + 118,172 KiB file RSS for codex Isolated 2 GiB cgroup OOM

    All four successful cases logged loader completion with zero parse errors. For the 512 case the journal reported CONSTRAINT_MEMCG, scope Result=oom-kill, MemoryPeak=2147639296, and MemorySwapPeak=0; both the diagnostic Python supervisor and codex were killed. No sample/result file survived that case. The scope elapsed 2.739 s including fixture generation and both backend startups, not a measured resume-only duration.

    The 2 GiB scope figure includes file cache and supervisor memory; it is not presented as backend RSS. The normal desktop app and its original backend stayed running.

    Run one case at a time on Linux with a hard diagnostic limit:

    systemd-run --user --scope --unit=codex-rollout-repro-128 \
      -p MemoryMax=2G -p MemorySwapMax=0 \
      python3 reproduce-rollout-allocation.py /path/to/codex --snapshots 128

    Change 128 to 512 and use a fresh unit name for the failing case. The script attempts an RSS-based early stop, but a fast allocation burst can hit the cgroup limit before that monitor responds. The hard limit is required for the failing case.

    Self-contained Python reproduction
    #!/usr/bin/env python3
    """Offline synthetic history test; never accesses an existing CODEX_HOME.
    
    Example: python reproduce-rollout-allocation.py /path/to/codex --snapshots 128
    Run under a 2 GiB cgroup limit. The RSS monitor attempts an early stop,
    but a fast allocation burst can reach the hard limit first (observed at 512).
    No model turn is started. All text and thread state are created from scratch.
    """
    import argparse
    import hashlib
    import json
    import os
    from pathlib import Path
    import queue
    import subprocess
    import tempfile
    import threading
    import time
    
    
    def main():
        ap = argparse.ArgumentParser()
        ap.add_argument('codex', type=Path)
        ap.add_argument('--snapshots', type=int, default=128)
        ap.add_argument('--text-mib', type=int, default=1)
        ap.add_argument('--stop-rss-mib', type=int, default=1200)
        ap.add_argument('--result', type=Path)
        args = ap.parse_args()
        binary = args.codex.resolve()
        if not 1 <= args.snapshots <= 512 or not 1 <= args.text_mib <= 4:
            ap.error('snapshots must be 1..512 and text-mib must be 1..4')
        result = {'snapshots': args.snapshots, 'text_mib': args.text_mib,
                  'version': subprocess.check_output([str(binary), '--version'], text=True).strip(),
                  'binary_sha256': hashlib.file_digest(binary.open('rb'), 'sha256').hexdigest()}
        with tempfile.TemporaryDirectory(prefix='codex-synthetic-rollout-') as td:
            root = Path(td)
            home = root / 'codex-home'
            home.mkdir()
            (home / 'config.toml').write_text('[features]\nmemories = false\n')
            env = {'PATH': os.environ.get('PATH', '/usr/bin:/bin'), 'HOME': str(root),
                   'CODEX_HOME': str(home), 'XDG_CONFIG_HOME': str(root / 'xdg'),
                   'RUST_LOG': 'codex_rollout=debug'}
    
            def launch():
                err = tempfile.TemporaryFile()
                proc = subprocess.Popen([str(binary), 'app-server', '--stdio'], cwd=root,
                                        env=env, stdin=subprocess.PIPE, stdout=subprocess.PIPE,
                                        stderr=err)
                messages = queue.Queue()
    
                def read():
                    for line in proc.stdout:
                        try:
                            messages.put(json.loads(line))
                        except ValueError:
                            pass
                    messages.put(None)
    
                threading.Thread(target=read, daemon=True).start()
    
                def send(message):
                    proc.stdin.write((json.dumps(message) + '\n').encode())
                    proc.stdin.flush()
    
                def response(request_id):
                    deadline = time.monotonic() + 15
                    while time.monotonic() < deadline:
                        msg = messages.get(timeout=max(.01, deadline - time.monotonic()))
                        if msg is None:
                            raise RuntimeError('backend exited before response')
                        if msg.get('id') == request_id:
                            if 'error' in msg:
                                raise RuntimeError(json.dumps(msg['error']))
                            return msg['result']
                    raise RuntimeError('response timeout')
    
                send({'id': 1, 'method': 'initialize', 'params': {
                    'clientInfo': {'name': 'synthetic-rollout-diagnostic', 'version': '1'},
                    'capabilities': {'experimentalApi': True}}})
                response(1)
                send({'method': 'initialized'})
                return proc, err, messages, send, response
    
            def stop(proc):
                if proc.poll() is None:
                    proc.terminate()
                    try:
                        proc.wait(timeout=3)
                    except subprocess.TimeoutExpired:
                        proc.kill()
                        proc.wait()
    
            proc, err, messages, send, response = launch()
            try:
                send({'id': 2, 'method': 'thread/start', 'params': {
                    'cwd': str(root), 'approvalPolicy': 'never', 'sandbox': 'read-only',
                    'config': {'features.memories': False, 'mcp_servers': {}}}})
                thread = response(2)['thread']
                thread_id = thread['id']
            finally:
                stop(proc)
                err.close()
            rollouts = list(home.glob('sessions/**/*.jsonl'))
            if len(rollouts) > 1:
                raise RuntimeError('expected at most one freshly generated rollout')
            if rollouts:
                rollout = rollouts[0]
            else:
                # thread/start can defer file creation until its first turn.
                rollout = home / 'sessions/2026/10/01' / (
                    'rollout-2026-10-01T00-00-00-' + thread_id + '.jsonl')
                rollout.parent.mkdir(parents=True)
                meta = {'timestamp': '2026-10-01T00:00:00.000Z', 'ordinal': 0, 'type': 'session_meta',
                        'payload': {'id': thread_id, 'timestamp': '2026-10-01T00:00:00Z',
                                    'cwd': str(root), 'originator': 'synthetic-rollout-diagnostic',
                                    'cli_version': result['version'].split()[-1],
                                    'source': 'cli', 'model_provider': 'openai',
                                    'history_mode': 'paginated'}}
                rollout.write_text(json.dumps(meta) + '\n')
            payload = {'message': '', 'replacement_history': [
                {'type': 'message', 'role': 'user', 'content': [
                    {'type': 'input_text', 'text': 'x' * (args.text_mib * 1024 * 1024)}]}]}
            record = {'timestamp': '2026-10-01T00:00:00.000Z', 'type': 'compacted',
                      'payload': payload}
            with rollout.open('ab') as f:
                for n in range(args.snapshots):
                    record['ordinal'] = n + 1
                    line = (json.dumps(record, separators=(',', ':')) + '\n').encode()
                    f.write(line)
            result['rollout_bytes'] = rollout.stat().st_size
            result['final_replacement_text_bytes'] = args.text_mib * 1024 * 1024
            proc, err, messages, send, response = launch()
            started = time.monotonic()
            samples = []
            result['outcome'] = 'timeout'
            try:
                send({'id': 2, 'method': 'thread/resume', 'params': {
                    'threadId': thread_id, 'excludeTurns': True, 'cwd': str(root),
                    'config': {'features.memories': False, 'mcp_servers': {}}}})
                while time.monotonic() - started < 20:
                    try:
                        status = dict(l.split(':', 1) for l in
                                      Path(f'/proc/{proc.pid}/status').read_text().splitlines()
                                      if ':' in l)
                        sample = {'elapsed_s': round(time.monotonic() - started, 4),
                                  'rss_kib': int(status.get('VmRSS', '0').split()[0]),
                                  'anon_kib': int(status.get('RssAnon', '0').split()[0])}
                        samples.append(sample)
                        if sample['rss_kib'] > args.stop_rss_mib * 1024:
                            result['outcome'] = 'stopped_at_rss_limit'
                            break
                    except (OSError, ValueError):
                        pass
                    try:
                        msg = messages.get(timeout=.01)
                    except queue.Empty:
                        continue
                    if msg is None:
                        result['outcome'] = 'backend_exited'
                        break
                    if msg.get('id') == 2:
                        result['outcome'] = 'error' if 'error' in msg else 'resume_completed'
                        result['response_error'] = msg.get('error')
                        break
            finally:
                result['elapsed_s'] = round(time.monotonic() - started, 4)
                stop(proc)
                result['backend_exit'] = proc.returncode
                err.seek(0)
                log = err.read().decode(errors='replace')
                result['loader_completed'] = 'Resumed rollout with' in log
                result['loader_parse_errors_zero'] = 'parse errors: 0' in log
                err.close()
            result['peak_rss_kib'] = max((s['rss_kib'] for s in samples), default=0)
            result['peak_anon_kib'] = max((s['anon_kib'] for s in samples), default=0)
            result['samples'] = samples
        if args.result:
            args.result.write_text(json.dumps(result, indent=2) + '\n')
        print(json.dumps({k: v for k, v in result.items() if k != 'samples'}, indent=2))
    
    
    if __name__ == '__main__':
        main()

    Specific fix and regression-test targets

    At the verified release commit:

    A bounded memory-extraction reader should skip irrelevant records before retaining them, apply byte/record budgets during reading, and handle oversized media/tool output before cloning and serialization. A defensive size rejection should cover decompressed input, rather than relying only on compressed file size. Resume replay also needs bounded handling of superseded snapshots while preserving thread semantics.

    The synthetic fixed-final-history cases provide a regression shape: memory growth should not track all superseded snapshots merely because the archive grew. No implementation fix has been applied or validated here.

    The bot suggested #51947, and #29510 describes the same general large-history failure family. This report adds a confirmed background memory-job trigger plus a synthetic reproduction on an exact official release; the precise caller in #51947 remains unproven.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    appIssues related to the Codex desktop appbugSomething isn't workingperformance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions