Repository navigation
[Linux Desktop 26.1007] Background memory extraction of inactive 7 GiB chat OOMs native backend; no chat opening required #53039
Description
Activity
- addedbugSomething isn't workingSomething isn't workingappIssues related to the Codex desktop appIssues related to the Codex desktop app
on Oct 11, 2026 github-actions commented
on Oct 11, 2026 on Oct 11, 2026 – with GitHub ActionsContributorMore actionsPotential duplicates detected. Please review them and close your issue if it is a duplicate.
Powered by Codex Action
- changed the title
[-][Linux Desktop 26.1007] Work session OOM-killed after 6h: app-owned scope peaks at 19.9 GiB RAM + 17.75 GiB swap[/-][+][Linux Desktop 26.1007] Oversized saved rollout OOMs native backend: 7 GiB history, 6 GiB RSS within 4s on metadata-only resume[/+]on Oct 11, 2026 - changed the title
[-][Linux Desktop 26.1007] Oversized saved rollout OOMs native backend: 7 GiB history, 6 GiB RSS within 4s on metadata-only resume[/-][+][Linux Desktop 26.1007] Background memory extraction of inactive 7 GiB chat OOMs native backend; no chat opening required[/+]on Oct 11, 2026 SpinGiantCRM commented
on Oct 11, 2026 AuthorMore actionsFurther investigation adds an exact release match, a credential-free synthetic reproduction, and a content-free breakdown of the real oversized history.
Exact binary/source match
The installed bundled backend is byte-for-byte identical to the official
rust-v0.162.0-alpha.17.2Linux x86_64 musl release. I downloaded the official release archive, verified its published SHA-256, decompressed it and compared the binary hash:Installed binary SHA-256: 73b73b048546993298234a295de01ef2424e7a8b29def9d0ecb1ca7cb77500d4 Official release binary SHA-256: 73b73b048546993298234a295de01ef2424e7a8b29def9d0ecb1ca7cb77500d4 Release tag commit: 740e5af33c71225640e0c1c1555c514c2c93ab74The installed loader logs also identify
rollout/src/recorder.rs:1108(load starts) and:1161(load completes), matching that release's source.What made the original history large
A streaming scan of the original 7,507,994,138-byte file, without exporting conversation contents, measured:
Record category Count Serialized bytes Custom tool outputs 11,933 5,165,329,265 Compaction snapshots 171 1,656,273,161 Other records 94,531 686,391,712 Embedded
data:image/...URLs account for 6,577,033,437 bytes (87.60%) across 4,868 occurrences. That includes 5,105,045,283 bytes in custom tool outputs and 1,453,847,178 bytes in compaction snapshots. Local hash equality checks found 2,066 unique image URLs and 1,543,511,104 bytes of exact repeated image URLs. Only counts and byte sizes are disclosed; no image data or hashes are posted.This is serialized image storage, not an estimate of decoded pixel memory. It does not establish a leak in screenshot capture. It explains why loading this persisted history is unusually costly.
Synthetic reproduction: no private history or database required
The script below creates a fresh temporary CODEX_HOME, starts a new thread without starting a model turn, and writes repeated valid paginated compaction records with ordinals. Every compaction's replacement history contains the same 1 MiB text; the final replacement history stays constant while the stored archive grows. It then cold-resumes the thread with
excludeTurns=true.Memories are disabled in the synthetic fixture. No auth files, original chats or original databases are used. Thus the synthetic test exercises the shared history/resume path; the separately recorded memory job establishes the real second shutdown's background caller.
Observed on the identical official release binary:
Saved snapshots Approximate file size Sampled peak backend RSS Outcome 1 1.00 MiB 342.42 MiB Resume completes, 0.025 s 32 32.01 MiB 494.03 MiB Resume completes, 0.175 s 128 128.02 MiB 791.89 MiB Resume completes, 0.599 s 256 256.05 MiB 1,058.05 MiB Resume completes, 1.167 s 512 ~512.10 MiB Kernel recorded 1,542,640 KiB anonymous RSS + 118,172 KiB file RSS for codexIsolated 2 GiB cgroup OOM All four successful cases logged loader completion with zero parse errors. For the 512 case the journal reported
CONSTRAINT_MEMCG, scopeResult=oom-kill,MemoryPeak=2147639296, andMemorySwapPeak=0; both the diagnostic Python supervisor andcodexwere killed. No sample/result file survived that case. The scope elapsed 2.739 s including fixture generation and both backend startups, not a measured resume-only duration.The 2 GiB scope figure includes file cache and supervisor memory; it is not presented as backend RSS. The normal desktop app and its original backend stayed running.
Run one case at a time on Linux with a hard diagnostic limit:
systemd-run --user --scope --unit=codex-rollout-repro-128 \ -p MemoryMax=2G -p MemorySwapMax=0 \ python3 reproduce-rollout-allocation.py /path/to/codex --snapshots 128
Change 128 to 512 and use a fresh unit name for the failing case. The script attempts an RSS-based early stop, but a fast allocation burst can hit the cgroup limit before that monitor responds. The hard limit is required for the failing case.
Self-contained Python reproduction
#!/usr/bin/env python3 """Offline synthetic history test; never accesses an existing CODEX_HOME. Example: python reproduce-rollout-allocation.py /path/to/codex --snapshots 128 Run under a 2 GiB cgroup limit. The RSS monitor attempts an early stop, but a fast allocation burst can reach the hard limit first (observed at 512). No model turn is started. All text and thread state are created from scratch. """ import argparse import hashlib import json import os from pathlib import Path import queue import subprocess import tempfile import threading import time def main(): ap = argparse.ArgumentParser() ap.add_argument('codex', type=Path) ap.add_argument('--snapshots', type=int, default=128) ap.add_argument('--text-mib', type=int, default=1) ap.add_argument('--stop-rss-mib', type=int, default=1200) ap.add_argument('--result', type=Path) args = ap.parse_args() binary = args.codex.resolve() if not 1 <= args.snapshots <= 512 or not 1 <= args.text_mib <= 4: ap.error('snapshots must be 1..512 and text-mib must be 1..4') result = {'snapshots': args.snapshots, 'text_mib': args.text_mib, 'version': subprocess.check_output([str(binary), '--version'], text=True).strip(), 'binary_sha256': hashlib.file_digest(binary.open('rb'), 'sha256').hexdigest()} with tempfile.TemporaryDirectory(prefix='codex-synthetic-rollout-') as td: root = Path(td) home = root / 'codex-home' home.mkdir() (home / 'config.toml').write_text('[features]\nmemories = false\n') env = {'PATH': os.environ.get('PATH', '/usr/bin:/bin'), 'HOME': str(root), 'CODEX_HOME': str(home), 'XDG_CONFIG_HOME': str(root / 'xdg'), 'RUST_LOG': 'codex_rollout=debug'} def launch(): err = tempfile.TemporaryFile() proc = subprocess.Popen([str(binary), 'app-server', '--stdio'], cwd=root, env=env, stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=err) messages = queue.Queue() def read(): for line in proc.stdout: try: messages.put(json.loads(line)) except ValueError: pass messages.put(None) threading.Thread(target=read, daemon=True).start() def send(message): proc.stdin.write((json.dumps(message) + '\n').encode()) proc.stdin.flush() def response(request_id): deadline = time.monotonic() + 15 while time.monotonic() < deadline: msg = messages.get(timeout=max(.01, deadline - time.monotonic())) if msg is None: raise RuntimeError('backend exited before response') if msg.get('id') == request_id: if 'error' in msg: raise RuntimeError(json.dumps(msg['error'])) return msg['result'] raise RuntimeError('response timeout') send({'id': 1, 'method': 'initialize', 'params': { 'clientInfo': {'name': 'synthetic-rollout-diagnostic', 'version': '1'}, 'capabilities': {'experimentalApi': True}}}) response(1) send({'method': 'initialized'}) return proc, err, messages, send, response def stop(proc): if proc.poll() is None: proc.terminate() try: proc.wait(timeout=3) except subprocess.TimeoutExpired: proc.kill() proc.wait() proc, err, messages, send, response = launch() try: send({'id': 2, 'method': 'thread/start', 'params': { 'cwd': str(root), 'approvalPolicy': 'never', 'sandbox': 'read-only', 'config': {'features.memories': False, 'mcp_servers': {}}}}) thread = response(2)['thread'] thread_id = thread['id'] finally: stop(proc) err.close() rollouts = list(home.glob('sessions/**/*.jsonl')) if len(rollouts) > 1: raise RuntimeError('expected at most one freshly generated rollout') if rollouts: rollout = rollouts[0] else: # thread/start can defer file creation until its first turn. rollout = home / 'sessions/2026/10/01' / ( 'rollout-2026-10-01T00-00-00-' + thread_id + '.jsonl') rollout.parent.mkdir(parents=True) meta = {'timestamp': '2026-10-01T00:00:00.000Z', 'ordinal': 0, 'type': 'session_meta', 'payload': {'id': thread_id, 'timestamp': '2026-10-01T00:00:00Z', 'cwd': str(root), 'originator': 'synthetic-rollout-diagnostic', 'cli_version': result['version'].split()[-1], 'source': 'cli', 'model_provider': 'openai', 'history_mode': 'paginated'}} rollout.write_text(json.dumps(meta) + '\n') payload = {'message': '', 'replacement_history': [ {'type': 'message', 'role': 'user', 'content': [ {'type': 'input_text', 'text': 'x' * (args.text_mib * 1024 * 1024)}]}]} record = {'timestamp': '2026-10-01T00:00:00.000Z', 'type': 'compacted', 'payload': payload} with rollout.open('ab') as f: for n in range(args.snapshots): record['ordinal'] = n + 1 line = (json.dumps(record, separators=(',', ':')) + '\n').encode() f.write(line) result['rollout_bytes'] = rollout.stat().st_size result['final_replacement_text_bytes'] = args.text_mib * 1024 * 1024 proc, err, messages, send, response = launch() started = time.monotonic() samples = [] result['outcome'] = 'timeout' try: send({'id': 2, 'method': 'thread/resume', 'params': { 'threadId': thread_id, 'excludeTurns': True, 'cwd': str(root), 'config': {'features.memories': False, 'mcp_servers': {}}}}) while time.monotonic() - started < 20: try: status = dict(l.split(':', 1) for l in Path(f'/proc/{proc.pid}/status').read_text().splitlines() if ':' in l) sample = {'elapsed_s': round(time.monotonic() - started, 4), 'rss_kib': int(status.get('VmRSS', '0').split()[0]), 'anon_kib': int(status.get('RssAnon', '0').split()[0])} samples.append(sample) if sample['rss_kib'] > args.stop_rss_mib * 1024: result['outcome'] = 'stopped_at_rss_limit' break except (OSError, ValueError): pass try: msg = messages.get(timeout=.01) except queue.Empty: continue if msg is None: result['outcome'] = 'backend_exited' break if msg.get('id') == 2: result['outcome'] = 'error' if 'error' in msg else 'resume_completed' result['response_error'] = msg.get('error') break finally: result['elapsed_s'] = round(time.monotonic() - started, 4) stop(proc) result['backend_exit'] = proc.returncode err.seek(0) log = err.read().decode(errors='replace') result['loader_completed'] = 'Resumed rollout with' in log result['loader_parse_errors_zero'] = 'parse errors: 0' in log err.close() result['peak_rss_kib'] = max((s['rss_kib'] for s in samples), default=0) result['peak_anon_kib'] = max((s['anon_kib'] for s in samples), default=0) result['samples'] = samples if args.result: args.result.write_text(json.dumps(result, indent=2) + '\n') print(json.dumps({k: v for k, v in result.items() if k != 'samples'}, indent=2)) if __name__ == '__main__': main()
Specific fix and regression-test targets
At the verified release commit:
- phase1::sample, line 267 loads the entire archive before filtering or applying the memory input token budget.
- load_rollout_items, lines 1105–1167 retains every decoded record in a vector.
- The V1 memory filter, line 397 discards compaction records after they have already been loaded. Both memory input versions discard these records.
- The memory output policy admits custom tool outputs; the sanitizer clones them. V1 then serializes the filtered collection. This is another possible amplification after the initial load, not a separately measured allocation stage in the actual crash.
A bounded memory-extraction reader should skip irrelevant records before retaining them, apply byte/record budgets during reading, and handle oversized media/tool output before cloning and serialization. A defensive size rejection should cover decompressed input, rather than relying only on compressed file size. Resume replay also needs bounded handling of superseded snapshots while preserving thread semantics.
The synthetic fixed-final-history cases provide a regression shape: memory growth should not track all superseded snapshots merely because the archive grew. No implementation fix has been applied or validated here.
The bot suggested #51947, and #29510 describes the same general large-history failure family. This report adds a confirmed background memory-job trigger plus a synthetic reproduction on an exact official release; the precise caller in #51947 remains unproven.
What version of the Codex App are you using (From “About Codex” dialog)?
ChatGPT-branded Linux desktop app
26.1007.21434(packagechatgpt-desktop-bin 26.1007.21434-1).Bundled backend:
codex-cli 0.162.0-alpha.17.2. Both verified locally. The bundled backend is byte-for-byte identical to the official release binary; its source commit is740e5af33c71225640e0c1c1555c514c2c93ab74.What subscription do you have?
Not collected.
What platform is your computer?
CachyOS Linux x86_64, kernel
7.2.9-2-cachyos-bore, KDE Plasma6.7.5/ Wayland, RTX 4080 SUPER / NVIDIA615.78.08. Approximately 30.5 GiB usable RAM and 30.5 GiB swap.What issue are you seeing?
The second shutdown was triggered by background memory extraction of an inactive, oversized chat. The user did not open that chat. The same underlying full-history loader also reproduces runaway allocation in an isolated backend.
The rollout is 7,507,994,138 bytes (6.99 GiB) with 106,635 JSONL records; its largest record is approximately 17.08 MiB. Large records include persisted compaction replacement histories. A content-free scan found 87.60% of the file's serialized bytes are embedded image URLs, including 1.54 GB of exact repeated URLs. The detailed breakdown is in the synthetic-reproduction follow-up. Contents and identifiers are withheld.
In both actual app failures, backend logs recorded
Resuming rollout from "<same oversized rollout>"shortly before the OOM kill:The second app instance had been running only 3 minutes 52 seconds. The user had two ordinary chats and one Work chat active and explicitly confirms they did not open the oversized old chat.
Background trigger evidence: a read-only query of the local
memories_1.sqlitejob table found the oversized chat's job with:kind = memory_stage1status = runningstarted_at = 1791696400(2026-10-11 05:26:40 UTC / 15:56:40 Adelaide)finished_at = NULLlease_until = 1791700000(one hour later)The job start matches the oversized rollout's backend load log to the second. Another small chat's memory job was claimed and loaded concurrently. This establishes background memory extraction as the second incident's trigger; it does not require navigation to the oversized chat. The first incident loaded the same oversized history, but its earlier job claim has been overwritten, so the exact first caller remains unverified.
The log text
Resuming rollout from ...is emitted by the genericload_rollout_itemsloader and does not prove a user-facingthread/resumerequest. Earlier framing of this as opening/resuming the old chat was too narrow.Local config has
features.memories = true,memories.generate_memories = true, andmemories.use_memories = true.What steps can reproduce the bug?
A self-contained synthetic reproduction is now available in this follow-up. It needs no private history or database. With a constant 1 MiB final replacement history, increasing saved snapshots from 1 to 256 raised sampled backend RSS from 342 MiB to 1,058 MiB. A ~512 MiB synthetic archive exhausted an isolated 2 GiB diagnostic cgroup. The follow-up includes the full script, successful parse checks, kernel evidence, exact release verification, and fix/test targets.
The original private-history reproduction follows:
A controlled local reproduction used the installed bundled backend, a separate CODEX_HOME, an independent copy of the rollout and a filtered copy of the local thread metadata. The live chat/database was not changed. No turn was started and no authentication files were copied.
codex app-server --stdioinside a separate systemd scope withMemoryMax=6GandMemorySwapMax=0.{"id":2,"method":"thread/resume","params":{"threadId":"<fixture-thread-id>","excludeTurns":true,"cwd":"<isolated-fixture-directory>","config":{"mcp_servers":{}}}}Results with the 6.99 GiB rollout:
CONSTRAINT_MEMCG; victim: codex; anonymous RSS at termination: 6,269,660 KiB.Result=oom-kill; the normal desktop app remained running.Control: the same metadata-only resume procedure with a different 28,905,963-byte rollout completed in 0.42 seconds, with sampled peak RSS 405,472 KiB and exit code 0.
Thus
excludeTurns=truedoes not protect this installed backend from the oversized rollout load.What is the expected behavior?
Background memory extraction and chat resume must not materialize an unbounded archive into backend memory. Extraction needs streaming/filtering and a bounded input budget before retaining records. Oversized inputs should be skipped with a recoverable diagnostic. A metadata-only resume should also avoid system-wide memory exhaustion.
Additional information
Actual watchdog evidence:
MemoryPeak=16026378240andMemorySwapPeak=15437815808(approximately 14.93 / 14.38 GiB)./usr/lib/chatgpt/ChatGPT; after relaunch the equivalent scope also contains the app-server/helpers.During subsequent bounded sampling, the normal app-server stayed approximately 368–463 MiB RSS, and process counts fell 32 → 20 as helpers exited. The isolated failing test directly identifies the native backend as an allocation owner.
The verified release's source is consistent with the measured behavior: RolloutRecorder::load_rollout_items parses rollout records and retains them in a vector; get_rollout_history uses that loader. The installed binary hash matches the official release built from this tagged source. This is source inspection, not a captured allocation backtrace.
The verified release's source also identifies the background path: memory startup asynchronously launches phase 1 for eligible root sessions; phase1::sample first calls
RolloutRecorder::load_rollout_itemsand only afterwards filters or applies the memory input token budget. The pipeline documentation confirms that eligible idle chats are processed in the background. These references are pinned to the official release commit matching the installed backend.This closely matches #29510, which already describes huge rollout histories causing app-server growth to tens of GB. Please investigate the continuing failure on this newer backend, particularly automatic memory extraction of inactive chats and cold resume with
excludeTurns=true.The original report also noted repeated access-denied hydration requests and IAB errors. Those observations remain valid, but the new reproduction supplies a much stronger cause; their causal role has not been established.
The saved chat and its original history remain intact. Public evidence contains only metrics, record counts and sanitized log signatures.