Skip to content

radlink: drain shared barrier workers before reusing pass state (follow-up to #842) - #937

Open
honkstar1 wants to merge 112 commits into
EpicGames:devfrom
honkstar1:codex/shared-pool-barrier-drain-pr-20260904
Open

radlink: drain shared barrier workers before reusing pass state (follow-up to #842)#937
honkstar1 wants to merge 112 commits into
EpicGames:devfrom
honkstar1:codex/shared-pool-barrier-drain-pr-20260904

Conversation

@honkstar1

@honkstar1 honkstar1 commented Sep 4, 2026

Copy link
Copy Markdown

September 8 refresh

This block supersedes older base-revision and validation notes below.

The incremental feature change remains three files: the barrier-worker drain and its regression coverage. Compressed-object, memory-cleanup and separate #939 changes were not folded into this PR.

Older text is retained below for historical context. Already-integrated upstream changes are behind the refreshed merge base rather than new standalone fixes in this refresh.


Scope and review order

One-commit correctness follow-up to #842. Review the incremental change here:

Diff against #842: 1 commit, 3 files, +27/-9.

This fixes the shared governor/barrier implementation introduced by #842. It does not depend on the compressed-object or later optimization PRs. GitHub targets EpicGames:dev because the dependency branch is in a fork, so the default Files changed view also includes #842. Please apply this follow-up after #842, or fold commit 7000a421 into that PR. #842 itself is unchanged.

Current merge status (September 4): GitHub reports conflicts against dev@706abca6. A local merge-tree check of both #842's head and this head reports the same ten conflicting files. The dependency stack needs an upstream refresh before direct merge; this PR intentionally keeps that separate from the one-commit correctness fix. No upstream CI result is claimed for this conflicted stack.

Problem

For a shared barrier pass, the last task posted main_semaphore before its background worker left tp_run_tasks_shared. The dispatcher could consume that notification and return while a worker was still about to read/decrement the task queue. A subsequent pass could then reset the queue or change the pass kind underneath that worker.

Ordinary governor dispatch already drains its outstanding grants before returning; task completion alone did not provide the equivalent guarantee for barrier dispatch.

Fix

  • Initialize an atomic countdown for the barrier pass's woken background workers.
  • Each worker acknowledges completion after leaving the task loop; the last worker posts the existing completion semaphore.
  • Main leaves its own task loop and waits for that acknowledgement before returning or ending the cohort.
  • Extend shared_pool_governor_repeated_passes to exercise four consecutive barrier passes per bracket, followed by ordinary governor dispatch, across 96 iterations.

No new semaphore, thread, polling loop or allocation. Normal governor scheduling, the cross-process budget protocol, cleanup options and stall assertions are unchanged. The added work is one counter initialization and one atomic decrement per woken worker per barrier pass. Waiting for a delayed worker is intentional: its previous early handoff was unsafe.

Validation

Fresh builds of this exact PR head (7000a421), using LLVM/Clang/LLD 20.1.8:

Configuration Passed Failed / crashed Excluded
Release 158 0 / 0 12
Debug 158 0 / 0 12

The 12 exclusions are existing MSVC-invoking cases; no MSVC compiler/linker/assembler was used. These are local supported-subset results, not a claim that the full upstream CI matrix passes.

Additional validation of the identical patch in combined development commit 348ae713:

  • An investigation-only scheduling hook paused the last background task before its next queue read. The old dispatcher returned while the worker was paused (probe failed); the fixed dispatcher waited (probe passed). The hook is not in this PR or the supplied binary. The committed test provides repeated-transition coverage without timing thresholds.
  • Release and Debug: 163 passed, 0 failed/crashed, 12 excluded in each.
  • A real shader-module link with compressed objects passed with one-shot cleanup both enabled and disabled; DLL/PDB bytes matched the previous binary's outputs.
  • Small Windows A/B test, 12 interleaved links per version, four workers, isolated shared pool, identical flags: median process lifetime 151 ms before / 149 ms after; observed peak working set ~71.6 MiB and process commit ~222.6 MiB for both. This detects no regression in that small test, not a speedup or a large-link/Wine performance guarantee.

Investigation was prompted by rare 0xc000001d failures under Wine, but there is no dump/fault address and this has not been established as their cause. No Linux/Wine/UBA or large-link soak validation has been performed for this fix. This PR corrects the independently reproduced handoff race; it does not disable assertions, extend timeouts, or claim to resolve every farm failure.

NikitaSmith057 and others added 30 commits August 28, 2026 10:45
codec produces byte identical binaries for the tests
Assembly code is >25% faster than intrinsic implementations on msvc.
The image buffer was push'd on the shared link arena and only reclaimed in the
single-threaded process rundown at exit -- a multi-second kernel page-reclaim tail
(observed: one thread 100% in-kernel, zero user frames).

Allocate it as a standalone reserve_memory/commit_memory region and
release_memory() it the instant the background image-write thread joins (image is
on disk, no later reader). VirtualFree(MEM_RELEASE) returns fast; the kernel zeroes
the ~1GB on its background thread, overlapping the parallel input-view release +
exit instead of blocking rundown. Discard early so the kernel cleans up while the
app still runs -- don't defer to exit.

Gated 65/65 linker torture (determ_test + p2r_determinism: image correct).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit e7a0e5e)
…before PDB emit (opt-in, default off)

After type merging, prune any merged TPI/IPI record not transitively reachable
from a surviving symbol. Roots are the type indices referenced by the symbols
that survive /OPT:REF (plus inlinee call-site types); the type graph is then
closed over and everything unreached is dropped before the streams are written.
Runs only under /OPT:REF -- it is the debug-info analogue of dead-section
stripping, and is otherwise transparent (no visible type is removed).

Implementation notes:
  - parallel transitive closure (bulk-synchronous rounds, atomic mark/expand)
  - fwdref<->definition pairing via a per-unique-name ring so a live forward
    reference keeps its definition (and vice-versa)
  - compaction is in place with the remap kept in scratch, so peak memory is
    unchanged

Numbers (UnrealEditorFortnite-Engine.dll, /OPT:REF /OPT:ICF, hashing NONE):
  PDB  5315 -> 5081 MB   (type-GC alone: -234 MB)

Default OFF: shrinks the PDB but a pruned type can't be cast-to in the debugger watch window.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 024f1ed)
The closure re-scanned every merged type each round (O(rounds * total types)) to
find marked-but-unexpanded leaves. In a full-link trace of UnrealEditorFortnite
that round-rescan dominated the type-GC: lnk_gc_expand_task ~13.3 s of CPU.

Replace it with a frontier worklist: the atomic mark now gates a single append per
leaf, and each round expands only the slice newly marked by the previous round, so
total work is O(reachable types) instead of O(rounds * total types). Drops the
per-round `expanded` bitmap and full-array sweeps.

Output is unchanged -- same reachable set, PDB byte-identical (5081 MB on the same
input), and debugger-fidelity checks (addr->symbol 100%, core types resolve) match
the pre-change build.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 58d0c27)
The frontier mark did an interlocked op on every reference edge. Add a plain
non-atomic check first (the mark bit only ever goes 0->1, so a stale "already
set" read is safe), so the interlocked op runs once per leaf at its 0->1
transition instead of once per edge. Output identical (PDB 5081 MB).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 9872144)
Resolve MSVC C++ header-unit IFC debug records (LF_IFC_RECORD 0x1522) by
merging the .ifc .msvc.trait.debug-records CodeView stream and redirecting
each record to its real type -> fixes VS debugger AV stepping into header-
unit code (BAD 921->0, 44/44 MSVC name-match, live debug verified).

(cherry picked from commit 4b7f8bf, ICF leader-keying half dropped: superseded
by upstream d0d8b59's own /OPT:ICF)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The IFC apply phase merged the per-worker nonblob_complete sets and parsed
.ifc inputs serially on the main thread. Run the set merge and the ifc
parse as pool tasks. Applied record set and output bytes unchanged (gated
byte-identical).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 86c48d3)

radlink: header-unit IFC debug-record resolution

Resolve MSVC C++ header-unit IFC debug records (LF_IFC_RECORD 0x1522) by
merging the .ifc .msvc.trait.debug-records CodeView stream and redirecting
each record to its real type -> fixes VS debugger AV stepping into header-
unit code (BAD 921->0, 44/44 MSVC name-match, live debug verified).

(cherry picked from commit 4b7f8bf, ICF leader-keying half dropped: superseded
by upstream d0d8b59's own /OPT:ICF)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

radlink: skip redundant library re-searches in the resolution fixpoint

lnk_link_inputs resolves libraries to a fixpoint: an outer pass loops over
every library, and each library is re-searched (a full tp_for_parallel over
all workers, scanning every undefined/weak symbol in search_chunks) once per
drained input batch until nothing new resolves. On the UE editor link that is
~2733 dispatches, each waking and joining ~60 workers -- and the phase is
barrier-bound, so that wake/join is the cost, not the scan.

Most re-searches are redundant: search_chunks only grows during the loop
(symbols are never removed until the end) and member-queue dedup is idempotent,
so a re-search can only queue new members if the undefined/weak symbol set grew
or anti-dep searching was just enabled since this library was last searched.
Stamp each LNK_Lib with the search_chunks symbol count + anti-dep mode at its
last search and skip the dispatch when neither changed.

~24% fewer dispatches (2733 -> ~2089) and ~0.2s of wake/join wall-time removed.
Output is byte-identical and reproducible (which dispatch coalesces is timing
dependent, but a skipped one provably queued nothing, so the result is
unchanged -- verified relink-twice byte-identical across many runs).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 3054138)

radlink: cache symbol interp so the lib search stops re-parsing resolved symbols

(squashed with: exact per-obj bitset filter for ifc_redirect_hm lookups; parallelize the IFC apply discovery/read+parse/resolution -- the three were one logical change split by a rebase)
Several O(input) per-obj setup sweeps at the head of
lnk_make_code_view_input ran serially on the main thread while the pool
idled between the parallel parse phases. Dispatch them over the pool;
results land in per-obj slots, so downstream iteration order and output
bytes are unchanged (gated byte-identical).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 69672ab)
After merge-types reaches the per-thread scratch high-water (~9GB of tctx
scratch arenas stay committed but idle through the PDB peak), release the
committed-but-unused scratch pages back to the OS before the PDB build
re-grows them. Drops recorded peak working set on the monolithic
UnrealEditorFortnite-Engine.dll link.

- arena_decommit_unused(): decommit committed pages strictly above each
  block's live pos in the active chain, and the unused bodies of free-list
  blocks (keeping the header page). Reservation kept; push path re-commits
  on demand, so reuse is transparent and output byte-identical.
- tctx_scratch_decommit(): decommit the calling thread's two equipped
  scratch arenas.
- lnk_scratch_decommit_worker + tp_for_parallel(worker_count) with an
  in-task barrier: every worker (worker 0 IS the main thread) decommits its
  own scratch exactly once between lnk_merge_types and lnk_build_pdb.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit b1d8a4a)
…it peak

The per-source leaf-dedup probe tables (bucket_arr) are dead once the
unique-leaf set has been extracted, but their pages stayed committed through
the merge-types commit peak. Release them as soon as extraction is done so
they come off the peak footprint. Memory-release timing only; output bytes
unchanged (gated byte-identical).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit c03287a)
…ostic, byte-neutral)

When RADLINK_PHASE_LOG is set, lnk_log_timers writes machine-parseable raw
per-phase microseconds (Image/PDB/RDI/Lib/Debug + TOTAL) to that path, for
automated perf A/B. Env-unset -> identical code path, DLL/PDB byte-identical.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit f491549)
…governor + fair-share cohorts)

Share one worker budget across concurrent radlink processes: a cross-process
governor grants budget slots fairly (fair-share cohorts for barrier passes),
capping total running workers near core count during farm convoys instead of
workers x links oversubscription. The pool is dual-path: without the switch
the upstream barrier implementation runs verbatim, so default behavior is
unchanged. Soak-validated at 3216 concurrent shared-pool links with zero
hangs; output byte-identical in both modes.

Squashed from 2 series commits; their original messages follow.

* radlink: cross-process shared thread pool (governor + fair-share barrier passes)

Collapsed net of our 5-commit pool series (9fd1db9 -> d28970d revert ->
7434720 reapply -> e373f03 fair-share -> 35703bd grant-race fix), reconciled
onto upstream's reworked thread pool (9b63131 "ditch semaphores in favor of
barriers").

Upstream rewrote tp_run_tasks to bracket each pass with an entry+exit
barrier_wait(pool->barrier) sized to the full worker_count, and dropped
main_semaphore/task_semaphore. That model is incompatible with our shared
governor, where a path-A pass wakes only a governor-chosen SUBSET of workers
(main always runs, woken workers return their budget slot on drain): a
full-width entry barrier would block forever waiting for parked workers. We
therefore keep our semaphore-completion tp_run_tasks (no built-in barrier) which
already implements BOTH modes correctly, and discard upstream's barrier-bracket
internals. This is transparent to callers -- tp_broadcast/tp_sum_u64/barrier_wait
semantics are preserved, and barrier passes rendezvous a cohort-sized barrier.

Net design (= e373f03 + 35703bd final):
 - per-process governor thread + NAMED budget_semaphore (machine core budget)
 - parked workers woken one-per-granted-slot; each returns its slot on drain
 - path A (tp_for_parallel): main always runs, governor opportunistically
   recruits workers as budget frees; granted++ is published BEFORE the
   pass_active re-check (no grant-race hang), aborting with granted-- if the
   pass already ended
 - path B (tp_for_parallel_reserve + tp_barrier_begin/end): fair-share cohort =
   1 + budget slots free RIGHT NOW (never amasses the machine, deadlock-free);
   output is width-independent so any cohort is byte-identical to full width
 - off-by-default: non-shared mode is the prior semaphore pool, zero overhead

Carries the base_threads helpers the pool needs (semaphore_drop_n,
semaphore_drop_if_room, semaphore_take_n on win32 + linux); upstream's
semaphore_drop_count/semaphore_drop shim are kept untouched alongside.

Caller conversions (lnk.c, lnk_debug_info.c): the ICF-refine region, /OPT:REF
mark, scratch-decommit, p2r, rrt type-data, and PDB move-globals/write-modules
barrier passes now use tp_for_parallel_reserve inside a tp_barrier_begin/end
cohort bracket; lnk_icf_refine_region_task indexes per-lane scratch by task_id
(the contiguous lane), and PDB obj indices are redistributed per pass to the
pinned cohort C via lnk_build_pdb_distribute_obj_indices.

* radlink: dual-path thread pool (non-shared=upstream barrier, shared=governor)

Reconcile the thread pool into two independent code paths selected by
pool->is_shared (== /RAD_SHARED_THREAD_POOL name.size>0):

  NON-SHARED: restore origin/dev's barrier implementation VERBATIM. Workers
  park on a kernel barrier between passes (zero-syscall steady state).
  tp_run_tasks brackets the work loop in barrier_wait at entry+exit;
  tp_worker_main loops calling it; tp_for_parallel just inits state and
  joins as worker 0. No governor thread, no budget/wake/governor/main
  semaphores allocated -- zero overhead, zero new threads vs upstream.

  SHARED: keep OUR cross-process governor. Parked workers woken one-per-grant
  by a per-process governor thread borrowing a NAMED global budget semaphore;
  each woken path-A worker returns its slot on drain; granted++ published
  before the pass_active re-check (grant-race fix); main always runs;
  main_semaphore completion. Path-B barrier passes use the fair-share cohort
  (tp_for_parallel_reserve + tp_barrier_begin/end: cohort = 1 + budget-free-now,
  pinned, width-independent), NOT a full-width barrier.

Dispatch branches on is_shared: tp_for_parallel runs the upstream barrier
dispatch for non-shared and the governor dispatch for shared; the work loop
splits into tp_run_tasks (upstream barrier, non-shared) and
tp_run_tasks_shared (main_semaphore completion, shared). tp_alloc allocates
governor/budget state and launches the governor thread only when is_shared.
tp_barrier_begin/end and tp_for_parallel_reserve are no-ops in non-shared
mode (degrade to the upstream full-width barrier pass), so the lnk.c /
lnk_debug_info.c cohort callers stay transparent there.

Validated: build clean. Non-shared output content-identical to the d566da3
baseline (soak_big input set; only the irreducible /Brepro debug-GUID + PE
stamp/checksum bytes differ, equal to the same-binary control floor); no
governor thread spawned in non-shared. Shared re-soak 3216 concurrent
shared-pool links (48-way, mixed big/small, 64-core budget), zero hangs;
shared output (64-way and 8-way) content-identical to non-shared
(width-independent).

Pool name is copied into the config arena at parse (rsp-backed string lifetime).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ites

The reloc patcher (lnk_obj_reloc_patcher) walked each section's relocs in
on-disk COFF table order, scattering RMW patch writes randomly across the
~739MB image buffer (no HW prefetch, poor write locality). Sort each section's
relocs by apply_off (orig_idx tiebreak for a deterministic total order) into a
per-section scratch copy before patching, so the write stream is monotone
forward. Per-section relocs are independent and each writes its own disjoint
field, so final bytes are unchanged.

byte-identical: soak_big DLL+PDB diff == base-vs-base control floor (34/18,
GUID/timestamp/checksum churn only; no .text/.data/reloc-region diffs).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 57484d8)
The ~1GB image buffer is filled (align/zero fill + per-contrib body copies),
then streamed straight to disk; the filling threads never re-read it. Normal
stores pollute L2/L3 with ~1GB of write-once data. Use SSE2 non-temporal
stores (_mm_stream_si128, baseline on x86-64 -- no -mavx needed) for the large
(>=256B) fills via lnk_stream_set / lnk_stream_copy, with a scalar head/tail and
a MemoryCopy/MemorySet fallback for small or unaligned spans (identical bytes
either way). _mm_sfence() after the align-fill and at the end of each fill
worker makes the NT stores globally visible before the reloc-patch /
checksum passes read the image back.

byte-identical: soak_big DLL+PDB diff == base-vs-base control floor (34/18,
GUID/timestamp/checksum churn only; identical diff offsets, no image-content
diffs).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit b735928)
ICF refine region: chunk_nc/chunk_kp/chunk_sc/chunk_max[worker_count] are
written per-lane as chunk_X[wid] (24B/64B stride); adjacent lanes shared a
cache line on every per-round scan/radix reduction. Stride each lane entry to
a full 64B line (LNK_ICF_CL_STRIDE_U64, index wid*stride) so each lane owns
its line; chunk_max's [0..2] round-total scratch stays inside lane-0's line.

Symbol table: pad LNK_SymbolHashTrieChunkList to 64B so symtab->chunks /
search_chunks (per-worker arrays indexed [worker_id]) no longer false-share
on the parallel insert.

Pure layout change (more memory, no logic change); indexed values identical.
byte-identical, soak_big cmp 0 (DLL 0 / PDB 0).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 79c98e3)
LNK_SymbolHashTrie descent (insert + search) dereferenced name -> String8 ->
name bytes (2-3 cache-line misses) per level just to call str8_match. Store
the full key hash in the node at insert and fast-reject on hash mismatch
before touching the name string; str8_match still gates the real match, so
hash is fast-reject only. Same trie topology / contents / match results.

byte-identical, soak_big cmp 0 (DLL 0 / PDB 0).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 85cf482)
The multi-GB LEAF_BUCKETS probe-table arena release at the end of type
dedup cost ~370-950ms of SERIAL main-thread kernel time
(MiDeleteVaDirect/MiDecommitFreePage under VirtualFree(MEM_RELEASE)),
sitting on the critical path between dedup and the type-index fixup
passes. Hand the arena to a background reaper thread instead; the main
thread proceeds into the fixups while the kernel tears the pages down
concurrently.

Rejected alternative (measured): chunked MEM_DECOMMIT across the thread
pool does not help -- decommit serializes in the kernel on the process
address-space lock (~14 GB/s aggregate regardless of thread count;
8 GiB took 860 ms chunked-parallel vs 371 ms serial release). Same
reason the per-worker scratch-decommit pass before the PDB build stays
as-is: redistributing it is not faster, and backgrounding it is unsafe
because workers re-push into those scratch arenas during the PDB build
(documented in lnk_scratch_decommit_worker).

Reaper protocol: all pointers into the arena are dropped before launch
(bucket_arr slots zeroed), at most one reaper is in flight (joined
before relaunch and at end of link next to the image-write join), and
the handle is not detached at launch (thread_detach right after
thread_launch frees the W32_Entity the entry point is about to read).

Editor-scale (UnrealEditorFortnite-Engine.dll, /OPT:ICFSTATIC /BREPRO
/RAD_WORKERS:64):
- main-thread MiDeleteVaDirect  909 -> 163 ms (Superluminal, whole link)
- main-thread NtFreeVirtualMemory incl 1219 -> 553 ms
- lnk_merge_types main-thread incl 6.48 -> 5.10 s
- Debug phase 13.88/13.92 -> 12.97/13.63 s (RADLINK_PHASE_LOG, 2 warm
  runs each); reaper carries the 947 ms release off-thread
- PeakWorkingSet64 unchanged (49.91 GB both) -- peak is during dedup
  while the buckets are still live, so the delayed release is past it
- byte-identity: DLL base-vs-head diff == control (18B /Brepro
  timestamp+GUID band), OTHER=0; PDB streams identical in count, size,
  and name (MSF page placement shuffles, as it already does run-to-run)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit ca8c7f1)
… pages clean)

Debug sections (.debug$*) are not copied into the image, so lnk_obj_reloc_patcher
used to commit relocations directly into the memory-mapped input (FILE_MAP_COPY),
dirtying copy-on-write pages the kernel then has to reclaim at exit. Instead, copy
each reloc-patched debug section of debug-info-surviving objs into private memory
and patch the copy; readers fetch section bytes through lnk_obj_get_sect_data,
which prefers the copy. Objs excluded from debug info never reach the PDB/RDI
path, so their debug relocs are skipped entirely instead of patched pointlessly.
Sections without relocs are consumed straight from the (clean) input view.

Same reloc math, same phase order; GSI dedup and PDB module copies see identical
post-fixup bytes, so DLL and PDB outputs are byte-identical by construction.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 074e5c2)
…t the source $$S

The stripped-PDB pass zeroed CV_SymProc32.itype in the SOURCE symbol record and
then copied it out, mutating $$S bytes (and dirtying their backing pages) after
the full PDB was already built. Write the record first and zero the itype field
in the destination copy instead; the source stays untouched. Stripped-PDB output
bytes are unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit acc3929)
…mate

On dup-heavy input (PCH/type-server fan-out) the unique leaf count is a
small fraction of the total, and sizing the dedup probe tables from the
total leaf count oversizes them 5-10x, costing tens of seconds of
demand-zero page faults on 64B-apart random probes.

Estimate the distinct-hash count per CV_TypeIndexSource from the already
produced debug_h hashes (presence bitmap + linear counting) and size the
tables from that, clamped to the old total-based caps. The bitmaps are
filled with commutative atomic ORs over deterministic input hashes, so
the estimate -- and therefore the caps -- are identical run to run.

Winner selection per hash class (min lnk_leaf_ref_compare) and the sorted
extraction are cap-independent: only probe sequences change, not probe
logic, so output bytes are unchanged.

Overflow safety: the table-full case (probe wraps without a slot, which
previously only fired a debug Assert) now sets a shared flag; the passes
bail out early and the whole dedup is redone once with the always
sufficient total-based caps. Whether that happens is a pure function of
the input, so the retry path is deterministic too.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit ffe2aa1)
lnk_move_global_symbols_to_gsi hashed globals, proc refs, and publics in
parallel but then funneled every gsi_push_ through a task 0 serial loop
while the other workers parked on the barrier.

Shard the inserts instead: worker i owns buckets [i*B/W, (i+1)*B/W) and
walks the full symbol sequence in global order, inserting only symbols
whose hash % bucket_count lands in its range. Each bucket has exactly one
owner and receives its inserts in global sequence order, so per-chain
order -- which is serialized into the PDB -- is byte-identical to the
serial loop, for any worker count. No locks or atomics; gsi->symbol_count
is bumped once by task 0 behind a barrier.

The publics walk previously consumed the per-worker lists in place
(clearing node->next during iteration), which would race with concurrent
shard walkers; flatten the lists into a global-order node/hash array
first, then shard-insert from that.

Barrier passes walk the fixed lane partition strided by the pinned cohort so no lane is dropped under /RAD_SHARED_THREAD_POOL fair-share.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit c820b64)
…escriptor tables

cv_get_leaf/symbol_type_index_offsets built an arena-pushed CV_TypeIndexInfo
node per type index per record, and hot loops (leaf hashing, TI fixup, type
GC) pointer-chased those lists; the two functions showed ~13s combined
exclusive CPU across workers on editor-scale links.

New CV_TiOffsets view returns, with zero allocation for the common kinds:
- fixed-shape leaves/symbols: pointer into a static per-kind {source,offset}
  table (POINTER picks between two static variants off attribs)
- count-stride kinds (ARGLIST, SUBSTR_LIST, BUILDINFO, VFTPATH,
  CALLERS/CALLEES/INLINEES): inline {run_base, run_count} descriptor,
  offset(i) = run_base + i*4
- member-walk kinds (FIELDLIST/METHODLIST/inlinee lines) keep their walk and
  materialize a flat CV_TiOff array (doubling growth, footprint comparable
  to the old per-node list)

Emission order is preserved exactly (incl. FUNC_ID IPI-before-TPI and
UDT_SRC_LINE TPI-then-IPI asymmetries); the blake3 leaf-hash stream and all
consumers see identical offset sequences, so output bytes are unchanged.

All hot consumers (lnk_hash_cv_leaf(_deep), lnk_fixup_cv_type_indices,
lnk_gc_visit_offsets/expand, ifc closure, pdb_builder TI patch) now iterate
the flat view; legacy list API kept as a thin shim over the new one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit b36f595)
The align-byte fill is the first touch of the freshly committed ~751MB image
buffer (every page a demand-zero fault) and ran serially on the main thread
while all workers parked.

Range-split the fill into a flat (dst, byte, size) task list and dispatch it
via tp_for_parallel: sections are cut into ~4MB chunks with PAGE-ALIGNED split
points inside the page-aligned image reservation, so no two workers ever touch
the same 4K page. Each task keeps the NT-store lnk_stream_set fill and sfences
its own stores before signalling completion, so everything is globally visible
after the join (replacing the single main-thread sfence).

Measured on a FN editor DLL link (751MB image, 64 workers), via a throwaway
timer around the fill block: serial 24.5-29.1ms -> parallel 4.7-5.0ms, i.e.
about -20ms wall. Smaller than the profile-estimated 0.3-0.8s: Windows fault
clustering + NT-store bandwidth make the serial first-touch much cheaper here
than the estimate assumed.

Writes are value-identical to the serial loop and byte-disjoint -> output is
byte-identical by construction (verified: base-vs-head diff is the PE checksum
byte + debug-directory GUID/timestamp band only, matching the same-binary
control).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 51af745)
Bump the commit quantum from the 64KB default to 2MB for the three arena
families that take essentially all of a 50GB editor link's commit churn:

  - thread-pool per-worker arenas (tp_arena_alloc)
  - per-thread scratch arenas (tctx_alloc)
  - the linker's huge debug-info arena (lnk_get_huge_arena)

via their arena_alloc params only -- the global 64KB default is untouched, so
small arenas do not bloat. arena_decommit_unused is page-granular and tracks
cmt exactly, so it stays correct with the larger quantum.

Measured on a FN editor DLL link (751MB image, ~50.7GB committed, 64 workers),
via a throwaway atomic counter in commit_memory:

  VirtualAlloc(MEM_COMMIT) calls:  80.5-81.1K -> 8.3K   (-90%)
  committed bytes:                 50.7GB     -> 51.3GB (+0.6GB quantum slack)
  peak working set:                56.4GB     -> 56.8GB (+0.3-0.5GB)
  kernel time (GetProcessTimes):   no detectable movement (210-282s run noise
                                   swamps it; faults, not commit syscalls,
                                   dominate kernel time on this link)

So this is a syscall/address-space-lock relief change, not a measurable
wall/kernel win on this workload; kept because the slack cost is bounded and
small relative to the 56GB peak.

Commit sizes only affect when pages are committed, never what is written ->
output is byte-identical (verified: base-vs-head diff is the PE checksum byte
+ debug-directory GUID/timestamp band only, same as the same-binary control).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 11ac6b1)
The pre-dedup unique-leaf estimator swept every debug_h hash into a presence
bitmap (~6.6s of worker thread-time on editor-scale links). Sample every 8th
leaf POSITION per obj instead: position-based sampling is a pure function of
the input (schedule-independent), the presence bitmap shrinks 8x (cheaper
cache footprint per update), and the sampled distinct count is scaled back up
by a calibrated factor before the existing 1.9x safety + Min(fallback-cap)
clamp. SCALE=5.0 satisfies SCALE*1.9 >= K=8, covering the worst-case
sampled-to-true ratio for every duplication pattern, so the deterministic
overflow-retry stays off; an undershoot would still be caught by that retry
(exercised live during calibration at SCALE=1).

Measured on the FN editor-scale link: estimate block 182.6 -> ~38 ms wall,
caps byte-for-byte identical to the unsampled estimator (64M TPI / 16M IPI,
load factors 0.351 / 0.387), output byte-identical.

Also logs the estimate/caps/load factors under /RAD_LOG:TIMERS (output-neutral).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 857a941)
lnk_build_pdb_distribute_obj_indices round-robined objs across lanes by index,
ignoring per-obj debug$S size, so a lane drawing several giant objs held the
barrier pass at its final barrier while other lanes idled.

Distribute by greedy LPT instead: objs taken in weight-descending order
(obj_idx tie-break), each assigned to the least-loaded lane. Weights are O(1)
per obj -- symbols-subsection total_size for the GSI pass, total debug$S size
for the module-write pass. The partition is output-neutral: per-obj results
land in per-obj module streams or in GSI bucket chains that are content-sorted
at serialization (gsi_symbol_is_before), so any deterministic assignment
produces byte-identical PDB bytes.

Measured on the FN editor-scale link (6 quiet interleaved samples): Write
Modules wall 174-191 -> 151-168 ms (~-12%, samples fully separated);
Move Global Symbols wall unchanged (~352 ms) -- its heavy sub-phases are
partitioned by symbol_input_ranges/symtab chunks, not obj_indices.

Also logs both phase walls under /RAD_LOG:TIMERS (output-neutral).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 5690f99)
The per-bucket GSI sort tie-breaks same-name globals on CV_Symbol.offset,
which was the compacted deduper slot index -- CAS-arrival order in
cv_symbol_deduper_insert_or_update. Same-name different-content records
(duplicate S_UDTs with distinct type indices) could swap symrec positions
whenever the lane->worker schedule changed (fair-share cohorts under
/RAD_SHARED_THREAD_POOL) or probe chains contended. Key on the content
hash of the full raw record instead: order becomes a pure function of
record bytes (hash -> kind -> data-bytes fallback in gsi_symbol_is_before;
byte-identical records are folded by the deduper before the sort).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 59862fa)
…c tree HashMaps, skip name decodes

- per-worker lossy open-addressing cache (obj input idx, symbol idx) -> final resolved ref;
  the resolve chain (interp parse + trie search per hop) repeated per referencing reloc
- cycle detection + per-walk visited-section set: flat arrays with linear scan instead of
  arena-backed tree HashMap nodes (same first-revisit semantics)
- lnk_resolve_symbol + walk unpack sites use lnk_parsed_symbol_from_coff_symbol_idx_no_name;
  name (string-table decode + strlen) only where a by-name symbol-table search happens
- test-and-test-and-set on is_live flags to keep already-live cachelines in shared state

Live-section set and warning behavior are byte-for-byte unchanged by construction.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit c990a45)
The remove pass was a serial O(all sections) walk on task 0 (self-labeled TODO: thread).
Section flags are per-obj so writes are disjoint; stride the obj list across tasks via
objs_by_idx. Stats accumulate per task and reduce on task 0, keeping the /OPT:REF debug-log
totals identical regardless of cohort width or schedule.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 2f86507)
Reuse the merged-leaf pointer array for sizing metadata so materialization avoids a second journal/header lookup. Store assigned-TI hash and index in adjacent 12-byte entries, preserving the former 8+4 byte footprint while reducing random cache-line loads during fixups. FN PGO A/B: leaf metadata wall median 17.35 -> 17.10 s with memory flat; packed assigned table user CPU 265.7 -> 262.5 s and merge user CPU 88.5 -> 86.5 s with wall/memory neutral. DLL/PDB byte identity and repeat determinism pass.
The color hash key's old_color is dead after group indexing and split counting. Reuse it for the assigned output color instead of retaining a separate U64 in every refinement slot. This shrinks the 16.8M-slot FN table from 48 to 40 bytes per slot (~128 MiB). Four PGO pairs: ICF wall 1.019 -> 0.990 s, user CPU 40.85 -> 38.76 s, kernel CPU 9.45 -> 8.25 s. DLL/PDB byte identity and repeat determinism pass.
After colored-hash insertion completes, the 128-bit key hash is dead. Store the old-color table slot index in those bytes instead of retaining a dedicated U32 field. Together with output-color reuse this shrinks FN's 16.8M color slots from 40 to 32 bytes (~128 MiB more). Four PGO pairs vs the 40-byte version: ICF wall 1.003 -> 0.966 s, user CPU 39.99 -> 38.64 s, kernel CPU 8.21 -> 7.10 s. DLL/PDB byte identity and repeat determinism pass.
A cached relocation target uses either a live color pointer (Regular symbols) or a static ID, never both. Union those fields and use the existing interpretation tag to select the active member, shrinking targets from 24 to 16 bytes. Four PGO pairs: ICF wall 0.959 -> 0.948 s, user CPU 38.67 -> 37.30 s, and faults 750.5K -> 732K pages. DLL/PDB identity and repeat determinism passed for the union layout.
Store resolved object input index and U32 symbol index in one U64 and recover the object through the existing objs_by_idx table. This reduces each worker's 32K-slot memo from 768 to 512 KiB (~16 MiB total). Four PGO pairs: REF wall 360.5 vs 361.5 ms and CPU neutral within noise, with ~4K fewer faults. DLL/PDB byte identity and repeat determinism pass.
Identity radix passes do not swap source and destination. When an odd number of non-identity passes runs, the final sorted values therefore remain in task.src. Copying task.dst back into the caller array restored stale partial ordering, which could leave type records unmaterialized and later crash PDB construction.\n\nCopy from the active source buffer and enable the existing object/leaf ordering check in debug builds.
On entry-point inference errors, print object and command-line default libraries plus the libraries successfully loaded. This distinguishes a lost .drectve/defaultlib from a failed archive open or symbol-index lookup in remote Wine builds.
Parallel library parsing stored invalid results one slot past the array and reported later-batch paths with a global index into a batch-local input array. Correct both indices so a malformed or transiently unreadable archive produces a deterministic error instead of memory corruption.
Recover the EpicGames#892 diagnostic omitted from the earlier clean-stack refresh and adapt it to upstream's unresolved-symbol reporting. Extend link_undef.tst with a failing missing-import case and stderr expectations.
Refresh against dev@47de6bf3. Retain regression-proven archive/override rescans, integrate ICF target offsets and the single-task fast path, and preserve migrated DEFAULTLIB/MSF/ICF test coverage.
Refresh against dev@47de6bf3 without adding the compressed-object or memory-cleanup PR layers.
C99 inline semantics after the dev portability refresh otherwise leave this address-taken callback undefined in Linux builds.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

3 participants