Skip to content

Commit 3b13ec3

Browse files
proggeramlugRalph Küpper
andauthored
perf(codegen, runtime): stop recording parameter-guard visits that can never be consulted (#8238)
* perf(codegen, runtime): stop recording parameter-guard visits that can never be consulted `js_param_type_guard` kept its visited set unconditionally, so every container it touched paid a linear scan of up to 64 inline entries and, past that, a `HashSet` insert — per ARRAY ELEMENT. The compiler owns the descriptor graph and can decide which visits are worth recording, so it now does, and the runtime reads the answer instead of recomputing it. Refs #8202. * fix(runtime): bound the param guard's cumulative visits The visit-tracking analysis reasons about the descriptor graph; value-level duplication can still re-enter an untracked node with the same address, and nesting it multiplies. MAX_DEPTH bounds depth, not work. Claude-Session: https://claude.ai/code/session_01AHvBYz7E6wWKv8kmvLLGpj --------- Co-authored-by: Ralph Küpper <ralph@skelpo.com>
1 parent 09bbf03 commit 3b13ec3

3 files changed

Lines changed: 581 additions & 14 deletions

File tree

Lines changed: 79 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,79 @@
1+
### perf(codegen, runtime): the parameter validator stops recording visits it can never consult
2+
3+
`js_param_type_guard` runs on every unproven call into a guarded
4+
ordinary-parameter clone. #8201 routed the scalar descriptors to the typed-abi
5+
leaf guards; what stayed on the interpretive validator is the structural half,
6+
and #8202 priced its **fixed** per-call component. Two of those three costs are
7+
removable without changing what the guard decides.
8+
9+
**The visited set was unconditional.** Every container the walk touched paid a
10+
linear scan of up to 64 inline entries and, past that, a `HashSet` insert — per
11+
ARRAY ELEMENT. Validating `p: { toks: Token[], pos: number }` on every `peek(p)`
12+
therefore recorded one entry per token, and not one of them could ever be
13+
consulted: `Token` lies on no descriptor cycle and is reachable by exactly one
14+
path, so a second arrival at the same `(address, node)` pair is impossible.
15+
16+
The set is load-bearing for exactly two facts, and both are properties of the
17+
immutable compiler-emitted graph rather than of the value being validated:
18+
19+
* **termination** — a value cycle (`env.parent === env`) can only walk forever
20+
through a node that reaches itself;
21+
* **no re-walk blowup** — a node the traversal can enter twice with the same
22+
address must memoize, or a shared graph re-walks exponentially.
23+
24+
So the compiler decides it. `visit_tracking_bits` runs Tarjan over the graph it
25+
just built and propagates a saturating "ways in" count from the root, then sets
26+
the high bit of the op byte on exactly the container nodes that need recording;
27+
the runtime masks the op byte and reads the bit instead of recomputing the
28+
answer per call. On `interp.ts`: `peek(p: Parser)` — 6 nodes, **0 tracked**,
29+
where every token used to be recorded; `asNum(v: Value)` — 123 nodes, **15
30+
tracked**, exactly the recursive `Node`/`Env` cluster. Descriptor length is
31+
unchanged; the bit rides in a byte that only ever held ops 0–16. The magic goes
32+
`PGT1``PGT2` so a mismatched compiler/runtime pair fails closed on the magic
33+
(guard returns 0, caller takes the generic function) rather than reading a v1
34+
blob as one that opts out of tracking everywhere.
35+
36+
**`GuardState` zeroed 1 KB of stack per call.** `inline_visited` is now
37+
`MaybeUninit`; only `[..inline_visited_len]` is ever read, and after the change
38+
above most guarded calls never write a slot at all.
39+
40+
Measured on the 19-program corpus (instructions retired, best-of-3, stdout
41+
byte-exact, `iso_miss` still `misses 0`, both arms built from their own tree
42+
with the same `-p` set and `PERRY_RUNTIME_DIR` pinned per arm). **Exactly 2 of
43+
the 19 rows emit a `js_param_type_guard` call site**`interp` and `iso_miss`,
44+
two each (`asNum`, `peek`) — and both improve: `interp` **−1.63%**, `iso_miss`
45+
**−1.36%**, peak RSS unchanged on both. Differencing against the same runtime
46+
archive so binary-layout effects cancel (a `PGT1` blob under a `PGT2` runtime
47+
fails every guard), the validator's own cost falls `interp` 1.671 B → 1.422 B
48+
(**−14.9%**, 12.05% → 10.45% of the program) and `iso_miss` 1.611 B → 1.400 B
49+
(**−13.1%**, 9.76% → 8.60%).
50+
51+
★ The other 17 rows are **not attributable in either direction**. The two arms'
52+
`libperry_runtime.a` differ in exactly two functions out of 11,185 in the
53+
crate's codegen unit — `js_param_type_guard` (808 → 316 bytes) and
54+
`GuardState::matches` (+28) — with every other function byte-identical, and
55+
those 17 programs execute neither. Their movement (`pipeline` −3.9%,
56+
`retain_wide1` +0.6%, `deeplist` −0.5%, the rest within ±0.1%) is address-layout
57+
noise: two `main` builds from identical source came out byte-identical (archive
58+
and `perry` binary alike) and repeat runs of one binary spread ~0.1%, so the
59+
build is deterministic and `pipeline`'s ±4% is what an address-hash-sensitive
60+
program does when the heap moves. None of it is claimed here.
61+
62+
★★ #8202's premise — that the fixed per-call overhead dominates — does not hold.
63+
It is ~15% of the validator's cost; the structural walk is the other ~10.5pp of
64+
`interp`. Measured separately (see the issue), a diagnostic runtime whose guard
65+
always accepts and one whose guard always rejects land within 0.3% of each
66+
other, 12% below `main`: on these two rows the specialization the validator
67+
gates is worth ~0.2% while running the validator costs ~12%. That is a policy
68+
question for #8094/#8079, not a per-call-overhead one.
69+
70+
Review follow-up: the analysis decides tracking from the DESCRIPTOR graph, but
71+
"entered twice with the same address" is a property of the VALUE. One object
72+
held at several fields re-enters an untracked node at `entries == 1`, and
73+
nesting that duplication multiplies — `d` levels of a two-way share re-walk
74+
`k^d` times where the unconditional memo ran once. The realistic sharing shapes
75+
are safe (a recursive type is on a cycle, a diamond has two ways in), but
76+
`MAX_DEPTH` bounds depth, not total work. `MAX_VISITS` now caps cumulative
77+
visits and fails the guard to the generic function — the same safe direction as
78+
the depth cap, and the better choice on its own terms past a million checks.
79+
Covered by `nested_value_duplication_through_untracked_nodes_is_bounded`.

0 commit comments

Comments
 (0)