Found by bug-hunt wave 23 (read-heavy service simulation — the workload shape the library is actually for; prior soaks were write-heavy). M4-class host, 14 logical cores, dim=128, measured per-call with perf_counter.
1. HIGH — search concurrency is capped by RAYON_NUM_THREADS, not by caller threads. Every search on an index of ≥8192 vectors goes through the one process-local rayon pool (turbovec-python/src/lib.rs:438, :801 → with_pool at :1192). pool.install() from an external thread injects the job into that shared pool and blocks the caller, so concurrent searches are capped by pool size and Python-side concurrency contributes nothing.
n=200k, nq=1, unmasked, calls/s:
| RAYON_NUM_THREADS |
1 py thread |
4 py threads |
14 py threads |
| 1 |
1919 |
1919 |
1902 (0.99x) |
| 2 |
3420 |
3464 |
3661 |
| 4 |
5491 |
6846 |
6014 |
| 14 |
5010 |
10192 |
13087 |
With RAYON_NUM_THREADS=1, 14 concurrent request threads deliver exactly the throughput of one — and that's the setting an operator naturally picks to avoid oversubscription in a threaded server, and the one the README documents for pinned runs (README.md:167/191/291). The GIL is not the cause; it's released. Same flat line for nq=8 (0.99x) and masked (1.03–1.08x). Even at the best setting efficiency is 49% (13,087 q/s vs a 26,866 q/s ceiling), with per-query core-time roughly doubling (521 → 1070 µs).
2. MEDIUM-HIGH — unmasked read-heavy scaling saturates at 2.4x on 14 cores while p999 degrades 32x.
| config |
T=1 |
T=4 |
T=8 |
T=14 |
| unmasked nq=1 q/s |
5324 |
10459 |
12332 |
12884 (2.42x) |
| p50 / p999 µs |
177 / 410 |
332 / 2742 |
468 / 3923 |
542 / 13194 |
| unmasked nq=8 q/s |
11535 |
18596 |
20685 |
21468 (1.86x) |
| p50 / p999 µs |
679 / 1079 |
1441 / 9002 |
2009 / 24005 |
2289 / 67939 |
Throughput is flat from T=8→T=14 (+4%) while p999 triples — the added load buys nothing but tail. Head-of-line blocking at the shared pool's injection queue, same root as (1).
Important counterpoint for #294: in a service shape masked search scales better than unmasked — masked nq=1 goes 1756 → 10,485 q/s (5.97x) vs unmasked's 2.42x, and at T=14 masked nearly catches unmasked despite being 3.0x slower single-threaded. Unmasked is the path with the hard ceiling. Worth knowing before "fixing" #294 by making masked behave like unmasked.
3. MEDIUM — undocumented latency cliff at exactly n=8192: a 0.4% larger index makes nq=1 search 3.4x slower (p50) and 5–8x slower (p99). single_query_parallelizes() (search.rs:19,26) routes nq=1 into the pool at 256 blocks = 8192 vectors:
| n |
blocks |
pooled |
p50 µs |
p99 µs |
| 8160 |
255 |
no |
24.7 |
86.5 |
| 8192 |
256 |
yes |
84.1 |
16431 |
| 8224 |
257 |
yes |
74.0 |
296.6 |
| 262144 |
8192 |
yes |
497.0 |
2370.7 |
The ~70 µs pool handoff (acknowledged at lib.rs:1222) is 2.8x the entire inline search at that size. Fitting both regimes — inline ≈ 2.31 µs/1k vectors, pooled ≈ 70.7 µs + 1.63 µs/1k — puts break-even at n ≈ 104,000, so the gate fires about 12x too early. Every nq=1 search on an index between 8k and ~100k pays up to 3.4x more latency than the inline path it just left. Nothing in docs/ or README mentions 8192 or 256 blocks. Thread scaling also inverts across the boundary (n=8000: 2.17x at T=14; n=8224: 4.93x).
4. MEDIUM — one writer doing single-row adds halves read throughput. n=200k, one writer looping add(one_row):
| T readers |
writer |
reads/s |
p50 |
p99 |
p999 |
| 4 |
off |
11047 |
320 |
930 |
1242 |
| 4 |
on |
6633 (−40%) |
468 |
2056 |
2934 |
| 8 |
off |
11341 |
502 |
3046 |
4581 |
| 8 |
on |
5673 (−50%) |
1178 |
4466 |
6530 |
An uncontended 1-row add is 6.2 µs and — good news — flat in n (6.7/6.2/6.8/6.2 µs at n=50k/200k/500k/1M), so hold time is O(batch), not O(n). At ~4000 writes/s that's ~2.4% of one core of work destroying ~48% of an 11k q/s read service. The RwLock is fair/queueing rather than read-preferring, so writers aren't starved — but each queued writer barriers all subsequent readers, and because the read guard is held across the entire pooled search (lib.rs:408/:768 — lock_read then with_pool_if inside the guard), every write insertion drains and refills the pool. Distinct from #289/#319: there the GIL is the problem, here the GIL is released and the lock itself is the cost.
Bulk case, n=1M, 1 reader: add(50k) holds the lock 14.2 ms and produces a 14.5 ms reader stall against a 722 µs baseline — a 20x spike. So the operational rule is "a search can be blocked for as long as your largest add batch"; a 50k-row ingest is a 14 ms hiccup on every in-flight query.
🤖 Generated with Claude Code
Found by bug-hunt wave 23 (read-heavy service simulation — the workload shape the library is actually for; prior soaks were write-heavy). M4-class host, 14 logical cores, dim=128, measured per-call with
perf_counter.1. HIGH — search concurrency is capped by
RAYON_NUM_THREADS, not by caller threads. Every search on an index of ≥8192 vectors goes through the one process-local rayon pool (turbovec-python/src/lib.rs:438, :801 →with_poolat :1192).pool.install()from an external thread injects the job into that shared pool and blocks the caller, so concurrent searches are capped by pool size and Python-side concurrency contributes nothing.n=200k, nq=1, unmasked, calls/s:
With
RAYON_NUM_THREADS=1, 14 concurrent request threads deliver exactly the throughput of one — and that's the setting an operator naturally picks to avoid oversubscription in a threaded server, and the one the README documents for pinned runs (README.md:167/191/291). The GIL is not the cause; it's released. Same flat line for nq=8 (0.99x) and masked (1.03–1.08x). Even at the best setting efficiency is 49% (13,087 q/s vs a 26,866 q/s ceiling), with per-query core-time roughly doubling (521 → 1070 µs).2. MEDIUM-HIGH — unmasked read-heavy scaling saturates at 2.4x on 14 cores while p999 degrades 32x.
Throughput is flat from T=8→T=14 (+4%) while p999 triples — the added load buys nothing but tail. Head-of-line blocking at the shared pool's injection queue, same root as (1).
Important counterpoint for #294: in a service shape masked search scales better than unmasked — masked nq=1 goes 1756 → 10,485 q/s (5.97x) vs unmasked's 2.42x, and at T=14 masked nearly catches unmasked despite being 3.0x slower single-threaded. Unmasked is the path with the hard ceiling. Worth knowing before "fixing" #294 by making masked behave like unmasked.
3. MEDIUM — undocumented latency cliff at exactly n=8192: a 0.4% larger index makes nq=1 search 3.4x slower (p50) and 5–8x slower (p99).
single_query_parallelizes()(search.rs:19,26) routes nq=1 into the pool at 256 blocks = 8192 vectors:The ~70 µs pool handoff (acknowledged at lib.rs:1222) is 2.8x the entire inline search at that size. Fitting both regimes — inline ≈ 2.31 µs/1k vectors, pooled ≈ 70.7 µs + 1.63 µs/1k — puts break-even at n ≈ 104,000, so the gate fires about 12x too early. Every nq=1 search on an index between 8k and ~100k pays up to 3.4x more latency than the inline path it just left. Nothing in docs/ or README mentions 8192 or 256 blocks. Thread scaling also inverts across the boundary (n=8000: 2.17x at T=14; n=8224: 4.93x).
4. MEDIUM — one writer doing single-row adds halves read throughput. n=200k, one writer looping
add(one_row):An uncontended 1-row add is 6.2 µs and — good news — flat in n (6.7/6.2/6.8/6.2 µs at n=50k/200k/500k/1M), so hold time is O(batch), not O(n). At ~4000 writes/s that's ~2.4% of one core of work destroying ~48% of an 11k q/s read service. The RwLock is fair/queueing rather than read-preferring, so writers aren't starved — but each queued writer barriers all subsequent readers, and because the read guard is held across the entire pooled search (lib.rs:408/:768 —
lock_readthenwith_pool_ifinside the guard), every write insertion drains and refills the pool. Distinct from #289/#319: there the GIL is the problem, here the GIL is released and the lock itself is the cost.Bulk case, n=1M, 1 reader:
add(50k)holds the lock 14.2 ms and produces a 14.5 ms reader stall against a 722 µs baseline — a 20x spike. So the operational rule is "a search can be blocked for as long as your largest add batch"; a 50k-row ingest is a 14 ms hiccup on every in-flight query.🤖 Generated with Claude Code