You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
at naive implementation tier; reduces to rank-1 small-M GEMV + epilogue).
1328
+
1329
+
**Cross-fire note**: PR #1922 (F-WEDGE-SMALL-M-GEMV-WALL, landed concurrently) also
1330
+
falsifies the rank-1 small-M GEMV wedge as BW-bound (real ceiling 1.19x, not 5-10x).
1331
+
Combined with this fire, the 2026-05-27 ranked-wedge table's top-2 candidates are now
1332
+
both 🔴 — the LM-head decode pipeline has no obvious single-kernel wedge above 1.10x
1333
+
under naive hand-emit. Next direction = (a) tiled GEMM-style design with register top-K
1334
+
epilogue, or (b) honestly retire the entire decode wedge family pending mma.sync /
1335
+
warp-cooperative redesign.
1336
+
1268
1337
## 2026-05-28 — F-WEDGE-SMALL-M-GEMV-WALL fire — 🔴 FALSIFIED (empirical seal on Round 10 analytical retirement)
1269
1338
1270
1339
본 fire 는 2026-05-27 ranked-wedge #1 (small-M GEMV, 추정 ceiling 5-10×) 의 pre-registered falsifier `F-WEDGE-SMALL-M-GEMV-WALL` 을 hand-emit GEMV vs cuBLAS 직접 head-to-head 로 검증. Round 10 (line 1021) 은 동일 wedge 를 honest HBM-roofline 분석으로 retire 했고 — 본 fire 는 그 retirement 의 **empirical seal** (분석 retirement 와 측정 retirement 의 일치).
Round 10 (analytical) + 본 fire (empirical) 의 일치 = **roofline reference 는 AI-aware 해야 한다** (`feedback_closure_is_physical_limit` g0 instance #2). compute-peak gap 을 memory-bound op 에 적용하면 phantom wedge 가 생긴다. Methodology 가 cheap-first oracle 의 ceiling 추정 안에서 *roofline kind* 를 명시하도록 다음 ranked-wedge 표는 "ceiling (vs compute peak | vs BW peak)" 두 칸 분리 권장.
1309
1378
1310
1379
cycle 결과: F-WEDGE-SMALL-M-GEMV-WALL = 🔴 closed-negative. 다음 active wedge = top-K fusion (rank 1 new).
1380
+
1381
+
**Cross-fire note**: PR #1925 (F-WEDGE-TOPK-FUSED-WALL, landed concurrently) also falsifies the rank-1-new top-K fusion wedge at the naive hand-emit tier (0.066x at M=8 LLaMA = 15x slower than cuBLAS+cub). Combined with this fire, the 2026-05-27 ranked-wedge table's top-2 candidates are both 🔴 — LM-head decode has no obvious single-kernel wedge above 1.10x under naive hand-emit.
-[]**No PyBind11 / no ATen dispatch overhead** — cuBLAS via PyTorch goes through Python → C++ Tensor → ATen → CUDA stream → cuBLAS handle. Hexa-emit directly compiled into the binary
685
-
-[]**Static kernel selection** — cuBLAS-LT runtime heuristic picks an algorithm; hexa compile-time selects + bakes the algorithm
686
-
-[]**Single-shot binary** — no shared library boundary, no `cudaGetSymbolAddress`, no driver-level dispatch table
684
+
-[x]**No PyBind11 / no ATen dispatch overhead** — cuBLAS via PyTorch goes through Python → C++ Tensor → ATen → CUDA stream → cuBLAS handle. Hexa-emit directly compiled into the binary (**R16 2026-05-28** evidence-flip: §5l ldd fire — standalone host binary 6 dyn libs · libcuda.so.1 + libc/libm/libdl/libpthread/librt + ld-linux, zero libcudart/libcublas/Python; `archive/fires/gpu_multiarch_fatbin_probe_2026_05_28/ldd_standalone.txt`)
685
+
-[x]**Static kernel selection** — cuBLAS-LT runtime heuristic picks an algorithm; hexa compile-time selects + bakes the algorithm (**R16 2026-05-28** evidence-flip: R14 RoPE PTX `tool/artifacts/rope_f64_trigfix_2026_05_28.ptx` 가 Taylor 계수 9개를 `0d3CE952C77030AD4A` 등 hex immediate 로 PTX 에 직접 베이크 — runtime dispatch / heuristic-select 0건. 동일하게 R15 `cvt.rn.f64.s64` instruction도 compile-time 확정; ptxas JIT 가 SASS 로 lowering 만 함, 알고리즘 선택은 hexa codegen 에서 끝남)
686
+
-[x]**Single-shot binary** — no shared library boundary, no `cudaGetSymbolAddress`, no driver-level dispatch table (**R16 2026-05-28** evidence-flip: §5l Standalone cubin embed silicon-validated `F-GPU-STANDALONE-CUBIN` — `unop_cubin_data.h` xxd-embed into 22176 B host binary, `cuModuleLoadData` from libcuda.so.1 직접 호출, shared-lib 경계 부재. `cuModuleGetFunction` 으로 entry 잡으므로 driver-level dispatch table 도 미사용. `archive/fires/gpu_standalone_cubin_probe_2026_05_28/{result.json, ldd_standalone.txt}`)
687
687
688
688
### 5g — Operator-specific surgical override
689
689
@@ -703,7 +703,7 @@ PR #189/#190/#191 fires used direct one-shot bash; sustained automation needs he
703
703
704
704
-[ ]**`.so` blob vs source emit** — cuBLAS is closed-source binary; hexa users see + modify the emit path. Bug-fix loop: cuBLAS = file ticket + wait; hexa = patch source + rebuild
705
705
-[ ]**`hexa gpu disasm`** — view exact SASS via `cuobjdump`; cuBLAS too but harder to correlate to user code (the high-level mapping is lost in the closed binary)
706
-
-[]**Single-language stack** — host + device + autograd all in hexa-lang; cuBLAS-using stacks need Python/C++/CUDA polyglot
706
+
-[x]**Single-language stack** — host + device + autograd all in hexa-lang; cuBLAS-using stacks need Python/C++/CUDA polyglot (**R16 2026-05-28** evidence-flip: 구조적 사실 — `stdlib/flame/*.hexa` host training + `compiler/codegen/nvptx_target.hexa` device emit + `stdlib/flame/ag_tape.hexa` autograd 가 전부 hexa source. R15 NN-primitive 4-surface 빌트인 (softmax/swiglu_vec/layer_norm/rope_pair) 도 hexa codegen + hexa runtime; Python/C++/CUDA polyglot 부재. README "🔥 flame + 🔧 forge" 섹션 의 architecture 도식 참조)
"notes": "Hand-emit fused kernel achieves correct top-K (8/8 + 32/32) but is 15x-44x slower than cublasSgemm + cub::DeviceSegmentedRadixSort, far below the 1.10x SUPPORTED-NUMERICAL threshold and the 1.02x FALSIFIED threshold. Falsifies the cheap-first oracle's implicit assumption that hand-emit fusion can approach the 1.10-1.30x realistic ceiling. The naive design lacks shared-memory tiling of B and warp-level cooperative K-reduction; cuBLAS already optimizes these to ~7.9% of FP32 peak even in small-M regime, and cub::DeviceSegmentedRadixSort is far more efficient than thrust::sort for the top-K stage. The 2026-05-27 oracle's 1.802x ceiling assumed FREE fusion of top-K on top of cuBLAS GEMM; replacing both with naive hand-emit overshoots both stages by an order of magnitude. Closed-negative finding: top-K fusion is NOT a viable wedge with naive hand-emit. A real wedge would need to (a) match cuBLAS small-M GEMM throughput first, then (b) fuse top-K. The ranking implied by the 2026-05-27 oracle that 'top-K fusion is rank 2 behind small-M GEMV' is essentially refuted at the implementation tier."
0 commit comments