AKA is an end-to-end Agent project for GPU kernel implementation, analysis, profiling, and iterative optimization. It helps an Agent turn PyTorch logic or an existing kernel into a high-performance GPU kernel through a structured, profile-driven workflow.
- [2026-07] We helped Qwen3.8 rank No. 1 on the SOL-ExecBench FlashInfer operator optimization leaderboard. [Leaderboard]
- [2026-07] We released Atrex Kernel Agent v0.2.0 with a dual-route optimization system, an orchestrated clean-session loop, native SOL-ExecBench operator workflow, Triton-to-Gluon conversion support, and a fuller NVIDIA profiling toolchain. [Release]
- [2026-07] We released the Atrex paper: Are LLM-Generated GPU Kernels Production-Ready? A Trace-Driven Benchmark and Optimization Agent.
- [2026-06] We released Atrex Kernel Agent v0.1.0 as the initial open-source version, with the interactive
gpu-kernel-optimizerSkill route, GPU Wiki knowledge base, profile-driven optimization workflow, profiling tools, and reference templates. [Release]
- Creates an isolated optimization workspace under
kernel_opt_<name>/. - Looks up target hardware specs from the local
gpu-wikiknowledge base. - Runs Roofline analysis and sets auditable performance targets.
- Implements a correct baseline kernel before entering optimization.
- Runs the profile-driven optimization loop: profile with
ncuorrocprofv3, extract bottleneck evidence, querygpu-wiki/ reference projects / web sources for relevant optimization knowledge, write an evidence-based plan, apply one optimization category, validate correctness and performance, record memory, commit, then repeat until Stop Conditions are met. - Records plans, profile artifacts, structured memory, reports, and Git commits for every accepted iteration.
For the full architecture and workflow design, see docs/design.md.
See the Quick Start guide for prerequisites, installation, and complete runnable paths for both the interactive Skill route and the orchestrated loop route.
| Route | Driver | Termination | Best for |
|---|---|---|---|
| Route 1: Interactive Skill | gpu-kernel-optimizer Skill + hooks, invoked inside a coding session |
In-session judgment, guarded by hooks | Hands-on, interactive optimization from a coding runtime |
| Route 2: Orchestrated Loop | orchestrator/optimize.py, spawning fresh clean sessions per iteration |
Mechanical (max iterations / token budget / target utilization) | Unattended, budget-bounded, batch optimization |
Both routes share the same knowledge base (gpu-wiki/), reference projects, tools (tools/), and structured memory format (memory/v<N>.json).
This route installs the gpu-kernel-optimizer Skill and workflow hooks into your coding runtime. You then drive the optimization interactively from a coding session, and the hooks keep the workflow on track (memory reads, plan reads, correctness gates, stop-condition checks).
The optimization workspace kernel_opt_<name>/ is created in the current working directory where you run the session, so all artifacts stay next to where you are working.
Internal users should configure git insteadOf URL redirect rules so that submodules and dependencies resolve against the internal network before running git submodule update. External users can skip this step entirely.
The install path is optional; defaults to ~/aka_kernel_opt.
Common installer options:
bash install.sh --prefix ~/my_path # Install to a custom directory
bash install.sh --hooks-only # Install or update hooks only
bash install.sh --without-github # Skip GitHub-hosted reference repos
bash install.sh --uninstall # Remove hooks installed by this scriptThe installer detects supported runtime home directories and prepares local hooks when available. It ships only the Skill route; the orchestrator route (Route 2) runs from the source repo and is pruned from the installed skill directory.
This route runs the optimization loop from the source repo without installing anything into your coding runtime. orchestrator/optimize.py owns the outer loop and spawns a fresh, clean Claude, Qoder, Codex, or Pi CLI session for each iteration. Select the backend with --agent-cli claude|qodercli|codex|pi (default: claude). State crosses the session boundary only through disk (memory/v<N>.json, plans/, profiles/, and git), and HEAD is always the best kernel. Codex runs use codex exec --json --ephemeral; repository-scoped skills are prepared under each campaign's .agents/skills/ without modifying the user's global Codex installation.
For single-operator SOL and atrex-bench campaigns, the default outer flow first collects evaluator-faithful, production-visible runtime signatures in the sandbox. These signatures contain only explicit non-tensor arguments and tensor shape/stride/dtype/layout metadata—never tensor contents or evaluator-only workload values. The workload inspector runs in a data-minimized temporary workspace containing only those signatures and writes an exact, disjoint workload_buckets.json; every bucket boundary must therefore be reproducible by the no-sync production dispatcher, and indistinguishable signatures cannot be split. Every bucket then runs the original optimization loop concurrently in an independent Git workspace. The first ten iterations are an aggregation warmup: improvements are recorded but do not edit the main kernel. Once every bucket has reached at least V10 and has a committed improvement, the orchestrator deterministically copies every bucket's committed kernel and generates an exact runtime dispatcher—no coding-agent/LLM aggregation is used. Every later bucket improvement replaces only that bucket module and regenerates the dispatcher. Every candidate is accepted only after a separate full-workload, multi-seed correctness run and full-workload geomean benchmark beat the main incumbent. Dispatcher sources, visibility policy, provenance, pending improvements, accepted kernels, and rejected attempts are auditable in the main workspace's Git history, dispatch_signatures.json, aggregate_dispatch.json, and aggregation_state.json.
Correctness/performance validation and profiling run on an atrex-gpu-gateway sandbox selected by
--sandbox-hardware. The gateway worker receives code and test/profile inputs only: optimizer memory/, plans,
edits, and Git state remain local. Structured test results and profile analysis artifacts are returned to the
local session. Evaluation is selected by input format: native atrex-bench operators (shapes.json) use the
canonical atrex-bench/scripts/run_eval.py, with workspace test_kernel.py acting only as an immutable
result adapter; SOL operators (definition.json + workload.jsonl) continue using SOL-ExecBench unchanged.
The same transport can be used directly:
python tools/sandbox.py --hardware REMOTE_GPU --no-sync -- python test_kernel.py --no-memory
python tools/sandbox.py --hardware REMOTE_GPU --sync profiles/v1 -- \
bash tools/profile_nvidia.sh kernel.py --output-dir profiles/v1 --source
# Same interface on the bundled localhost FIFO scheduler
# Start it first with: python tools/local_gateway.py serve
python tools/sandbox.py --hardware local --url http://127.0.0.1:8000 \
--no-sync -- python test_kernel.py --no-memoryLocal gateway mode preserves the request/packaging/result interface but is not a security sandbox:
submitted commands run directly as the server user. The bundled scheduler serializes jobs by default,
persists their status in SQLite, and speaks the same public agate dev/jobs API. See
docs/local_gateway.md for startup, queue, cancellation, and compatibility details.
Termination is mechanical, not left to in-session judgment: the loop stops on a hard budget (max iterations or token budget) or a target-utilization short-circuit on a committed, correctness-passing iteration.
Everything op-specific (workspace name, reference, and full workload/shape set) is read from --op-dir.
Ground-truth files are never edited. Bucket workspaces receive derived filtered workload.jsonl or
shapes.json copies, while the main workspace retains and validates the complete set. --platform is
required. In the default leaderboard mode, --framework may select one framework explicitly; when omitted,
the orchestrator launches independent campaigns in parallel for Triton/CuteDSL/Cuda on NVIDIA,
Triton/FlyDSL on AMD, or Triton on unknown hardware.
Key options:
--max-iters N # Hard cap on optimization iterations
--max-workload-buckets N # Inspector bucket cap (default 8)
--aggregate-min-improvement-pct PCT # Full-workload gain required for aggregate acceptance
--no-workload-bucketing # Restore the legacy single-workspace SOL flow
--token-budget N # Hard token cap across all sessions (0 = no cap)
--agent-cli CLI # Optimization session backend: claude (default), qodercli, codex, or pi
--optimization-mode MODE # leaderboard (default) or production
--framework DSL # One explicit DSL; omit to parallel-dispatch all supported DSLs
--target-util PCT # Peak-utilization %% short-circuit (default 90)
--sandbox-hardware GPU # agate selector/alias; independent of the logical --platform name
--sandbox-profile P # Optional pre/prod endpoint; default uses agate config
--sandbox-url URL # Explicit endpoint; use http://127.0.0.1:8000 with hardware=local
--sandbox-timeout S # Remote command timeout, max 600 seconds
--workspace DIR # Working directory for the campaign (default: current directory)
--max-stall N # Stop after N consecutive no-commit iterations (0 = disabled)
--convert-after N # Triton only: after N stalls, require Gluon conversion until it succeeds (default 3)
--arch ARCH # Override auto-detected runtime arch, e.g. sm_103 or gfx942Auto-dispatched main campaigns use flat framework/hardware suffixes; for example,
<workspace>/kernel_opt_<name>_triton_h20 and
<workspace>/kernel_opt_<name>_cutedsl_h20. Each main workspace owns its full-workload kernel.py,
bucket manifest, aggregation history, and ignored workload_buckets/ directory containing the
independent bucket Git workspaces. Each bucket receives its own full iteration and
token budgets. Explicit --framework campaigns use the same naming convention.
--optimization-mode leaderboard preserves the existing permissive CLAUDE.md workflow: sessions may
use a different/mixed implementation or third-party kernel libraries when profiling evidence supports it.
--optimization-mode production also supports omitted --framework: the orchestrator auto-dispatches the
hardware-supported frameworks and binds every child campaign to its assigned framework. V0 may remain the
PyTorch correctness baseline, but every accepted optimized candidate must be implemented directly and
exclusively in that child's framework. Third-party kernel/operator imports, calls, and solution dependencies are forbidden. A mechanical
post-session gate rejects and reverts non-compliant kernel commits, records a production_policy_rejection,
and refuses to package a non-compliant final kernel. A production Triton campaign escalates to the same
toolchain's Gluon DSL after three consecutive stalls. Once triggered, conversion is mandatory and retries
immediately until correctness and performance parity pass; later iterations remain in Gluon.
python orchestrator/optimize.py \
--op-dir /path/to/op --platform TARGET_GPU --sandbox-hardware REMOTE_GPU \
--optimization-mode production --framework Triton--platform is a logical optimization target while --sandbox-hardware is the gateway selector. The
orchestrator deliberately does not compare their names or reported GPU models because gateway inventory
may be aliased or desensitized. Runtime architecture probing remains authoritative when an omitted
--framework requires vendor-specific dispatch.
All four backends run non-interactively with clean session state and the same workspace-local skills,
prompts, sandbox constraints, and quality gates. Authenticate the selected CLI first with
claude auth status, qodercli status, codex login status, or pi --list-models. Provider-specific
settings can be supplied through ATREX_CLAUDE_SESSION_SETTINGS, ATREX_QODER_SESSION_SETTINGS,
ATREX_CODEX_SESSION_SETTINGS, or ATREX_PI_SESSION_SETTINGS; ATREX_SESSION_SETTINGS remains the
generic fallback. For Codex,
the setting value must be either a JSON object or a JSON array of literal key=value strings and is
translated to repeatable codex exec -c arguments, for example:
export ATREX_CODEX_SESSION_SETTINGS='{"model":"gpt-5.6-sol","model_reasoning_effort":"xhigh"}'Pi uses --mode json, a unique persisted session id, and the configured Pi provider/model. Optional
provider/model selection is restricted to non-secret CLI values:
export ATREX_PI_SESSION_SETTINGS='{"provider":"anthropic","model":"claude-opus"}'
python orchestrator/optimize.py ... --agent-cli piCodex JSONL turn.completed.usage is included in token-budget accounting. Cache and reasoning
sub-counters are not double-counted. Pi finalized message usage, including cache read/write counters,
is aggregated after agent_settled. Some Qoder models report zero token usage in stream JSON; in
that case --token-budget cannot be enforced and --max-iters remains the hard campaign bound.
.
├── SKILL.md # Route 1: gpu-kernel-optimizer router manifest
├── install.sh # Route 1 installer / uninstaller
├── orchestrator/ # Route 2: clean-session optimization orchestrator
│ ├── optimize.py # Outer optimization loop driver
│ └── prompts/ # Per-session prompts (setup, framework baseline, iteration, convert)
├── agents/ # Subagent definitions used by both routes
├── docs/ # Detailed project design docs
├── reference/ # Workspace, plan, memory, and profiling templates
├── skills/ # Baseline, optimizer, restart, and output-contract modules
├── tools/ # Profiling, utilization, memory, and measurement tools
└── gpu-wiki/ # Local GPU knowledge base
This project builds on and references many excellent open-source works. We gratefully acknowledge the authors and communities behind them.
Reference kernel projects (reference-projects/):
- CUTLASS — CUDA Templates for Linear Algebra Subroutines
- cutex — CUDA Template Extensions
- cuLA — inclusionAI CUDA Linear Algebra
- flash-attention — Flash Attention
- FlashInfer — Kernel library for LLM serving
- FlyDSL — ROCm FlyDSL
- Triton — Triton language and compiler
- DeepGEMM — DeepSeek DeepGEMM
- LeetCUDA — CUDA learning kernels
- FlashMLA — DeepSeek FlashMLA
- Composable Kernel — ROCm Composable Kernel
- cute-gemm — CuTe GEMM examples
- hpc-ops — Tencent HPC Ops
- aiter — ROCm AIter
- quack — Dao-AILab Quack
- tilelang — TileLang
Knowledge base and tooling (gpu-wiki/3rdparty/, 3rdparty/):
- KernelWiki — GPU kernel knowledge base
- modern-gpu-programming-for-mlsys — Modern GPU programming for MLSys
- ncu-report-skill — Nsight Compute report parsing skill
- humanize — Plan generation plugin
- AKO4ALL — AKO4ALL
- KDA — Kernel Design Agents
Please cite our paper if it is helpful to your research.
@misc{atrex2026,
title = {Are LLM-Generated GPU Kernels Production-Ready? A Trace-Driven Benchmark and Optimization Agent},
author = {Lingyun Yang and Yuxiao Wang and Shenghao Liang and Linfeng Yang and Daocheng Ying and Chunbo You and Rui Zhang and Luping Wang and Yinghao Yu and Guodong Yang and Liping Zhang},
year = {2026},
eprint = {2607.14541},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2607.14541}
}Licensed under the Apache License 2.0.


