Select checks from the actual risks, publish reviewable evidence, and distinguish local qualification from hosted CI results.
Pull requests cancel superseded revisions and run cheap contracts by default.
Master finishes active full verification and keeps the newest pending push; all
retained master pushes select the full development matrix. Nightly fallback runs
at 06:00 UTC in a separate concurrency group without publication operations, and
skips expensive jobs only with matching available successful full evidence for
its exact revision and policy.
Release packaging first runs the full development matrix on the exact candidate
revision through build.yml's revision workflow-call input.
CI measurements batch completed runs hourly using trusted default-branch code and
inert artifacts. Cancellations do not create measurement workflows, and cheap-only
observations are excluded. Explicitly requested verification on drafts is included.
Batches retain attempt identities to avoid remeasuring
completed runs and attempt at most ten new measurements per batch. feedback.json reports p90 after ten successful matching-inventory
PR samples, with explicit gaps when there are fewer samples. Metrics report queue delay, active execution time, idle
gaps, aggregate runner minutes, feedback time and cache observations separately.
Per-job queue timestamps are estimates, not proof of runner saturation. Overlapping
jobs count once in wall execution time and separately in runner minutes.
Compare ten successful runs with matching event and exact coverage inventory using
python3 .github/scripts/ci_run_metrics.py --before before.json --after after.json.
Report workload reductions separately from execution savings. The initial service
objective for explicitly requested affected-PR verification is p90 feedback below
15 minutes and queue delay below two minutes; full verification and releases may
take longer. Cheap-contract feedback is not engine verification feedback. Missing
samples and missing inventories cannot establish improvements.
GCC 13 builds its client and applicable transport programs once. Runtime consumers reuse same-run artifacts with matching source/configuration provenance. Compiler build jobs publish programs before separate CLI/scripting regressions; native and browser consumers can start without waiting for those regressions. Golden-only changes build once per selected platform and distribute programs to sweep consumers. Primary generator sweeps are complete; secondary platforms retain golden rows and telemetry equivalence under enabled tiers. Native inventory audits require every selected engine case to belong to exactly one shard.
ELF dependency collection records runtime package owners and rejects unresolved or
unowned libraries. Generator, CLI and browser consumers install runtime libraries
rather than compiler/header environments. Enable CI_RUNTIME_PACKAGES_ENABLED=true
for ordinary engine shards only after Validate clean Linux runtime images passes
both Ubuntu container images. Until then engine shards retain their prior package
installation; the clean-image dispatch also exercises CLI/scripting consumers.
Pinned SDL3, WebAssembly and Android dependency prefixes are cached independently from application objects. Restored prefixes require matching inputs and complete file hashes; invalid content rebuilds. Android tooling and Playwright installations use pinned tool/package identities. Default-branch jobs publish shared caches; ordinary PRs restore them. The cleanup sweep removes closed-PR caches and retains three generations per master compiler/object family. Cache timings and quota pressure must be measured before expanding retained cache data.
Native runners optionally accept --timing-profile and --auxiliary-jobs.
Nonempty profiles assign jobs longest first to the least-loaded shard, with label
and shard-number tie breaks; unknown cases use the median recorded duration.
Empty or omitted profiles preserve alphabetical slicing and existing auxiliary
ownership. Unit tests and auxiliary groups participate in the same load plan,
retaining their original commands, flags and timeouts. Every selected engine job
runs exactly once. Auxiliary artifacts follow their assigned group.
Successful job durations are retained alongside JUnit results. To propose updated
weights, gather observations and run
python3 test/build_ci_timing_profile.py observations --family ubuntu-24.04 --output test/ci-timings/ubuntu-24.04.json.
Only platform-matched jobs with ten successful samples enter a profile. Review the
resulting diff before shipping it; weights never change during a run. Profiles
remain empty until measurements are available, rather than using invented data.
Relevant local or VM testing is the standard PR verification path. Maintainers accept evidence directly, including their own evidence; hosted CI success is not a merge prerequisite. Choose tests from the change's actual risks and justify coverage and omissions rather than reproducing the CI matrix. A VM supplies coverage for its actual OS, architecture and configuration; testing on one platform does not establish another platform's compatibility.
Record evidence in a PR comment using this template, with links accessible to
reviewers. Store generated files under ignored artifacts/ and temporary narratives
under docs/.work/, then upload or attach evidence for review; local paths alone
are insufficient. Keep secrets out of uploaded logs.
Local / VM verification
- Tested commit SHA:
- Base revision and integration state (PR head or merge with base):
- Environment: OS, architecture, VM/container image if applicable, compiler/runtime versions:
- Dependencies, build configuration and flags:
- Coverage rationale: changed behavior and risks addressed:
- Exact build/test commands and results (including counts and exit status):
- Omitted checks and why; limitations:
- Evidence: accessible links to logs and applicable checksums, saves, replays or screenshots:
- Maintainer acceptance: sufficient evidence for this revision, accepted by <name>:Evidence must describe the tested source and binaries; reuse built artifacts only when source, compiler, flags and dependency inputs match. Refresh evidence when later edits affect tested behavior, dependencies or integration. Fetch current master before final validation and resolve actual conflicts. Unrelated master advancement alone does not invalidate evidence; changes in the same components, dependencies or CI configuration require renewed integration assessment.
Focused coverage does not waive affected simulation determinism, save/load,
replay/network, platform compatibility or simulation-version requirements in
AGENTS.md. Local and hosted results may jointly supply that coverage. Hosted
checks may be pending or unavailable when merging; known failures introduced by
the PR still require resolution.
Draft and ready PRs run the existing cheap contracts by default. Changes under tools/music/ also run the cheap music job (the pipeline's
Python unit tests in a venv from requirements.txt), on PRs without ci:run too; it
gates no engine verification. Community converter changes also select platform and
stack verification; shared WASM exports/build inputs select full verification
when hosted checks are requested. The music job includes a small C++ portable-file
round trip using libopusfile. Files inside a soundtrack set, data/zik/<set>/, select
the native and browser checks when hosted verification is requested, while
data/zik/SConscript stays on full CI. Becoming ready
starts no expensive jobs. ci:run requests hosted affected checks; ci:full
requests the complete development matrix, even in draft. ci:windows, ci:android
and ci:browsers expand requested coverage but do not start verification alone.
Label changes re-evaluate selection; removing the request labels restores
cheap-only selection. The aggregate summary rejects missing, failed, cancelled
and unexpectedly skipped selected jobs, while clearly distinguishing cheap-only
success from engine verification or acceptance of PR evidence.
Every retained master push runs the full development matrix regardless of tier settings. Active runs finish and only the newest pending push remains. Full master CI detects regressions asynchronously; existing master failures do not restrict PR merges. Retain failure artifacts, prioritize diagnosis and repair, document verification in repair PRs and confirm recovery with subsequent full master runs. Do not require master to become green before other PRs merge.
For focused macOS qualification, dispatch .github/workflows/ci-macos.yml
with an exact revision and coverage_profile=full (or compatibility for the
selected compatibility inventory). It runs the same native build, regression,
continuation, scripting and strict generator evidence steps as the reusable
workflow, with a 90-minute job ceiling. This is macOS evidence, not a full
development checkpoint.
Nightly is a fallback with a separate concurrency group. Expensive nightly jobs are skipped only when a successful full hosted run already covers the exact master SHA under the current coverage policy and its evidence is available. Missing, expired, mismatched or inaccessible evidence triggers the full matrix. A skipped nightly is not a new full checkpoint. Manual and release verification retain their existing behavior. Local evidence never substitutes for a hosted full checkpoint.
.github/scripts/ci_policy.py records proposed and effective selection, reasons,
changed paths, policy identity and selected command inventory in ci-selection.json.
The observation also identifies cheap-contracts, affected, full or
nightly-reused verification and the reused run ID when applicable. Native runners
retain eligible/assigned case inventories. Cheap-only PR runs do not enter ordinary
verification performance cohorts; requesting tests on a draft does not exclude
actual verification from those cohorts.
For explicitly requested affected PR verification, primary Linux keeps the
complete applicable native suite. Simulation/save/AI changes add older-GCC and
Windows compatibility cases plus native/browser per-tick and scripting comparisons.
Presentation changes retain software/WebGL and Firefox/WebKit coverage. Network
changes retain transport/server/deployment checks. Android changes retain arm64
builds and x86_64 emulator smoke. Shared headers, dependency/build configuration
and unknown paths select the full development matrix.
test/ci-compatibility.json owns repeated native compatibility cases; add suites
there when introducing a portability boundary.
Affected-PR tier reductions remain disabled until a full hosted master/nightly
matrix validates the current policy. CI_TIERED_COVERAGE_ENABLED=true activates
them with matching available full evidence; CI_TIER_BASELINE_RUN_ID may specify
a preferred baseline. Missing evidence or setting the flag false restores
conservative affected-PR coverage. These settings never reduce master coverage or
start expensive PR checks without an explicit request.