fix(tooling): a stale apr binary can no longer be resolved, hardcoded… #8447
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| # Sovereign CI — calls reusable workflow from paiml/.github | |
| # Change once in paiml/.github → applies to all repos | |
| # | |
| # Jobs provided by sovereign-ci.yml: | |
| # test: cargo test --lib (self-hosted clean-room) | |
| # lint: cargo clippy --all-targets -- -D warnings + cargo fmt --check | |
| # coverage: cargo llvm-cov + codecov upload | |
| # security: cargo audit (ubuntu-latest, continue-on-error) | |
| # provenance: SLSA attest-build-provenance | |
| # gate: aggregates test+lint results | |
| name: CI | |
| on: | |
| push: | |
| branches: [main, master] | |
| pull_request: | |
| branches: [main, master] | |
| # Merge queue: GitHub dispatches a `merge_group` event for each batch it forms. | |
| # Both required checks (`ci / gate` via the reusable sovereign-ci call, and the | |
| # top-level `workspace-test` job) live in THIS workflow, so triggering here makes | |
| # both run against the queued batch. Without this trigger, enabling the queue jams | |
| # every PR (no checks ever run on the merge_group ref → nothing merges). | |
| merge_group: | |
| workflow_dispatch: | |
| concurrency: | |
| group: ci-${{ github.event.pull_request.number || github.ref }} | |
| cancel-in-progress: true | |
| jobs: | |
| ci: | |
| uses: paiml/.github/.github/workflows/sovereign-ci.yml@main | |
| with: | |
| repo: ${{ github.event.repository.name }} | |
| # Phase 3 pilot (heavy workload) — build-performance.md §7 Phase 3. | |
| # APR-MONO monorepo: 879+ compile units, largest dep graph in the fleet. | |
| # Highest expected sccache hit-rate lift; without it each PR cold-compiles | |
| # in its per-PR-per-run target dir | |
| # (`/mnt/nvme-raid0/targets/aprender-ci/<PR>/run-<RUN_ID>`) | |
| # for ~34min, leaving only ~4min for tests inside the 40min timeout — | |
| # the entire merge queue saturates. | |
| # | |
| # 2026-05-15: target-dir path bumped from `aprender-ci/<PR>` to | |
| # `aprender-ci/<PR>/run-<RUN_ID>` to break the cancel-corrupt-state | |
| # race introduced by the prior per-PR fix (paiml/.github#31, | |
| # 2026-04-23). `concurrency.cancel-in-progress: true` + persistent | |
| # per-PR mount = the SIGTERM→SIGKILL window of a dying old cargo | |
| # corrupted `/workspace/target/debug/deps/` for the new run that | |
| # mounted the same host path. sccache stays on its own mount, so | |
| # cross-run cache effectiveness is preserved; only cargo-incremental | |
| # state (small fraction of total compile) is lost per new run. | |
| # | |
| # 2026-04-19: temporarily disabled — sovereign-ci:stable container image | |
| # was missing the `rustc-sccache` wrapper script. Fixed upstream in | |
| # paiml/infra commit f4fccf9 (PR #66, "use exec script not symlink"). | |
| # 2026-05-12: re-enabled — image verified to ship `/usr/local/bin/rustc-sccache` | |
| # (sccache 0.14.0), shared cache at `/home/noah/data/sccache` (warm, ~11GB). | |
| enable_sccache: true | |
| use_nextest: true | |
| # NOTE: coverage_min (the opt-in coverage ratchet from the build-system audit) | |
| # is intentionally NOT set on aprender. The pilot run exposed that aprender's | |
| # ROOT crate is a facade — the sovereign-ci coverage job runs `--lib` on the | |
| # root and exercises 0 tests ("test result: ok. 0 passed"); all real code + | |
| # tests live in workspace members, run by the separate `workspace-test` job. | |
| # So coverage_min has no lcov data to gate on. Making it meaningful here needs | |
| # test_workspace: true + GPU-member test_args exclusions first (the PMAT-159 | |
| # workspace blind-spot). Tracked as a follow-up. The coverage ratchet | |
| # MECHANISM is live fleet-wide via sovereign-ci (#37); a single-crate repo is | |
| # the natural first coverage pilot. aprender's blocking-quality pilot is the | |
| # diff-scoped mutation gate below. | |
| secrets: inherit | |
| # APR-MONO: Workspace-wide test (all 75 crates) | |
| # | |
| # 2026-05-13: Refactored from GH Actions `container:` syntax (which forces an | |
| # unconditional `docker pull` with only 3 retries / ~6s total backoff) to | |
| # explicit `docker run` steps with a 15-attempt linear-backoff pull retry. | |
| # The previous design conflated "image is required" with "registry must be | |
| # reachable at pull time" — when `localhost:5000` blipped (registry restart, | |
| # network reload), the pull failed and the whole job died after ~25s. This | |
| # refactor preserves the same execution semantics (same image, same volume | |
| # mounts, same env) but moves the pull into a step the workflow controls, | |
| # giving us up to ~13 minutes of retry headroom before declaring the | |
| # registry unreachable. Mirrored in the `mutants` job below. | |
| workspace-test: | |
| runs-on: [self-hosted, X64, Linux, clean-room] | |
| timeout-minutes: 85 # bumped to match 75min step + 10min overhead | |
| env: | |
| IMAGE: localhost:5000/sovereign-ci:stable | |
| PR_OR_REF: ${{ github.event.pull_request.number || github.ref_name }} | |
| steps: | |
| - name: Pre-checkout ownership restore (EACCES self-heal) | |
| # Five-whys: the end-of-job "Fix file ownership" step is | |
| # `if: always()`, but a hard-killed job (runner death, forced | |
| # cancel) skips even always() steps → root-owned | |
| # target/.rustc_info.json + target/package/ survive in the runner | |
| # checkout → the NEXT run's actions/checkout `git clean -ffdx` | |
| # fails with EACCES (observed 2026-07-02: 4 jobs across PRs | |
| # #2257/#2258 on runners 10/14/16, each needing a manual | |
| # `ssh intel sudo rm` sweep). Restoring ownership BEFORE checkout | |
| # makes every run self-healing instead of depending on the | |
| # previous run's clean exit. | |
| # Soundness of the cached-image gate: leftovers can only exist if | |
| # a previous docker job ran on this runner — which implies the | |
| # image is already in the local cache. So "image not cached ⟹ no | |
| # leftovers" and skipping is safe (also keeps this step | |
| # registry-outage-tolerant). | |
| run: | | |
| if docker image inspect "$IMAGE" > /dev/null 2>&1; then | |
| docker run --rm -v "${GITHUB_WORKSPACE}:/workspace" "$IMAGE" \ | |
| bash -c 'chown -R 1000:1000 /workspace 2>/dev/null || true' | |
| else | |
| echo "Image not cached — no prior docker job on this runner, nothing to restore" | |
| fi | |
| - uses: actions/checkout@v7 | |
| - name: Pull sovereign-ci image (with retry + local-cache fallback) | |
| # Self-hosted runner's local Docker registry at localhost:5000 is | |
| # occasionally restarting OR experiencing extended outages | |
| # (paiml/infra ops). Two layers of resilience: | |
| # 1. Check the local Docker daemon cache first — the image was | |
| # successfully pulled on a prior run, so it's almost certainly | |
| # still in the cache (Docker doesn't GC images unless `prune` | |
| # is run). If present, skip the pull entirely; this makes the | |
| # workflow registry-outage-tolerant. | |
| # 2. Otherwise, try to pull with 15-attempt linear-backoff retry | |
| # (~13min total) — plenty for any normal restart cycle. | |
| # The local-cache path accepts slight staleness as the price of | |
| # registry-outage tolerance. paiml/infra:machines/intel/sovereign- | |
| # ci/rebuild.sh rebuilds the stable tag nightly so any drift gets | |
| # corrected within 24h on the next successful pull. | |
| run: | | |
| if docker image inspect "$IMAGE" > /dev/null 2>&1; then | |
| echo "Image $IMAGE already cached locally — skipping pull" | |
| echo "(local cache is registry-outage-tolerant; nightly rebuild keeps it fresh)" | |
| exit 0 | |
| fi | |
| max_attempts=15 | |
| delay=4 | |
| for i in $(seq 1 $max_attempts); do | |
| if docker pull "$IMAGE"; then | |
| echo "Image pulled successfully on attempt $i" | |
| exit 0 | |
| fi | |
| if [ $i -eq $max_attempts ]; then | |
| echo "::error::Registry localhost:5000 unreachable after $max_attempts attempts (~13min) AND image not in local cache" | |
| echo "::error::Suggests paiml/infra runner-side registry restart + initial image seed needed" | |
| exit 1 | |
| fi | |
| echo "Pull attempt $i/$max_attempts failed; sleeping ${delay}s" | |
| sleep "$delay" | |
| delay=$((delay + 6)) # linear backoff: 4,10,16,22,28,34,... | |
| done | |
| - name: Pre-flight target-dir consistency check | |
| # Root cause (five-whys): | |
| # 1. Why does workspace-test sometimes fail with "no such file or | |
| # directory .rcgu.o" / extern location missing / cc-rs can't | |
| # create .o? Cargo's incremental state on the per-PR target | |
| # dir is inconsistent. | |
| # 2. Why inconsistent? A prior run was SIGKILL'd mid-compile and | |
| # left orphan .rmeta files (parts cargo had registered as built) | |
| # without the corresponding .rcgu.o codegen artifacts (which were | |
| # mid-write at the moment of the kill). | |
| # 3. Why was it SIGKILL'd? concurrency.cancel-in-progress (line 22) | |
| # cancels the previous run as soon as a new commit lands on the | |
| # branch (and "Update branch" / strict-up-to-date triggers this | |
| # every time aprender main moves forward). | |
| # 4. Why does this persist? The target dir is bind-mounted from a | |
| # per-PR persistent path /mnt/nvme-raid0/targets/aprender-ci/<PR>/, | |
| # so partial-compile state survives across runs. | |
| # 5. Root cause: cargo's incremental state is not atomic-on-kill, so | |
| # a persistent shared target dir + cancel-in-progress = damage. | |
| # Prevention (this step): BEFORE invoking cargo, check whether the | |
| # immediately-preceding workflow run on this branch was cancelled. If | |
| # yes, rm -rf the target dir contents. This is a one-time check at | |
| # job start — NOT a retry-on-failure pattern (which the operator | |
| # rejects under the "flake is not allowed" directive). | |
| run: | | |
| set -e | |
| if [ -z "${{ github.event.pull_request.number }}" ]; then | |
| echo "Not a PR run; skipping prior-cancel check" | |
| exit 0 | |
| fi | |
| # Find the immediately-preceding workflow run on this branch. | |
| # status=completed filter excludes the current in-progress run. | |
| PREV_CONCLUSION=$(gh api \ | |
| "repos/${GITHUB_REPOSITORY}/actions/runs?branch=${GITHUB_HEAD_REF}&status=completed&per_page=1" \ | |
| --jq '.workflow_runs[0].conclusion' 2>/dev/null || echo "") | |
| echo "Previous run conclusion on ${GITHUB_HEAD_REF}: ${PREV_CONCLUSION:-<none>}" | |
| if [ "$PREV_CONCLUSION" = "cancelled" ]; then | |
| echo "::warning::Previous run was cancelled; nuking target dir to prevent cargo cancel-damage" | |
| docker run --rm \ | |
| -v "/mnt/nvme-raid0/targets/aprender-ci/${PR_OR_REF}:/workspace/target" \ | |
| "$IMAGE" \ | |
| bash -c 'rm -rf /workspace/target/* /workspace/target/.[!.]* 2>/dev/null || true; ls -la /workspace/target/ || true' | |
| else | |
| echo "No cancel damage to clean (prior conclusion: ${PREV_CONCLUSION:-fresh-branch})" | |
| fi | |
| env: | |
| GH_TOKEN: ${{ github.token }} | |
| - name: Pre-build chown — fix per-RUN root ownership | |
| # Root cause (five-whys): | |
| # 1. Why do fresh runs sometimes fail with "failed to create | |
| # /workspace/target/debug" / "No such file or directory"? | |
| # Cargo (running as user 1000 inside the container) can't | |
| # write to /workspace/target/debug. | |
| # 2. Why can't it write? The bind-mount source dir on the host | |
| # (/mnt/nvme-raid0/targets/aprender-ci/<PR>/run-<RUN_ID>) is | |
| # owned by root:root. | |
| # 3. Why is it root-owned? Docker's bind-mount creates missing | |
| # host directories with the daemon's uid (root). Per-RUN | |
| # paths are always fresh, so this fires every run. | |
| # 4. Why didn't this happen before? Pre-#1693 the per-PR (not | |
| # per-RUN) path persisted across runs, and the downstream | |
| # "post-job cleanup" docker chown step fixed ownership for | |
| # the NEXT run's git-clean. Per-RUN paths invalidate that — | |
| # each run gets a brand-new root-owned dir. | |
| # 5. Root cause: the chown step runs AFTER cargo, not BEFORE. | |
| # First-runs always fail; reruns appear to work only when | |
| # the previous run's belated chown fixed the now-stale dir. | |
| # Fix (this step): docker run as root, chown the per-RUN target | |
| # dir + cargo registry to noah:1000 BEFORE the cargo step. | |
| # Idempotent — `|| true` tolerates dirs that are already | |
| # noah-owned (e.g. a rerun of the same run-id). | |
| run: | | |
| docker run --rm \ | |
| -v "/mnt/nvme-raid0/targets/aprender-ci/${PR_OR_REF}/run-${GITHUB_RUN_ID}:/workspace/target" \ | |
| -v "/mnt/nvme-raid0/cargo-ci/registry/${PR_OR_REF}:/usr/local/cargo/registry" \ | |
| "$IMAGE" \ | |
| bash -c 'chown -R 1000:1000 /workspace/target /usr/local/cargo/registry 2>/dev/null || true' | |
| - name: Workspace lib tests (25,300+) | |
| # Excluded: aprender-gpu (cuBLAS), aprender-cuda-edge (CUDA), aprender-compute (SIMD SIGSEGV at exit) | |
| # Timeout: 75min (was 55, was 40). | |
| # 2026-05-15: bumped to 75min after the P0 per-run target-dir fix | |
| # (#1693) eliminated cargo-incremental cross-run warmth. Cold | |
| # compiles now happen on every run; sccache covers codegen | |
| # (~80% hit rate on warm cache) but cargo's metadata + linking + | |
| # test binaries still cost ~40-50min cold. Under runner-pool | |
| # saturation (7+ concurrent CI runs) we observed 55min hits | |
| # exactly at the timeout — runs 25919246467 / 25919258460 / | |
| # sibling PRs failed simultaneously at 55:00.0. | |
| # | |
| # 2026-06-24 (perf/ci-nextest, EXPERIMENT): switched from | |
| # `cargo test --workspace --lib` to `cargo nextest run`. Profiling the | |
| # ~40min job showed ~80% is SERIAL test EXECUTION (cargo's libtest runs | |
| # one test binary at a time; ~27min of the wall clock is just running | |
| # tests), not codegen (sccache already covers that). nextest runs every | |
| # test in its own process across a thread pool, so the ~25,300 tests | |
| # execute in PARALLEL across crates — the big lever. The exact | |
| # `--exclude` list is preserved verbatim (aprender-compute stays | |
| # excluded; its SIGSEGV-at-exit is handled by the dedicated "Compute | |
| # tests" step below). `--profile ci` uses .config/nextest.toml | |
| # ([profile.ci]: retries=0, fail-fast). line-tables-only debuginfo now | |
| # comes from root Cargo.toml [profile.test]; the CARGO_PROFILE_*_DEBUG | |
| # env vars are kept as belt-and-suspenders. | |
| timeout-minutes: 75 | |
| run: | | |
| docker run --rm \ | |
| -e CI -e GITHUB_ACTIONS -e GITHUB_REF -e GITHUB_SHA -e GITHUB_REPOSITORY -e GITHUB_RUN_ID -e GITHUB_EVENT_NAME -e GITHUB_WORKFLOW \ | |
| -v "${GITHUB_WORKSPACE}:/workspace" \ | |
| -v "/mnt/nvme-raid0/cargo-ci/registry/${PR_OR_REF}:/usr/local/cargo/registry" \ | |
| -v "/mnt/nvme-raid0/targets/aprender-ci/${PR_OR_REF}/run-${GITHUB_RUN_ID}:/workspace/target" \ | |
| -v "/home/noah/data/sccache:/sccache" \ | |
| -w /workspace \ | |
| -e CARGO_TARGET_DIR=/workspace/target \ | |
| -e RUSTC_WRAPPER=rustc-sccache \ | |
| -e SCCACHE_DIR=/sccache \ | |
| -e CARGO_INCREMENTAL=0 \ | |
| -e CARGO_BUILD_JOBS=8 \ | |
| -e CARGO_PROFILE_TEST_DEBUG=line-tables-only \ | |
| -e CARGO_PROFILE_DEV_DEBUG=line-tables-only \ | |
| "$IMAGE" \ | |
| cargo nextest run --profile ci --workspace --lib --exclude aprender-gpu --exclude aprender-cuda-edge --exclude aprender-compute | |
| - name: Compute tests (tolerate SIGSEGV at exit — all tests pass but harness crashes on cleanup) | |
| run: | | |
| docker run --rm \ | |
| -e CI -e GITHUB_ACTIONS -e GITHUB_REF -e GITHUB_SHA -e GITHUB_REPOSITORY -e GITHUB_RUN_ID -e GITHUB_EVENT_NAME -e GITHUB_WORKFLOW \ | |
| -v "${GITHUB_WORKSPACE}:/workspace" \ | |
| -v "/mnt/nvme-raid0/cargo-ci/registry/${PR_OR_REF}:/usr/local/cargo/registry" \ | |
| -v "/mnt/nvme-raid0/targets/aprender-ci/${PR_OR_REF}/run-${GITHUB_RUN_ID}:/workspace/target" \ | |
| -v "/home/noah/data/sccache:/sccache" \ | |
| -w /workspace \ | |
| -e CARGO_TARGET_DIR=/workspace/target \ | |
| -e RUSTC_WRAPPER=rustc-sccache \ | |
| -e SCCACHE_DIR=/sccache \ | |
| -e CARGO_INCREMENTAL=0 \ | |
| -e CARGO_BUILD_JOBS=8 \ | |
| "$IMAGE" \ | |
| bash -c 'cargo test -p aprender-compute --lib 2>&1 | tee /tmp/compute-test.log; grep -q "test result: ok\." /tmp/compute-test.log && ! grep -q "test result: FAILED" /tmp/compute-test.log' | |
| - name: Integration tests | |
| # perf/ci-nextest (rank 3 — integration collapse): INTENTIONALLY SKIPPED. | |
| # The investigation ranked collapsing these 8 `cargo test -p X --test Y` | |
| # invocations into ONE `cargo nextest run -E '...'` as rank 3 (low lever: | |
| # this step is NOT the ~27min bottleneck — the lib step is). It also | |
| # carries a real risk: this chain includes | |
| # `cargo test -p aprender-compute --lib beat_nf4_bitsandbytes_equivalence`, | |
| # and aprender-compute SIGSEGVs at harness EXIT. Under cargo's libtest a | |
| # single-name filter exits cleanly here, but nextest's per-test process | |
| # model would observe the segfaulting process exit code and could fail | |
| # the run. Per the experiment brief ("if fiddly/risky, SKIP — don't block | |
| # the experiment"), this step is left AS-IS so the nextest signal stays | |
| # attributable to the lib step alone. Revisit after the lib-step | |
| # measurement lands. | |
| run: | | |
| docker run --rm \ | |
| -e CI -e GITHUB_ACTIONS -e GITHUB_REF -e GITHUB_SHA -e GITHUB_REPOSITORY -e GITHUB_RUN_ID -e GITHUB_EVENT_NAME -e GITHUB_WORKFLOW \ | |
| -v "${GITHUB_WORKSPACE}:/workspace" \ | |
| -v "/mnt/nvme-raid0/cargo-ci/registry/${PR_OR_REF}:/usr/local/cargo/registry" \ | |
| -v "/mnt/nvme-raid0/targets/aprender-ci/${PR_OR_REF}/run-${GITHUB_RUN_ID}:/workspace/target" \ | |
| -v "/home/noah/data/sccache:/sccache" \ | |
| -w /workspace \ | |
| -e CARGO_TARGET_DIR=/workspace/target \ | |
| -e RUSTC_WRAPPER=rustc-sccache \ | |
| -e SCCACHE_DIR=/sccache \ | |
| -e CARGO_INCREMENTAL=0 \ | |
| -e CARGO_BUILD_JOBS=8 \ | |
| "$IMAGE" \ | |
| bash -c 'cargo test -p aprender-core --test monorepo_invariants && cargo test -p aprender-core --test readme_contract && cargo test -p apr-cli --test cli_commands && cargo test -p aprender-core --test beat_sklearn_iris && cargo test -p aprender-core --test beat_sklearn_nmi && cargo test -p aprender-core --test beat_sklearn_metrics_parity && cargo test -p aprender-core --test beat_sklearn_gaussiannb_accuracy && cargo test -p aprender-core --test beat_sklearn_svc_accuracy && cargo test -p aprender-core --test beat_sklearn_pipeline_encoder && cargo test -p aprender-serve --test beat_fail_closed_garbage && cargo test -p aprender-compute --lib beat_nf4_bitsandbytes_equivalence && cargo test -p aprender-core --test beat_pytorch_autograd_grad && cargo test -p aprender-train-lora --lib beat_lora_merge_forward_equivalence && cargo test -p apr-cli --release --test beat_pytorch_deploy_footprint && cargo test -p aprender-serve --test beat_fail_closed_structural && cargo test -p aprender-serve --test ollama_http_compat && cargo test -p apr-cli --test ollama_ndjson_streaming && cargo test -p apr-cli --test falsification_chat_http_cli' | |
| - name: Build.rs crate-root escape check (v0.31.1 yank guard) | |
| # Static Poka-Yoke: flags build.rs files that panic on files outside | |
| # CARGO_MANIFEST_DIR, which break `cargo install` from crates.io. | |
| # See scripts/check_build_rs_paths.sh for the full rationale. | |
| run: | | |
| docker run --rm \ | |
| -e CI -e GITHUB_ACTIONS -e GITHUB_REF -e GITHUB_SHA -e GITHUB_REPOSITORY -e GITHUB_RUN_ID -e GITHUB_EVENT_NAME -e GITHUB_WORKFLOW \ | |
| -v "${GITHUB_WORKSPACE}:/workspace" \ | |
| -w /workspace \ | |
| "$IMAGE" \ | |
| bash scripts/check_build_rs_paths.sh | |
| - name: apr-format leaf sovereignty guard (#2231) | |
| # Poka-Yoke: prove the extracted `apr-format` leaf pulls no ML/GPU/ | |
| # tokenizer/framework crate (so consumers `cargo add apr-format` without | |
| # aprender-core + trueno/wgpu). Discriminating: PASSES on apr-format + | |
| # aprender-quant, FAILS on aprender-core. Also runs a publish dry-run to | |
| # catch dev-dep cycles. See scripts/check_format_sovereignty.sh. | |
| run: | | |
| docker run --rm \ | |
| -e CI -e GITHUB_ACTIONS -e GITHUB_REF -e GITHUB_SHA -e GITHUB_REPOSITORY -e GITHUB_RUN_ID -e GITHUB_EVENT_NAME -e GITHUB_WORKFLOW \ | |
| -v "${GITHUB_WORKSPACE}:/workspace" \ | |
| -w /workspace \ | |
| "$IMAGE" \ | |
| bash scripts/check_format_sovereignty.sh | |
| - name: Fix file ownership (container runs as root, runner as noah:1000) | |
| if: always() | |
| run: | | |
| # Five-whys: Docker container creates files as root on bind-mounted | |
| # workspace. Runner (noah:1000) can't git-clean them on next run | |
| # → checkout fails → CI breaks. This runs inside the container | |
| # (as root) restoring host ownership for subsequent bare-metal jobs. | |
| docker run --rm \ | |
| -e CI -e GITHUB_ACTIONS -e GITHUB_REF -e GITHUB_SHA -e GITHUB_REPOSITORY -e GITHUB_RUN_ID -e GITHUB_EVENT_NAME -e GITHUB_WORKFLOW \ | |
| -v "${GITHUB_WORKSPACE}:/workspace" \ | |
| -v "/mnt/nvme-raid0/cargo-ci/registry/${PR_OR_REF}:/usr/local/cargo/registry" \ | |
| -v "/mnt/nvme-raid0/targets/aprender-ci/${PR_OR_REF}/run-${GITHUB_RUN_ID}:/workspace/target" \ | |
| "$IMAGE" \ | |
| bash -c 'chown -R 1000:1000 /workspace || true; chown -R 1000:1000 /usr/local/cargo/registry || true; chown -R 1000:1000 /workspace/target || true' | |
| # Poka-yoke (aprender#2269): fail fast if ANY self-hosted job omits a | |
| # discriminating runner label, so a bare [self-hosted, X64, Linux] selector | |
| # can never again silently land a job on a GPU/dev runner without the | |
| # sovereign-ci registry. Pure text check; runs on the clean-room pool. | |
| guard-runner-labels: | |
| runs-on: [self-hosted, X64, Linux, clean-room] | |
| env: | |
| IMAGE: localhost:5000/sovereign-ci:stable | |
| steps: | |
| - name: Pre-checkout ownership restore (EACCES self-heal) | |
| # Same guard as workspace-test (:92) and mutants (:435). This job was | |
| # the ONLY clean-room job missing it, and that asymmetry took main red | |
| # on 2026-07-06 (run 28776891625): a hard-killed docker job left | |
| # root-owned target/.rustc_info.json + target/package/ in the runner | |
| # checkout, so this job's actions/checkout `git clean -ffdx` failed | |
| # with EACCES before a single step ran. Because `gate` hard-requires | |
| # guard-runner-labels (:379), a checkout-level EACCES here reads as a | |
| # required-check failure and blocks every merge — the andon stays red | |
| # until someone manually `ssh intel sudo rm`s the leftovers. | |
| # Cached-image gate is sound: leftovers can only exist if a previous | |
| # docker job ran on this runner, which implies the image is already in | |
| # the local cache. So "image not cached ⟹ no leftovers" and skipping is | |
| # safe (and keeps this step tolerant of a registry outage). | |
| run: | | |
| if docker image inspect "$IMAGE" > /dev/null 2>&1; then | |
| docker run --rm -v "${GITHUB_WORKSPACE}:/workspace" "$IMAGE" \ | |
| bash -c 'chown -R 1000:1000 /workspace 2>/dev/null || true' | |
| else | |
| echo "Image not cached — no prior docker job on this runner, nothing to restore" | |
| fi | |
| - uses: actions/checkout@v7 | |
| with: | |
| fetch-depth: 1 | |
| - name: Every self-hosted job must pin a discriminating label | |
| run: bash scripts/check_runner_labels.sh | |
| # Poka-yoke: a beat that no workflow executes reads as enforcement, is | |
| # counted as enforcement, and proves nothing. The Pillar-4 marquee decode | |
| # beat sat in ZERO workflows while being quoted as an enforced win (#2319). | |
| # Pure text check, no build, so it belongs in this job rather than paying | |
| # for its own runner. | |
| - name: Every beat must be executed by some workflow | |
| run: bash scripts/check_beats_gated.sh | |
| # Poka-yoke: PMAT-CI-PASSGREP-001 killed ONE `grep "0 failed"` that also | |
| # matched "10 failed". Two more instances of the same class survived that | |
| # fix, one of them a live contract falsifier. Probe every zero-count | |
| # pass-grep against an all-failing line instead of trusting review. | |
| - name: Every zero-count pass-grep must reject a failing line | |
| run: bash scripts/check_pass_grep_anchored.sh | |
| # Poka-yoke: a bare `apr` runs whatever PATH resolves. qwen-story-daily | |
| # installed 0.61.0 to ~/.cargo/bin and then executed a 24-day-old 0.60.0 | |
| # from ~/.local/bin, so every beat validated stale code while reporting | |
| # green. Text-only check, no build. | |
| - name: Every CI-surface `apr` invocation must be pinned | |
| run: bash scripts/check_apr_bin_pinned.sh | |
| # Poka-yoke: `set` in a SOURCED file mutates the caller's shell. apr_bin.sh | |
| # opened with `set -euo pipefail`; qwen-story.sh sources it and had chosen | |
| # `set -uo pipefail` deliberately (it must run every beat and tally the | |
| # failures). The leak turned errexit on underneath it and the nightly | |
| # story died after six lines inside an ADVISORY pmat hunt. Both files read | |
| # correctly on their own - only the combination is wrong, which is why | |
| # this is mechanical rather than a review note. Text-only, no build. | |
| - name: Sourced libraries must not mutate the caller's shell options | |
| run: bash scripts/check_sourced_libs_option_neutral.sh --self-test | |
| # Top-level gate: satisfies org ruleset "Green Main" which requires check named "gate". | |
| # The reusable workflow produces "ci / gate" but rulesets need exact match on "gate". | |
| gate: | |
| runs-on: [self-hosted, X64, Linux, clean-room] | |
| needs: [ci, workspace-test, mutants, guard-runner-labels] | |
| if: always() | |
| steps: | |
| - name: Check required jobs | |
| run: | | |
| if [ "${{ needs.guard-runner-labels.result }}" != "success" ]; then | |
| echo "guard-runner-labels failed: ${{ needs.guard-runner-labels.result }}" | |
| exit 1 | |
| fi | |
| if [ "${{ needs.ci.result }}" != "success" ]; then | |
| echo "ci failed: ${{ needs.ci.result }}" | |
| exit 1 | |
| fi | |
| if [ "${{ needs.workspace-test.result }}" != "success" ]; then | |
| echo "workspace-test failed: ${{ needs.workspace-test.result }}" | |
| exit 1 | |
| fi | |
| # Diff-scoped mutation gate (PMAT gap #1): blocking on PRs. | |
| # `skipped` is the expected result on push-to-main (the job has | |
| # `if: github.event_name == 'pull_request'`); treat it as pass so | |
| # main-branch pushes are not blocked by a job that intentionally | |
| # did not run. Only an explicit `failure` blocks. | |
| MUT="${{ needs.mutants.result }}" | |
| if [ "$MUT" = "failure" ]; then | |
| echo "mutants (diff-scoped mutation) failed: $MUT" | |
| exit 1 | |
| fi | |
| echo "mutants result: $MUT (success/skipped both pass)" | |
| echo "All required jobs passed" | |
| # Mutation testing — DIFF-SCOPED + BLOCKING on PRs (PMAT build-system audit gap #1). | |
| # | |
| # BEFORE: full-tree `cargo mutants -- --lib`, push-to-main only, and | |
| # `continue-on-error: true` at BOTH the job and step level → a surviving | |
| # mutant never blocked anything. New under-tested code merged silently, | |
| # contradicting the 80%-mutation / ZERO-tolerance rule. | |
| # | |
| # AFTER: scope mutation to the PR DIFF (`cargo mutants --in-diff`), run it on | |
| # pull_request events, and make it BLOCKING (no continue-on-error; wired into | |
| # the `gate` job). Diff-scoping is the key lever: full-tree mutation on a | |
| # 75-crate monorepo is hours-long and would choke the merge queue. Gating only | |
| # the lines a PR actually touches keeps it fast (minutes, proportional to diff | |
| # size) while still preventing NEW untested code from landing. A PR whose diff | |
| # contains no mutable code is a clean no-op pass (cargo-mutants reports 0 | |
| # mutants → exit 0). | |
| # | |
| # On a push to main (post-merge), the job is a no-op pass: there is no PR diff | |
| # to scope against, so we skip rather than fall back to the old hours-long | |
| # full-tree run. | |
| mutants: | |
| runs-on: [self-hosted, X64, Linux, clean-room] | |
| timeout-minutes: 60 | |
| needs: [ci, workspace-test] | |
| if: github.event_name == 'pull_request' | |
| env: | |
| IMAGE: localhost:5000/sovereign-ci:stable | |
| # Max surviving (missed) mutants tolerated on the PR diff. 0 = every | |
| # mutant introduced/touched by this PR must be caught by a test. Tune up | |
| # via repo variable MUTANTS_MAX_MISSED if a diff legitimately can't reach 0. | |
| MUTANTS_MAX_MISSED: ${{ vars.MUTANTS_MAX_MISSED || '0' }} | |
| steps: | |
| - name: Pre-checkout ownership restore (EACCES self-heal) | |
| # Same guard as workspace-test: a hard-killed docker job leaves | |
| # root-owned files that EACCES this job's `git clean` at checkout | |
| # (observed 2026-07-02 — the mutants job failed twice at checkout | |
| # with zero mutants actually run, reading as a gate failure). | |
| # Cached-image gate is sound: leftovers ⟹ a docker job ran here | |
| # ⟹ image is cached. | |
| run: | | |
| if docker image inspect "$IMAGE" > /dev/null 2>&1; then | |
| docker run --rm -v "${GITHUB_WORKSPACE}:/workspace" "$IMAGE" \ | |
| bash -c 'chown -R 1000:1000 /workspace 2>/dev/null || true' | |
| else | |
| echo "Image not cached — no prior docker job on this runner, nothing to restore" | |
| fi | |
| - uses: actions/checkout@v7 | |
| with: | |
| # Need history + base branch to compute the PR diff for --in-diff. | |
| fetch-depth: 0 | |
| - name: Compute PR diff for mutation scoping | |
| # cargo-mutants --in-diff takes a unified diff and mutates ONLY the | |
| # lines it adds/changes. We diff the PR head against the merge-base with | |
| # the target branch so the scope is exactly "what this PR introduces". | |
| run: | | |
| set -euo pipefail | |
| BASE_REF="${{ github.event.pull_request.base.ref }}" | |
| # NOT --depth=1: a shallow base commit has no shared ancestor with the | |
| # fetch-depth:0 PR head, so `git merge-base` finds none and (under | |
| # `set -e`) the step dies — failing the mutants gate on every PR. | |
| # Fetch the base branch's history so the merge-base is reachable. | |
| git fetch --no-tags origin "$BASE_REF" | |
| MERGE_BASE=$(git merge-base HEAD "origin/$BASE_REF" 2>/dev/null || true) | |
| if [ -z "$MERGE_BASE" ]; then | |
| echo "No merge-base with origin/$BASE_REF — scoping diff to base tip" | |
| MERGE_BASE="origin/$BASE_REF" | |
| fi | |
| echo "Base ref: $BASE_REF merge-base: $MERGE_BASE" | |
| git diff "$MERGE_BASE"...HEAD > pr.diff | |
| echo "Diff size: $(wc -l < pr.diff) lines" | |
| if [ ! -s pr.diff ]; then | |
| echo "Empty diff — no code to mutate" | |
| fi | |
| - name: Pull sovereign-ci image (with retry + local-cache fallback) | |
| # Same two-layer resilience as workspace-test — see that job for full context. | |
| run: | | |
| if docker image inspect "$IMAGE" > /dev/null 2>&1; then | |
| echo "Image $IMAGE already cached locally — skipping pull" | |
| exit 0 | |
| fi | |
| max_attempts=15 | |
| delay=4 | |
| for i in $(seq 1 $max_attempts); do | |
| if docker pull "$IMAGE"; then | |
| echo "Image pulled successfully on attempt $i" | |
| exit 0 | |
| fi | |
| if [ $i -eq $max_attempts ]; then | |
| echo "::error::Registry localhost:5000 unreachable after $max_attempts attempts AND image not in local cache" | |
| exit 1 | |
| fi | |
| echo "Pull attempt $i/$max_attempts failed; sleeping ${delay}s" | |
| sleep "$delay" | |
| delay=$((delay + 6)) | |
| done | |
| - name: Install cargo-mutants | |
| run: | | |
| docker run --rm \ | |
| -e CI -e GITHUB_ACTIONS -e GITHUB_REF -e GITHUB_SHA -e GITHUB_REPOSITORY -e GITHUB_RUN_ID -e GITHUB_EVENT_NAME -e GITHUB_WORKFLOW \ | |
| -v "${GITHUB_WORKSPACE}:/workspace" \ | |
| -w /workspace \ | |
| "$IMAGE" \ | |
| cargo install cargo-mutants --locked | |
| - name: Run diff-scoped mutation testing (BLOCKING) | |
| # No continue-on-error: a missed mutant on the PR diff fails the job, | |
| # which fails `gate`, which blocks merge. --in-diff pr.diff restricts | |
| # mutation to PR-touched lines. Empty diff → 0 mutants → clean pass. | |
| # We parse mutants.out/outcomes.json for the missed count and compare to | |
| # MUTANTS_MAX_MISSED so the threshold is explicit and tunable (rather | |
| # than relying solely on cargo-mutants' aggregate exit code). | |
| run: | | |
| set -euo pipefail | |
| if [ ! -s pr.diff ]; then | |
| echo "No PR diff content — nothing to mutate. Pass." | |
| exit 0 | |
| fi | |
| docker run --rm \ | |
| -e CI -e GITHUB_ACTIONS -e GITHUB_REF -e GITHUB_SHA -e GITHUB_REPOSITORY -e GITHUB_RUN_ID -e GITHUB_EVENT_NAME -e GITHUB_WORKFLOW \ | |
| -v "${GITHUB_WORKSPACE}:/workspace" \ | |
| -w /workspace \ | |
| -e MUTANTS_MAX_MISSED \ | |
| "$IMAGE" \ | |
| bash -c ' | |
| set -uo pipefail | |
| # --in-diff pr.diff: mutate only PR-touched lines. | |
| # cargo-mutants exits non-zero when mutants survive; we still | |
| # parse outcomes.json so the threshold (MUTANTS_MAX_MISSED) is | |
| # explicit and the failure message is actionable. | |
| cargo mutants --no-times --timeout 300 --in-place \ | |
| --in-diff pr.diff -- --lib | |
| MUT_EXIT=$? | |
| echo "cargo-mutants exit: $MUT_EXIT" | |
| OUTCOMES=mutants.out/outcomes.json | |
| if [ ! -f "$OUTCOMES" ]; then | |
| # No outcomes file means cargo-mutants found no mutants in the | |
| # diff (e.g. diff only touched non-Rust / non-mutable lines). | |
| echo "No mutants.out/outcomes.json — 0 mutants in diff. Pass." | |
| exit 0 | |
| fi | |
| MISSED=$(grep -o "\"summary\"[^}]*\"missed\":[0-9]*" "$OUTCOMES" \ | |
| | grep -o "\"missed\":[0-9]*" | grep -o "[0-9]*" | head -1) | |
| TIMEOUT=$(grep -o "\"timeout\":[0-9]*" "$OUTCOMES" \ | |
| | grep -o "[0-9]*" | head -1) | |
| MISSED=${MISSED:-0}; TIMEOUT=${TIMEOUT:-0} | |
| echo "Diff-scoped mutation result: missed=$MISSED timeout=$TIMEOUT (max allowed missed=$MUTANTS_MAX_MISSED)" | |
| UNCAUGHT=$((MISSED + TIMEOUT)) | |
| if [ "$UNCAUGHT" -gt "$MUTANTS_MAX_MISSED" ]; then | |
| echo "::error::$UNCAUGHT mutant(s) survived/timed-out on the PR diff (> $MUTANTS_MAX_MISSED allowed). New code is under-tested — add tests that kill these mutants. This would have merged SILENTLY before (PMAT gap #1)." | |
| exit 1 | |
| fi | |
| echo "All diff-scoped mutants caught (or within threshold). Pass." | |
| exit 0 | |
| ' | |
| - name: Upload mutation results | |
| if: always() | |
| uses: actions/upload-artifact@v7 | |
| with: | |
| name: mutation-results | |
| path: mutants.out/ |