Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

proef in CI

Everything proef’s differentiators buy — typed exit codes, JUnit, sharding, the rerun overlay, the regression gate — pays off in CI, and each piece is documented on its own page. This page is the missing last mile: one paste-ready workflow, then the pieces it composes.

A complete GitHub Actions workflow

name: e2e
on: [push, pull_request]

jobs:
  e2e:
    runs-on: ubuntu-latest
    strategy:
      fail-fast: false
      matrix:
        shard: [1, 2, 3]
    steps:
      - uses: actions/checkout@v4
      - name: Install proef
        run: |
          curl -LsSf https://github.com/cargo-bins/cargo-binstall/releases/latest/download/cargo-binstall-x86_64-unknown-linux-musl.tgz | tar -xz -C /usr/local/bin
          cargo-binstall proef --no-confirm
      - name: Run the suite
        env:
          # Secrets reach proef only through the environment — never files,
          # never flags (values would land in the process listing).
          PROEF_SECRET_APITOKEN: ${{ secrets.API_TOKEN }}
        run: |
          proef test tests/features \
            --env staging \
            --shard ${{ matrix.shard }}/3 \
            --junit auto \
            --meta commit=${{ github.sha }}

The pieces:

  • --junit auto writes report.junit.xml into the run dir — under GITHUB_ACTIONS only, so local runs stay clean. Failures also reach the job summary and the PR’s changed-files gutter as ::error annotations automatically; no extra step.
  • --ctrf report.ctrf.json writes the same verdicts as CTRF JSON — the format that carries what JUnit XML cannot: retries and flakiness natively (a pass-after-retry lists its real failed attempts with messages), tags, and a file path per test. Point a CTRF consumer such as ctrf-io/github-test-reporter at it. The two files never disagree: a quarantined failure is skipped with a message in both (ADR-0019), because a dashboard reading “failed” beside exit 0 would contradict itself. Like --junit, the file is written even when [run] setup aborts the run — a job gating on it must never see it missing.
  • --shard I/N partitions by a stable hash of (file, scenario): adding a scenario never re-buckets the others, so shard timings stay comparable across commits. Every matrix job runs the same expression and the shards partition exactly. To balance by time instead of by count, see below.
  • --meta commit=… records provenance the run cannot harvest itself: proef never reads GITHUB_SHA or any CI variable (ADR-0020) — what the workflow hands over explicitly is what the record carries.
  • PROEF_SECRET_<NAME> supplies ${secret:name} values. They never appear in artifacts, events, logs, or reports — including base64/hex/ percent-encoded reflections.

Balancing the matrix by duration

A hash split balances by count, and a matrix finishes when its slowest shard does — so one long scenario can leave three runners idle. Every run that reaches its suite writes a small timings.json into its run directory; archive it once and hand it to the next run’s matrix:

      - name: Run the suite
        run: |
          proef test tests/features \
            --shard ${{ matrix.shard }}/3 --shard-weights timings.json
      # one job publishes the file the next run's matrix reads
      - uses: actions/upload-artifact@v4
        if: matrix.shard == 1
        with:
          name: proef-timings
          path: .proef-runs/*/timings.json

Every job must read the same file. The tempting shortcut — letting each job weight by its own local run history — is silently wrong: matrix jobs run on different machines with different (usually empty) runs-dirs, so each would compute a different assignment and scenarios would run twice or not at all while the suite still reported green.

A scenario the file does not mention falls back to the frozen hash, so a test added after the timings were captured still runs exactly once. A missing or malformed file is exit 2 — never a silent fall back to the unbalanced split. Full detail in CONFIG.md.

Gating on regressions between runs

proef diff compares two run records; --fail-on-regression makes it a gate. Download the base branch’s record (uploaded as an artifact by its own run) and compare:

proef test tests/features --junit auto        # today's run
proef diff path/to/base-events.jsonl          # vs the downloaded baseline
# exit 1 on a regression; new flakiness and perf deltas print either way

Continuing a cancelled run

A run stopped by --max-fail, a runner timeout, or a docker stop records what it never reached. SIGTERM and SIGHUP take the same graceful path as Ctrl-C: in-flight batches finish, the rest record as skipped, [run] teardown runs, the reports are written, and the record closes with run_finished + cancelled (exit 1) — so a timed-out job still leaves a complete record and its JUnit. Only a second signal (exit 130) or a SIGKILL truncates it.

proef test --rerun re-runs the last run’s failures and the scenarios it never got to — and its JUnit and HTML report cover the whole suite via the rerun overlay (rerun_of in the record), so one report stands for the composed result, never a false green.

Flakiness over history

Keep a few records ([run] keep-runs in proef.toml sets the rotation) and proef flaky renders verdicts over them — flapping, passes-only-on-retry, always-failing — from the same records CI already produced.

It also audits @quarantine itself, which nothing else can: a quarantined scenario failing every run is DISABLED (switched off — its failures gate nothing, so no job ever reports them), and one green throughout is recovered (the tag can come off). --format json carries quarantined and the verdict key, so a scheduled job can gate on either without parsing the table.

--by <key> splits the same history per run context — --by env, or any [meta]/--meta key such as --by runner. A scenario that flaps in one environment and is solid in another is not flaky but context-dependent, and a pooled history cannot tell those apart: it reports the one conclusion the merged view can never reach, naming the scenarios whose verdict changes with where they ran. A run that never set the key is its own (unset) bucket rather than being folded in with the runs that did.

Verdicts are keyed by the run’s input fingerprint — inputs.json beside each record, a hash of the feature sources, the loaded macros and fragments, and the resolved ${url:…}/${vars:…} scope — so a pack or config edit starts a fresh window instead of mixing runs of different inputs; --by splits within it. Three guards, each a [flaky] key with a flag twin (CONFIG.md): a scenario seen in fewer than min-samples runs (default 10) reads insufficient-data rather than earning a verdict — a fresh matrix needs ten records before the table says anything, and --min-samples 2 restores the pre-0.18 floor; a flagged scenario stays flagged until recovery-runs (5) trailing clean runs, so it cannot flip between adjacent runs; and a run in which more than outage-rate (0.8) of the suite failed is an environment incident, excluded wholesale, so one staging outage cannot mark the suite broken.