proef in CI
Everything proef’s differentiators buy — typed exit codes, JUnit, sharding, the rerun overlay, the regression gate — pays off in CI, and each piece is documented on its own page. This page is the missing last mile: one paste-ready workflow, then the pieces it composes.
A complete GitHub Actions workflow
name: e2e
on: [push, pull_request]
jobs:
e2e:
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
shard: [1, 2, 3]
steps:
- uses: actions/checkout@v4
- name: Install proef
run: |
curl -LsSf https://github.com/cargo-bins/cargo-binstall/releases/latest/download/cargo-binstall-x86_64-unknown-linux-musl.tgz | tar -xz -C /usr/local/bin
cargo-binstall proef --no-confirm
- name: Run the suite
env:
# Secrets reach proef only through the environment — never files,
# never flags (values would land in the process listing).
PROEF_SECRET_APITOKEN: ${{ secrets.API_TOKEN }}
run: |
proef test tests/features \
--env staging \
--shard ${{ matrix.shard }}/3 \
--junit auto \
--meta commit=${{ github.sha }}
The pieces:
--junit autowritesreport.junit.xmlinto the run dir — underGITHUB_ACTIONSonly, so local runs stay clean. Failures also reach the job summary and the PR’s changed-files gutter as::errorannotations automatically; no extra step.--ctrf report.ctrf.jsonwrites the same verdicts as CTRF JSON — the format that carries whatJUnitXML cannot: retries and flakiness natively (a pass-after-retry lists its real failed attempts with messages), tags, and a file path per test. Point a CTRF consumer such asctrf-io/github-test-reporterat it. The two files never disagree: a quarantined failure is skipped with a message in both (ADR-0019), because a dashboard reading “failed” beside exit 0 would contradict itself. Like--junit, the file is written even when[run] setupaborts the run — a job gating on it must never see it missing.--shard I/Npartitions by a stable hash of(file, scenario): adding a scenario never re-buckets the others, so shard timings stay comparable across commits. Every matrix job runs the same expression and the shards partition exactly. To balance by time instead of by count, see below.--meta commit=…records provenance the run cannot harvest itself: proef never readsGITHUB_SHAor any CI variable (ADR-0020) — what the workflow hands over explicitly is what the record carries.PROEF_SECRET_<NAME>supplies${secret:name}values. They never appear in artifacts, events, logs, or reports — including base64/hex/ percent-encoded reflections.
Balancing the matrix by duration
A hash split balances by count, and a matrix finishes when its slowest shard
does — so one long scenario can leave three runners idle. Every run that reaches
its suite writes a small timings.json into its run directory; archive it once
and hand it to the next run’s matrix:
- name: Run the suite
run: |
proef test tests/features \
--shard ${{ matrix.shard }}/3 --shard-weights timings.json
# one job publishes the file the next run's matrix reads
- uses: actions/upload-artifact@v4
if: matrix.shard == 1
with:
name: proef-timings
path: .proef-runs/*/timings.json
Every job must read the same file. The tempting shortcut — letting each job
weight by its own local run history — is silently wrong: matrix jobs run on
different machines with different (usually empty) runs-dirs, so each would
compute a different assignment and scenarios would run twice or not at all while
the suite still reported green.
A scenario the file does not mention falls back to the frozen hash, so a test added after the timings were captured still runs exactly once. A missing or malformed file is exit 2 — never a silent fall back to the unbalanced split. Full detail in CONFIG.md.
Gating on regressions between runs
proef diff compares two run records; --fail-on-regression makes it a
gate. Download the base branch’s record (uploaded as an artifact by its own
run) and compare:
proef test tests/features --junit auto # today's run
proef diff path/to/base-events.jsonl # vs the downloaded baseline
# exit 1 on a regression; new flakiness and perf deltas print either way
Continuing a cancelled run
A run stopped by --max-fail, a runner timeout, or a docker stop records what
it never reached. SIGTERM and SIGHUP take the same graceful path as Ctrl-C:
in-flight batches finish, the rest record as skipped, [run] teardown runs, the
reports are written, and the record closes with run_finished + cancelled
(exit 1) — so a timed-out job still leaves a complete record and its JUnit. Only
a second signal (exit 130) or a SIGKILL truncates it.
proef test --rerun re-runs the last run’s failures and the
scenarios it never got to — and its JUnit and HTML report cover the whole
suite via the rerun overlay (rerun_of in the record), so one report stands
for the composed result, never a false green.
Flakiness over history
Keep a few records ([run] keep-runs in proef.toml sets the rotation) and
proef flaky renders verdicts over them — flapping, passes-only-on-retry,
always-failing — from the same records CI already produced.
It also audits @quarantine itself, which nothing else can: a quarantined
scenario failing every run is DISABLED (switched off — its failures gate
nothing, so no job ever reports them), and one green throughout is recovered
(the tag can come off). --format json carries quarantined and the verdict
key, so a scheduled job can gate on either without parsing the table.
--by <key> splits the same history per run context — --by env, or any
[meta]/--meta key such as --by runner. A scenario that flaps in one
environment and is solid in another is not flaky but context-dependent, and a
pooled history cannot tell those apart: it reports the one conclusion the
merged view can never reach, naming the scenarios whose verdict changes with
where they ran. A run that never set the key is its own (unset) bucket rather
than being folded in with the runs that did.
Verdicts are keyed by the run’s input fingerprint — inputs.json beside
each record, a hash of the feature sources, the loaded macros and fragments,
and the resolved ${url:…}/${vars:…} scope — so a pack or config edit starts
a fresh window instead of mixing runs of different inputs; --by splits within
it. Three guards, each a [flaky] key with a flag twin (CONFIG.md): a
scenario seen in fewer than min-samples runs (default 10) reads
insufficient-data rather than earning a verdict — a fresh matrix needs ten
records before the table says anything, and --min-samples 2 restores the
pre-0.18 floor; a flagged scenario stays flagged until recovery-runs (5)
trailing clean runs, so it cannot flip between adjacent runs; and a run in
which more than outage-rate (0.8) of the suite failed is an environment
incident, excluded wholesale, so one staging outage cannot mark the suite
broken.