Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

proef — documentation index

proef (Dutch: test, trial — and tasting) is a declarative, modular, multi-engine end-to-end test runner, mostly Rust. Tests are Gherkin .feature files in business prose; macro packs bind the prose to executable steps; a pluggable engine runs each step batch — embedded Hurl for API testing (the seam admits future engines; none are scheduled).

This folder is the project corpus: the product requirements, the decision log, the normative technical spec, the milestone plan, and the testing strategy. Written 2026-07-28 from a validated research round (a working spike ran 5/5 scenarios green under both a prototype native runner and stock hurl 8.0.1 on identical generated artifacts); implementation has since delivered milestones M0–M5 and everything after them — through v0.19.0 (the 0.18 survey waves, then the gates that could not see what they covered). Only M6 (a second engine) is unscheduled. The repo-root CLAUDE.md carries the live status. This corpus is also published as a website: https://emrecdr.github.io/proef/.

Reading order

#DocumentWhat it answersAudience
0WRITING-SCENARIOS.mdWrite prose against a vocabulary somebody else maintains: see the sentences, the dry-run loop, the two errors you will hitP1 test authors
0GETTING-STARTED.mdYour first suite in ten minutes — including the packs behind itP2 pack maintainers
0AUTHORING.mdThe pack/feature reference from the author’s seatP2 pack maintainers
0EDITORS.mdWiring proef lsp into Neovim/Helix/Emacs for live diagnostics, jump-to-macro, completionP1/P2, whoever sets up the editor
0TROUBLESHOOTING.mdExit codes, glyphs, frequent failures, digging into runseveryone
0CONFIG.mdEvery proef.toml key with defaultsP2 pack maintainers
0DIAGNOSTICS.mdThe greppable index of every diagnostic codeP2 pack maintainers
0EVENTS.mdThe events.jsonl wire schema for CI consumersCI engineers
1PRD.mdWhat are we building, for whom, and how do we know it works?everyone
2adr/ — ADR-0001 onwardWhy is it built this way? Each decision, alternatives, consequencesengineers
3TECH-SPEC.mdHow exactly is it built? Types, pipeline, schemas, verified seam factsimplementers
4IMPLEMENTATION-PLAN.mdIn what order, with what acceptance criteria? M0–M6 task breakdown, risks, runbooksimplementers
5TESTING-STRATEGY.mdHow is every layer verified?implementers
—IMPROVEMENT-PLAN.mdPost-M5 competitive review: the feature roadmap, each item carrying a Status column (14 of 16 shipped)maintainers
—OPEN-FINDINGS.mdThe worklist. Every open defect and gap, whichever review found it, plus what shipped against eachmaintainers
—RELEASING.mdVersioning policy and the release runbookmaintainers
—CONTRIBUTING.mdSetup, gates, and the rules that are easy to trip overcontributors
—SECURITY.mdThreat model and vulnerability reportingeveryone
—CHANGELOG.mdPer-release change log (SemVer)everyone
—../CLAUDE.mdRepo-root guidance for Claude Code: constraints, seam facts, commands, statuscoding agents

Decision log (ADR index)

ADRDecisionStatus
0001Embed hurl’s crates in-process as the API engineAccepted
0002Multi-engine core: EngineFactory/EngineSession seam, step-kind routing, batchingAccepted
0003Exact pins + thin zero-diff fork as patch vehicle + upgrade canaryAccepted
0004Packs = YAML skeleton + embedded raw Hurl blocksAccepted
0005${…} author-time / {{…}} run-time variables; World; secretsAccepted
0006Engine traits are sync + dyn; no async machinery in v1Accepted
0007Cooperative cancellation at batch boundaries + budgets (hurl has none)Accepted
0008Serde event enum = run record; decorator reporter stack; libtest-mimic modeAccepted
0009User/TestFailure/System → exit 2/1/3; miette at the CLI edgeAccepted
0010Emitted .hurl artifacts are the executed input (same bytes) + sidecarsAccepted
0011Fixture server is synchronous tiny_http, not axum (tokio-runtime ban)Accepted
0012Project config & environments in proef.toml ([url]/[vars]/[env.*], --env, deep-merge)Accepted
0013Typed macro parameters (params name→type map; best-effort literal-args lint)Proposed (defer)
0014Suite-level setup/teardown ([run] setup/teardown features, CLI-edge orchestration)Accepted
0015Injected run-relative timestamps + worker id (sink-stamped, sans-IO core) for the HTML timelineAccepted
0016OpenAPI → suite generator: one-shot seed allowed under a bright line; oracle/drift mode permanently rejectedProposed (defer)
0017proef lsp language server: sync lsp-server, whole-suite wholesale recompute, injectable-provider + collect-all front-end refactorAccepted
0018Named hurl fragments: ref: as a second macro body form, # @proef <name> in real .hurl files, explicit bind: scopesAccepted
0019Reserved tags and the authored skip: @skip[:reason] at the CLI edge, reasons in every sink, authored-vs-mechanical split for --rerun, quarantine aligned in JUnitAccepted
0020Run metadata is explicit-injection-only: --meta/[meta]/[env.<name>.meta], run_started gains env/metadata/shuffled, harvested-vs-handed-over is the boundaryAccepted
0021Run discovery and rotation are separate questions: a directory is a record because it holds one, rotation still deletes only uuid-named dirs, ordering follows the uuid-v7 timestamp, then mtimeAccepted

Naming & identifiers

Project/binary proef · crates proef-core, proef-engine-hurl, proef-cli, proef-fixture, proef-harness, proef-lsp · run records .proef-runs/ · persistent World .proef-state.json · config proef.toml. The crates.io names proef, proef-core, proef-engine-hurl, and proef-lsp are published and owned.

Provenance

Grounded in two verified sources: (1) hurl master (Orange-OpenSource) — source-level verification of every library seam used, with file:line citations in TECH-SPEC §5; (2) a working spike proving the front end and artifact contract end to end (the spike predates this repo). Ecosystem practices are drawn from cargo-nextest, cucumber-rs, sqlx, rustls, probe-rs, and current (2026) Rust guidance.

Installing proef

Pick whichever fits your machine:

# Homebrew (macOS or Linuxbrew, arm64 or x86_64):
brew install emrecdr/proef/proef

# Prebuilt binary via cargo-binstall (any Rust dev environment):
cargo binstall proef

# From source via crates.io:
cargo install proef --locked

Or grab a prebuilt archive from a GitHub Release — five targets ship per tag (macOS arm64/x86_64, Linux arm64/x86_64-gnu, Windows x86_64-msvc), each with SLSA build provenance (gh attestation verify <archive> --owner emrecdr) and a .sha256 sidecar (sha256sum -c proef-<tag>-<target>.tar.gz.sha256 from the download directory). The Windows zip bundles the libcurl/libxml2 DLLs the binary needs; Linux binaries expect the distro’s libcurl4 and libxml2 (present on virtually every system, or apt install libcurl4 libxml2).

Prebuilt binaries need no Rust toolchain and no build prerequisites. Building from source (the cargo install/cargo binstall-fallback path) needs the native libraries the embedded hurl engine links: apt install build-essential pkg-config libssl-dev libcurl4-openssl-dev libxml2-dev libclang-dev on Linux; the Xcode command-line tools suffice on macOS.

Completions and the man page

Shell completions and a man page travel in every release archive (completions/, proef.1) — or generate them from any installed binary:

proef completions zsh > "${fpath[1]}/_proef"   # also bash, fish, powershell, elvish
proef man > /usr/local/share/man/man1/proef.1

In CI

cargo binstall proef resolves the release archives, so any workflow with a Rust toolchain installs in seconds; without one, download an archive and its .sha256 from the Release directly. A complete GitHub Actions workflow — secrets, JUnit, sharding, the regression gate — is in CI.

First run

proef doctor verifies the environment (embedded engine, native libraries, project layout, runs-dir writability). Then Getting started takes you from proef init to a green run.

A note on the dev fixture

The tutorial’s zero-network first run uses the dev fixture API (cargo run -p xtask -- fixture), which lives in this repository and is not part of the installed binary — it needs a checkout and a Rust toolchain. With an installed binary alone, point [url] base at any API you own instead; the proef init scaffold passes as-is against the fixture, and its routes (/health, /search, /version) are one file away from targeting yours.

Getting started — your first suite in ten minutes

proef runs end-to-end API tests written as plain Gherkin prose. The prose stays readable by anyone; a YAML pack binds each sentence to real HTTP work. This walkthrough builds a two-file suite from nothing and runs it.

Install first (Installing), then verify:

$ proef doctor
engine `hurl`:
  [ok  ] embedded hurl            hurl 8.0.1 (exact pin, ADR-0003)
  ...

0. Or scaffold it: proef init

$ proef init
  created ./proef.toml
  created ./suite/case.feature
  created ./suite/packs/api.yaml
  created ./hurl/api.hurl
  created ./.gitignore
  created ./suite/packs/proef-pack.schema.json
  ok ./suite/packs/api.yaml (modeline added)

created 6 file(s), skipped 0
next: proef test --dry-run  (then point ${url:base} at your API — the scaffold's routes are placeholders)

proef init writes the three files this walkthrough builds by hand below — proef.toml, suite/case.feature, suite/packs/api.yaml — plus a one-entry hurl/api.hurl, the pack JSON Schema and a .gitignore. The .hurl file is there because a step has two body forms: the scaffold’s pack shows an inline hurl: block and a ref: naming that file’s # @proef entry, so the form an adopter with an existing hurl corpus wants is visible from the first command rather than only in the docs. It is a deliberately smaller suite than the one this tutorial builds: no Then the first hit is record "r-1" step, no firstHit expect: macro, no ${secret:apiToken}, and search targets a plain /search route instead of the tutorial’s /api/v1/admin/search/records — while adding one step the tutorial does not, And the corpus reports its version, whose macro is the scaffold’s ref: example. Pasting the tutorial’s Then line into the scaffold’s feature file fails with bind::unbound_step — read on to build the fuller suite by hand, or extend the scaffold yourself once you understand the pieces.

1. A suite is three files

suite/
  case.feature        # the prose — what the test says
  packs/
    api.yaml          # the vocabulary — what each sentence does
proef.toml            # the configuration — where URLs and variables live (§3.5)

Layout is convention, not configuration. proef test suite takes one path and discovers everything under it: every *.feature file is a test file (at any depth), and every *.yaml/*.yml file directly inside a directory named packs is a macro pack — the packs directory itself may sit at any depth. All packs merge into one vocabulary shared by all feature files — grow the suite by adding files; nothing needs registering.

2. Write the prose

suite/case.feature:

Feature: Directory search
  Scenario: A known record is found
    Given the service is healthy
    When the operator searches for "Acme"
    Then the first hit is record "r-1"

The feature is pure prose — no URLs, no environment data. The target host and any variables live in proef.toml (§3.5) and reach the packs as ${url:…} / ${vars:…}.

3. Bind the prose

suite/packs/api.yaml:

macros:
  health:
    match: the service is healthy
    steps:
      - hurl: |
          GET ${url:base}/health
          HTTP 200

  search:
    params: [term]
    match: the operator searches for {term}
    steps:
      - name: search records for ${term}
        hurl: |
          GET ${url:base}/api/v1/admin/search/records
          Authorization: Bearer ${secret:apiToken}
          [Query]
          q: ${term}
          HTTP 200
          [Captures]
          recordId: jsonpath "$[0].id"

  firstHit:
    params: [id]
    match: the first hit is record {id}
    expect:
      - hurl: |
          jsonpath "$[0].id" == "${id}"

Three macro shapes are on display: a fixed sentence (health), a parameterized one ({term} captures the quoted word — quotes are shed), and an assert-only expect: macro whose lines merge into the previous request’s asserts. The hurl: blocks are raw Hurl — validated with the real parser the moment the pack loads.

3.5 Where URLs and variables live (proef.toml)

Variables are declared in proef.toml — never in the .feature files — and referenced as ${url:…} / ${vars:…}. The pack’s ${url:base} above resolves from here; proef finds the nearest proef.toml, searching up from the working directory (like cargo/git):

# proef.toml
[run]
suite = "suite"                    # `proef test` needs no path argument
# fragments = "hurl"               # uncomment to `ref:` entries of real .hurl files (§3.6)

[url]
base = "${env:PROEF_BASE_URL:-http://127.0.0.1:8787}"   # → ${url:base} (env override wins)

[vars]
apiVersion = "v1"                  # → ${vars:apiVersion}

[env.staging.url]                  # per-environment overrides (mirror the base tables)
base = "https://staging.example.com"

[env.prod.url]
base = "https://api.example.com"
[env.prod.http]                    # an env may override a runner setting too
timeout-ms = 60000

A macro then reads GET ${url:base}/api/${vars:apiVersion}/… with nothing declared in the feature. Pick an environment at run time:

$ proef test                       # base [url]/[vars]; runs `[run] suite` — "suite/" in the file above (`tests/` is only the fallback when the key is unset)
$ proef test --env prod            # [env.prod.*] deep-merged over the base (or set PROEF_ENV=prod)

The rule is uniform: under [env.<name>], url.* / vars.* override variables and http.* / run.* override runner settings — anything unlisted inherits the base, so an environment names only what changes (the Cloudflare-Wrangler / Cargo-profile model). Secrets never live here — they stay in the encrypted store (${secret:…}, §5).

3.6 A step’s body has two forms

Everything above uses an inline hurl: block, which is complete and permanent. The other form is ref: <name>, which points at one entry of a real .hurl file marked # @proef <name>, with values supplied by a bind: table:

Uncomment fragments = "hurl" in §3.5’s proef.toml (a second [run] table would be a TOML error — one table, both keys), put the annotated file under hurl/, and point a step at its entry:

    steps:
      - ref: admin.search          # one entry of hurl/admin.hurl
        bind: { q: "${term}" }

That file stays valid hurl, so the same bytes run under stock hurl and under proef — useful when a corpus already exists and you would rather annotate it once than transcribe it. Choose by capability, not taste: inline splices text (so ${docstring} can carry a multi-line body, which no binding can express), while ref: buys a name, reuse, standalone runnability, and a static check that every variable is supplied. AUTHORING.md §“hurl: or ref:” has the full comparison.

4. Validate without a network

$ proef test suite --dry-run

This binds every sentence, resolves every ${…}, emits the artifacts, and parse-validates them — no request is sent. Typos in prose, packs, or hurl blocks all fail here, with the file, line, and a “did you mean” where one exists. Wire your editor too: proef schema --add-to suite/packs/api.yaml gives packs autocomplete via the JSON Schema.

5. Provide the secret and a target, then run

Point PROEF_BASE_URL at any HTTP API you can reach — or start proef’s own dev fixture in a second terminal: from a checkout, cargo run -p xtask -- fixture binds the default base port (8787), so no PROEF_BASE_URL is needed (it prints a PROEF_BASE_URL line to export only when it ends up somewhere other than 8787 — the port was busy, or you passed a different ... -- fixture <port>). It also prints the fixture’s own token line: export PROEF_SECRET_APITOKEN=fixture-token — the value the next step needs, because every /api/v1/ route rejects anything else with a 401. Then:

$ proef secret set apiToken        # paste fixture-token; or: export PROEF_SECRET_APITOKEN=fixture-token
$ proef test suite
running 1 scenario(s) with 8 job(s) — run 019f…

  Scenario: A known record is found (suite/case.feature)
    ✓ suite/case.feature:3 — the service is healthy (2ms)
    ✓ suite/case.feature:4 — the operator searches for "Acme" (5ms)
    ✓ suite/case.feature:5 — the first hit is record "r-1" (0ms)
    ✓ scenario A known record is found

summary: 1 passed · 0 failed · 0 skipped

Secret values never appear anywhere — not in artifacts, events, logs, or reports.

6. When it fails

A failing assert names the feature line, the artifact line, and hands you a reproduce command (absolute paths; captures ride along in a .vars file):

  ✗ suite/case.feature:5 — assert failure (artifact case--a-known-record-is-found.hurl:12)
  reproduce: hurl --test /abs/path/.proef-runs/<run-id>/artifacts/case--a-known-record-is-found.hurl --variables-file /abs/path/….vars

Because this suite binds ${secret:apiToken}, the replay also needs the secret — proef never writes its value anywhere, so add it yourself: --secret apiToken=fixture-token (the artifact’s # replay: header says so).

Every run leaves a record under .proef-runs/<run-id>/: events.jsonl (the machine-readable event stream), run.log (the console mirror), and artifacts/ — the exact .hurl files that were executed, byte for byte. proef explain summarizes the latest run from the record, proef diff compares two runs — surfacing what regressed, what got fixed, and which steps turned flaky or slower — and proef report writes a self-contained HTML page of a run (scenario tree, timings, and deep-links to the executed artifacts).

Where next

  • AUTHORING.md — the full pack reference: composition with use:, retries, optional steps, guards, captures across scenarios, fakes.
  • TROUBLESHOOTING.md — exit codes, the glyph legend, and the frequent failures; DIAGNOSTICS.md indexes every error code you might hit.
  • proef flows suite lists every scenario with tags; --format json feeds the nextest harness (one IDE test per scenario).
  • proef test suite --watch reruns on every file change.
  • Tag a scenario @skip / @skip:reason to park it — still counted, its reason in every report — or @quarantine so a flaky one stops gating CI while you fix it (AUTHORING.md · Reserved tags).
  • Exit codes are a contract: 0 pass · 1 test failure · 2 your input · 3 environment — safe to wire straight into CI, with --junit for reports.

Writing scenarios

For the person who writes tests, not the person who wires them up.

You describe what should happen, in sentences. Somebody on your team maintains the vocabulary — the list of sentences proef understands — in files called packs. You never have to open one. This page is your whole loop.

If you also maintain the vocabulary, you want AUTHORING.md instead; this page deliberately stops at the boundary.

1. A test is sentences in a file

A test lives in a .feature file and looks like this:

Feature: Directory search
  Scenario: A known record is found
    Given the service is healthy
    When the operator searches for "Acme"
  • Feature — what area you are testing. One per file.
  • Scenario — one situation worth checking. A file can hold many.
  • Given / When / Then — the steps, in order. proef strips the opening word and matches only the rest of the sentence, so all of them behave identically; pick whichever reads best. Given for setup, When for the action, Then for the check. And and But work too, and read better than repeating yourself:
    When the operator searches for "Acme"
    Then the response status is 200
    And the value at "$.results" is a non-empty list

Each step is one sentence, and every sentence must be one the vocabulary knows. That is the only rule you have to satisfy.

2. See the sentences you may write

$ proef macros
builtin:core.yaml
  expectPresent                0×  the value at {path} is present  (builtin, unused here)
  expectStatus                 0×  the response status is {status}  (builtin, unused here)
  …
suite/packs/api.yaml
  health                       1×  the service is healthy
  search                       1×  the operator searches for {term}

10 macro(s) · 0 unused

The right-hand column is what you can say. The left is the internal name — you do not need it. The 1× counts how many steps across the suite use that sentence.

{term} is a blank you fill in. the operator searches for {term} means you write:

    When the operator searches for "Acme"

Quotes are the house style and make the value easy to see.

This command works even when your file has a mistake in it — which is exactly when you need it. If some step does not bind, you still get the list; only the 1× counts are withheld (they would be misleading), and they show as —.

Editor tip. If someone has set up proef lsp for your editor, these sentences appear as autocomplete while you type, and mistakes underline as you go. It is worth asking for — but nothing on this page needs it.

3. Check before you run

$ proef test --dry-run

This reads your file and tells you whether every sentence binds. It sends no requests and touches nothing, so it is always safe and takes well under a second. Run it constantly.

  ok suite/case.feature — 1 scenario(s), 2 step(s), 1 batch(es)

dry-run OK: 1 feature(s), 1 scenario(s), 2 step(s) …
next: proef test

When it says next: proef test, your sentences are good.

4. Run it

$ proef test
  Scenario: A known record is found (suite/case.feature)
    ✓ suite/case.feature:3 — the service is healthy (6ms)
    ✗ suite/case.feature:4 — the operator searches for "Acme" (0ms)

✓ passed · ✗ failed · ∅ skipped, because an earlier step in the same scenario failed.

A ✗ here is a real result — proef reached the system and the answer was wrong. That is the test doing its job. Take it to whoever owns the API, with the line it names.

5. The two mistakes you will actually make

Everything else is someone else’s problem. These two are yours.

“no macro matches …”

proef::bind::unbound_step

  × no macro matches `the operator serches for "Acme"` — did you mean
    `the operator searches for {term}`?
   ╭─[suite/case.feature:4:5]
 4 │     When the operator serches for "Acme"
   ·     ────────────────────────────────────
   ╰────

You wrote a sentence the vocabulary does not know — usually a typo, a plural, or a word order that drifted. proef points at the exact line and, when it can, guesses what you meant.

Your fix: say a sentence that exists. proef macros lists them.

The message also shows a macros: block. That is instructions for the person who maintains the vocabulary — hand it to them if the sentence you need genuinely does not exist yet. It is not something you need to write.

“url variable … is not set”

proef::resolve::missing_config_var

  × in macro `health`: url variable `bse` is not set — define `[url]` `bse`
    in proef.toml (or in the active `[env.<name>.url]`) — did you mean `base`?

A sentence you used needs an address that nobody has filled in. This one lives in proef.toml, which is a short settings file, not a pack:

[url]
base = "https://api.your-company.com"

Your fix: if you recognise the name, correct the typo it suggests. Otherwise this is a setup question for whoever configured the project.

One more you may meet once

If a brand-new project fails on its very first run and says it is still the proef init scaffold, nothing is broken — the starter files point at a placeholder address and placeholder routes on purpose. Somebody needs to point [url] base at the real API and replace the example routes. After that, the loop above is all yours.

6. The loop

read the sentences   →   proef macros
write a scenario     →   your .feature file
check it binds       →   proef test --dry-run
fix what did not     →   the two errors above
run it               →   proef test

That is the whole job. When you need something the vocabulary cannot say, that is a conversation with whoever maintains the packs — and AUTHORING.md is the page for them.

Where to go next

You want toRead
Understand a symbol, exit code, or failureTROUBLESHOOTING.md
Get autocomplete and live error underliningEDITORS.md
Maintain the vocabulary yourselfAUTHORING.md
Set up a project from scratchGETTING-STARTED.md

Authoring reference — packs and features from the author’s seat

Everything here is validated at load or at --dry-run time; nothing fails only at execution that could have failed earlier. Start with GETTING-STARTED if this is your first suite.

Feature files

Standard Gherkin: Feature:, Scenario:, Background: (prepended to every scenario), Rule:, Scenario Outline: + Examples: (expanded, #N-deduped — note the #N is positional: inserting an Examples row above renames every instance below it, which re-keys their JUnit history and re-buckets them across --shard. A column placeholder in the outline’s name (Scenario Outline: search finds <q>) keeps each instance’s identity tied to its data instead of its row number), data tables (rows become step arguments), and docstrings (delivered to the macro as the docstring param). Keywords (Given/When/Then/And) don’t affect binding — only the sentence text does. Prose between the Feature: line and the first scenario is the feature’s description: never executed, but proef flows prints it (JSON: featureDescription) so a file’s intent travels with its inventory.

An outline’s <column> placeholders substitute into the docstring too, not just the scenario name, step text and table cells. That is how a request body gets data-driven without leaving the feature file:

  Scenario Outline: Posting <label>
    When a record is posted
      """
      {"label": "<label>", "priority": "<priority>"}
      """
    Then the response status is 201

    Examples:
      | label | priority |
      | alpha | high     |
      | beta  | low      |

A <name> that is not an Examples column is a parse-time error wherever it appears, docstrings included.

Variables are declared in proef.toml ([url] / [vars]), never in the feature file, and referenced from packs as ${url:key} / ${vars:key} — see CONFIG.md. Feature files stay free of URLs and environment data.

Tags (@smoke) accumulate feature→scenario; proef test --tags <expr> selects scenarios by a boolean expression over them — and, or, not, and parentheses, with the @ optional (e.g. --tags "@api and not @slow" or --tags "(smoke or nightly) and not wip"). A bare tag is a valid expression; a selection matching nothing is an error, not a silent green run. Atoms may glob: * matches any run of characters, ? exactly one — anchored to the whole tag and case-sensitive — so --tags "JIRA-*" selects every ticket-tagged scenario, while a metachar-free atom stays plain equality.

Reserved tags

Two tag names carry behavior; every other tag is yours (selection via --tags, grouping, traceability):

  • @quarantine — the scenario runs and reports, but a test-failure does not gate the exit code, reaches JUnit as <skipped message="quarantined failure (non-gating): …">, and maps to # TODO in TAP. For flaky tests while they are being fixed — a User/System fault still fails the run.

    Quarantine is a holding pen, not a destination, and proef flaky is what keeps it one: over the retained records it separates a quarantined scenario that fails every run (DISABLED — switched off, and nobody is watching those failures because by design nothing reports them) from one that has been green throughout (recovered — the tag outlived the problem and is now suppressing the next real regression). Neither is visible any other way.

  • @skip / @skip:<reason-token> — the scenario is parked: never prepared or run, counted as skipped, and the reason (the tag spelling itself) appears in the console, JUnit, TAP, the record, the report, explain and flows. --tags "not @skip*" unselects both spellings when you want them gone from the totals too. All-skipped exits 0; --dry-run still validates a skipped scenario (skip is not a validation waiver). In [run] setup/teardown features, reserved tags have no effect.

A tag that is almost reserved — @quarantined, @skipped, @Skip — is an ordinary tag with no effect, and since 0.18 it warns (tags::reserved_tag_typo) with the spelling it likely meant, because a scenario its author believed quarantined would otherwise gate the build in silence.

Conditional, data-dependent skipping is a step concern and stays in packs: when: guards a step, optional: soft-fails one, retry: bounds one. That split is deliberate and permanent — scenario prose stays declarative; there is no scenario-level IF/WHILE/TRY and none is planned (ADR-0019).

Macros (macros: in a pack)

macros:
  name:
    match: the record {name} is resolved   # sentence pattern (optional)
    params: [name]                         # declared parameters
    defaults: { index: records }           # defaults for optional params
    description: One line for humans.
    tags: [Admin]
    steps: [...]                           # OR expect: [...] — never both
  • match: binds prose. Patterns need at least one literal word (no capture-only patterns), captures are {name}, quoted arguments shed their quotes, and matching is leftmost with ambiguity rejected — two macros that could claim the same sentence fail pack load.
  • Every {capture} must be a declared param; defaults: keys must be declared params; adjacent captures ({a} {b} with nothing between) are rejected.
  • A macro without match: is composition-only (reachable via use:).
  • proef macros lists every macro with its match: prose — the sentence a feature file may say — and its call count across the corpus, flagging pattern macros no scenario binds (dead prose bindings); use:-only helpers and unused builtins are listed but not flagged. It also flags near-duplicate pattern macros — two that differ only in their {capture} names (the same literal skeleton), which are confusable to authors. Both are advisory only: they never change the exit code, and --format json carries pattern, unused and nearDuplicateOf fields for a CI hygiene gate.
  • When a step does not bind, macros still lists the vocabulary (that is when you most need it) and keeps exit 2. Counts are withheld, not zeroed: calls/unused render as —/null, because an unbound feature contributes no calls and would make its own macros look dead. proef flows still refuses — it promises every scenario, and a silently partial list is a wrong answer.

Steps

steps:
  - name: human label (${…} resolves here too)
    optional: true                    # failure warns instead of failing
    retry: { count: 10, interval_ms: 300 }   # finite; 1..=10000
    delay: 250                        # ms before the request (capped at 1 hour)
    when: "${env:RUN_SLOW:-}"         # skips when empty or false/0 after resolution
    saveAs: { recordId: global }      # promote a capture to the global store
    hurl: |                           # the payload — raw hurl, one or more entries
      GET ${url:base}/api/v1/records/{{recordId}}
      HTTP 200
  - use: otherMacro                   # composition (cycle-checked, depth ≤ 32)
    with: { term: "${name}" }

retry:/delay: are baked into the entry’s [Options] so the emitted artifact replays with identical semantics under stock hurl. optional: steps run as their own batch so a failure cannot poison neighbours. Raw [Options] written inside a hurl: block are linted by the same rules — retry:/repeat: finite and at most 10 000, delay:/retry-interval:/max-time: at most one hour, and a key declared both in YAML and in the block is refused (pack::option_declared_twice) — and whatever the values, a batch’s watchdog budget never exceeds four hours (ADR-0007).

expect: macros carry no requests: status: 200 and/or raw hurl: assert lines merge into the previous request entry (a Then before any When is an error).

hurl: or ref: — two body forms, chosen by capability

A step’s body is an inline hurl: block or a ref: naming an entry in a real .hurl file (ADR-0018). They are not two spellings of one thing, so the choice is not a matter of taste:

hurl: |ref: name
variables${…} spliced before hurl parses{{…}} bound via bind:
can substituteanything, anywhere — including a whole multi-line docstring bodyonly what hurl can template
reusenone: the block has no nameany number of macros, each binding differently
runs under stock hurlnoyes, unchanged
unknown variablecaught when the artifact is parsedcaught at --dry-run, by name

Reach for inline for a request only this macro makes, and always when you need to splice something hurl cannot template — ${docstring} as a request body is the clearest case, since a bound value is a single-line scalar.

Reach for ref: when the same request serves several macros, when the hurl was written by somebody else, or when the file must stay runnable on its own.

bind:                              # pack scope — every macro in this file
  base:     ${url:base}
  apiToken: ${secret:apiToken}     # injected at run time, never into an artifact
macros:
  search:
    params: [q]
    match: "the operator searches for {q}"
    bind: { q: "${q}" }            # macro scope
    steps:
      - ref: admin.search
        bind: { index: records }   # step scope — the most specific wins

hurl: and ref: are chosen per step, not per suite. A step is one or the other, but a macro mixes them freely — which is what adopting an existing corpus looks like in practice: ref: the requests the corpus already has, write inline for the ones it doesn’t.

  archiveFirstResult:
    match: the operator archives the first result
    steps:
      - ref: admin.search        # the corpus already has this request
      - hurl: |                  # this one is new, and splices ${…}
          POST ${url:base}/api/v1/admin/records/{{recordId}}/archive
          HTTP 204

recordId is captured by the fragment and read by the inline step: the World threads captures across both forms, and contiguous same-engine steps batch together whichever form they were written in. Pick per step by capability — inline splices ${…} anywhere (including a multi-line ${docstring} body, which a single-line bound scalar cannot express); ref: keeps the file runnable on its own and checks its interface by name.

proef fragments lists the corpus: which entries exist, how many scenarios run each, and — the two questions a listing exists for — which are annotated but reached by nothing, and which carry no # @proef at all and so cannot be referenced. --check exits 1 on the first; add --require-annotated to include the second, which is opt-in because an unannotated entry is inert by design (pointing at a corpus you did not write costs nothing), and only a team mid-port means “not done yet” by it.

Set [run] fragments to the directory holding those files (see CONFIG.md). Every {{variable}} a fragment reads must be bound in one of the three scopes, captured by an earlier step, or supplied by the fragment itself; nothing is implicit, because hurl’s per-entry variable: assigns into one shared set rather than scoping, so an unbound name would quietly inherit an earlier entry’s value.

A fragment supplies its own value with an ordinary [Options] variable: line — which is how a corpus file stays runnable on its own, with fewer variables to pass in:

# @proef admin.search
GET {{base}}/api/v1/admin/search/{{index}}
[Options]
variable: index=records      # the file answers its own question
HTTP 200

Do not then also bind: that name. Both spellings reach the entry as variable: index=, hurl takes the last, and the fragment’s own line is last — so the bound value would never reach the request. Proef refuses the pair (pack::option_declared_twice) rather than picking one silently; delete whichever is not authoritative.

Bindings resolve once per scope instantiation — pack scope once per scenario, macro scope once per invocation, step scope per step — so one bind: entry is one value, and two are two. That is what makes a pack-scope ${fake:email} a single identity for the whole scenario.

A binding nothing can read is refused rather than dropped, at two levels:

  • The table — a bind: with no ref: in scope to read it (proef::pack::bind_without_ref), checked at all three scopes.
  • One key — a bind: entry no fragment in that scope reads (proef::pack::unread_bind_key), with did-you-mean over the names that are read. This is the one a typo produces: bind: { token: …, toekn: … } binds one real key and one that never arrives.

The key check is a union over the scope, never against a single fragment — a pack-scope table is the plumbing every macro in the file needs, so a key serving one macro and not its siblings is correct usage.

Note the one that surprises people: a macro-scope bind: does not reach a use: target — the target resolves its own pack and macro scopes — so the table belongs on the macro that actually carries the ref:.

You do not have to memorise a corpus you did not write: with proef lsp running, completing inside a bind: table offers the {{variables}} the fragments this pack ref:s actually read, each labelled with the fragment that wants it. The names are read off the .hurl file itself, so they cannot drift from it. If a name still goes unsupplied, proef::lower::unbound_placeholder names it at lower time — --dry-run is enough to surface that, no server needed.

The same check reads inside your bind values — with hurl’s own parser, so what hurl calls a function is never mistaken for a variable ({{newUuid}} and {{newDate}} need no supplier). A {{name}} that is a variable is templated when the entry runs, so it must be supplied by then: an earlier step’s capture (the usual shape: bind: { path: "records/{{recordId}}" }), the fragment’s own [Options] variable: (its lines evaluate before the injected ones), a secret in scope, or a sibling literal bind whose name sorts before this one — the injected lines are written and evaluated in name order, so base can feed q but not the other way around. And the reverse collision warns rather than fails: a literal bind: that re-uses a name an earlier step captured wins silently from that entry on (hurl’s variable: assigns into one shared set), so proef::lower::bind_shadows_capture names it — rename the binding if the capture was the point.

Where a file,…; asset lives

A file body — a request body, a multipart part, or a file,…; an assert compares against — sits beside the source that names it:

The step isThe file goesBecause
hurl: | (inline)beside the featurethe suite is one unit; tests/features/fixture.jpg serves tests/features/*.feature
ref: <name>beside the fragmentthat is where stock hurl looks, so the file keeps running on its own

This is hurl’s own rule (its --file-root defaults to the .hurl file’s directory), and the one Karate, pytest and Jest use for fixtures. You never set a root: proef stages each asset from its own source into the run’s artifacts/assets/<slug>/, and the emitted artifact carries the matching --file-root in its replay line, so the hand-off runs unchanged under stock hurl.

Two consequences worth knowing. The path must be a plain relative one — no leading /, no .. — because it names a file inside your suite, not a location on the machine. And a missing file is refused at --dry-run, naming the directory it was looked for in (proef::run::asset_unstageable) — so the gate CI runs before standing an environment up answers it, rather than a failing request minutes later, or hurl reporting an unreadable body against the artifact. Staging resolves beside the file the parser actually read, so the directory you run from does not matter; a symlink already at a staging destination is replaced, never written through; and two references that name one file on a case-insensitive filesystem (Data.json and data.json on macOS or Windows) are refused (proef::run::asset_unstageable) rather than silently last-writer-won.

Recipes — the three shapes every real suite needs

Log in, then use the token

The most common real-world flow: the API mints a token at a login endpoint, and every later request carries it. A capture crosses steps through the World, so the shape is one login macro capturing the token and any number of authed macros reading it:

macros:
  logIn:
    match: the operator logs in
    steps:
      - hurl: |
          POST ${url:base}/auth/login
          {"user": "${vars:user}", "password": "${secret:password}"}
          HTTP 200
          [Captures]
          token: jsonpath "$.token"

  listRecords:
    match: the records are listed
    steps:
      - hurl: |
          GET ${url:base}/records
          Authorization: Bearer {{token}}
          HTTP 200

{{token}} is hurl run-time vocabulary (the second tier): the capture is assigned when the login entry runs, and every later entry in the scenario reads it — no bind:, no config. To reuse one login across scenarios, add saveAs: { token: global } to the login step and read ${global:token}; note the promotion is refused if the value carries a secret (ADR-0005).

Wait for an eventually-consistent result

A POST answers 202 and the resource appears a moment later. Don’t sleep — put a finite retry: on the step that polls, and let its asserts be the condition:

  awaitRecord:
    params: [id]
    match: record {id} is eventually visible
    steps:
      - retry: { count: 10, interval_ms: 300 }
        hurl: |
          GET ${url:base}/records/${id}
          HTTP 200

The retry re-runs the entry until its asserts pass or the count runs out — and the count must be finite: hurl cannot be interrupted mid-call, so an unbounded retry is a hang the watchdog would have to abandon (the same reason retry: -1 is refused at load). The retry also bakes into the emitted artifact’s [Options], so a replay under stock hurl polls the same way.

Seed data before, clean up after

Three scopes, three mechanisms — pick by lifetime:

  • Per scenario: a Background: in the feature file runs its steps before every scenario in that file — provisioning prose, bound to macros like any other step.
  • Per suite: [run] setup / [run] teardown in proef.toml name feature files that run once around the whole pool — seed a database, then delete the run’s residue. Teardown runs even when the suite fails or is interrupted; a teardown failure is exit 3, never silent. The keys are documented in Configuration.
  • Across scenarios: a saveAs: { id: global } capture in setup (or any scenario) persists into later scenarios and runs as ${global:id} — the handle teardown needs to delete what setup created.

Asserting responses — the hurl vocabulary

Assertions live inside a step’s raw hurl: block (or an expect: macro), so the whole hurl 8.0 assert grammar is available untouched: proef resolves ${…} before the run and hands the rest to the embedded engine verbatim (ADR-0005). The block is parsed at pack-load time, so a grammar slip fails fast with a diagnostic — only JSONPath semantics (a query that lexes but never matches at run time) can slip through. An assert reads <query> [filters…] <predicate>; HTTP <status> is the implicit status assert. The authoritative list is hurl’s own asserting-response and filters docs; the common shape:

GET ${url:record}
Authorization: Bearer ${secret:apiToken}
HTTP 200
[Asserts]
jsonpath "$.id"        isUuid                 # type/shape checks — schema-lite
jsonpath "$.status"    == "active"
jsonpath "$.createdAt" isIsoDate
jsonpath "$.tags"      count == 3
header "Content-Type"  contains "application/json"
duration               < 1000                 # response-time budget (ms)
  • Queries (what to read): status, header "<n>", cookie "<n>", body, bytes, jsonpath "<expr>", xpath "<expr>", regex "<pat>", sha256, md5, url, redirects, variable "<n>", duration, certificate "<f>".
  • Predicates (the check): == != > >= < <=, startsWith, endsWith, contains, includes, matches "<regex>", exists / not exists, isEmpty, and the type family isBoolean isInteger isFloat isNumber isString isCollection isList isObject isDate isIsoDate isUuid isIpv4 isIpv6.
  • Filters transform the value before the predicate, chained left→right: count, nth <n>, first, last, split "<sep>", replace/replaceRegex, toInt/toFloat/toString, toDate "<fmt>", format "<fmt>", base64Decode/base64Encode, urlDecode/urlEncode, daysAfterNow/ daysBeforeNow, jsonpath "<expr>", regex "<pat>", utf8Decode.
[Asserts]
jsonpath "$.items"     count > 0
header "Set-Cookie"    split ";" nth 0 startsWith "session="

RFC 9535 JSONPath (hurl 8.0): filter expressions and functions are in scope — jsonpath "$.books[?(@.price < 10)]", length(), count(), match(), search(). Use the bracket form for names with hyphens ($['x-custom-id']), and assert a missing path with not exists (a non-matching path yields no value, not count == 0). The same query+filter grammar drives [Captures], threading a value into later steps as {{name}}.

Built-in shape macros

For the most common single-value shape checks, the built-in Core pack ships a small, product-neutral set of expect: macros so you rarely hand-write the predicate. Each reads a JSONPath ({path}) into the previous response and merges one assert:

When the record is fetched
Then the value at "$.id" is a uuid
And  the value at "$.name" is a string
And  the value at "$.tags" is a non-empty list

Available: the value at {path} is a string / … a number / … a boolean / … a uuid / … an ISO date / … present / … a non-empty list. Quote the path in prose when it contains spaces (the quotes are optional and shed). These are a convenience layer over the predicates above — they deliberately do not cover whole-body structural matching; reach for a raw expect: hurl: block for anything they omit.

Negative cases — one macro per malformation, one shared expectation

A validation suite is naturally combinatorial, and its cases differ structurally rather than by value: one omits a key, one empties it, one adds a key the caller may not set. A single parameterised macro cannot express that — the bodies are different shapes, not one shape with a hole in it. Outline placeholders do substitute into docstrings, but an Examples cell cannot practically hold JSON, and a raw body in the feature file defeats the prose the design exists to protect.

Name each malformation. One macro per case, each sentence saying what is wrong in business terms:

macros:
  createWithEmptyTitle:
    match: creating a task with an empty title is refused
    steps:
      - hurl: |
          POST ${url:base}/tasks
          Content-Type: application/json
          {"title": "", "priority": "high"}

  createWithServerOwnedField:
    match: creating a task that sets its own id is refused
    steps:
      - hurl: |
          POST ${url:base}/tasks
          Content-Type: application/json
          {"title": "ok", "id": "caller-chosen"}

Then let one expect: macro serve the whole catalogue. Because an expect: merges its asserts into the previous request entry, a single parameterised expectation covers every case in the set — typically the largest de-duplicator in a validation pack:

  expectErrorCode:
    params: [code]
    match: "the error code is {code}"
    expect:
      - hurl: |
          jsonpath "$.error.code" == "${code}"

The scenarios then read as the specification they are, and each one’s asserts land on its own request:

  Scenario: An empty title is rejected
    When creating a task with an empty title is refused
    Then the response status is 422
    And the error code is TITLE_REQUIRED
# emitted
POST http://127.0.0.1:8787/tasks
Content-Type: application/json
{"title": "", "priority": "high"}
HTTP *
[Asserts]
status == 422
jsonpath "$.error.code" == "TITLE_REQUIRED"

The trade, stated plainly: the pack grows with the malformation catalogue — one macro per case, where a value-driven test would have used one row. That is the cost of feature files that read as prose, and it buys a suite a non-engineer can review. The expectation side does not grow with it.

Variables — the two tiers

SyntaxResolvesWhenExamples
${…}proefat lowering (before execution)${param}, ${env:NAME:-default}, ${url:key}, ${vars:key}, ${run:id}, ${global:key}, ${secret:NAME}, ${fake:name}
{{…}}hurlat run timecaptures ({{recordId}}), secrets ({{apiToken}})

${…} is recursive (captured arguments may themselves contain ${…}, depth ≤ 8) and $${ escapes a literal ${. ${secret:NAME} never inlines the value — it lowers to {{NAME}} and the engine injects it through hurl’s redaction, so artifacts carry placeholders only.

Config variables — ${url:key} and ${vars:key} come from proef.toml’s [url] / [vars] tables, deep-merged with the active --env profile. This is how you keep URLs and settings out of the feature files entirely; a referenced-but-undefined one is a lower-time error. See CONFIG.md (ADR-0012).

Captures ([Captures] in a hurl block) flow forward within the scenario as {{name}}. saveAs: { name: global } additionally promotes the captured value into the persistent global store (.proef-state.json), where later scenarios and later runs read it at lowering time as ${global:name}.

Fakes are deterministic synthetic data: firstName, lastName, name, fullName, email, username, phoneNL, postCode, city, street, int, number, digits4, digits8, bool, word, uuid (unknown generators fail at load). Each ${fake:kind} reference gets its own value — an occurrence counter advances every time a scenario resolves one, so independent ${fake:email} references in the same scenario never collide, however many a step ends up resolving. A step’s name: label is the one deliberate exception: it is not independent of its own payload, so it replays from the start of the step’s own occurrence window instead of minting new ones — the label’s Nth ${fake:…} reference reuses whichever occurrence the payload’s (and when:’s) Nth reference consumed, matched by position, not by generator kind. When the label’s ${fake:…} references mirror the payload’s in kind and order — the common case, e.g. a label that names the same field the payload sends — this reproduces the payload’s own value exactly. When they diverge in kind (say the label’s first reference is ${fake:fullName} but the payload’s first reference is ${fake:email}), the label instead shows whatever that occurrence generated for its own kind — a value the request did not send. That includes a label with more ${fake:…} references than its payload: each extra one still reserves its own place in the sequence, so a later step can never be handed a value the label already displayed. That whole sequence is a pure function of ${run:id}: the same --run-id reproduces the same fakes, byte for byte, across runs. The counter restarts at zero for every scenario — known limitation: two different scenarios that each resolve ${fake:email} at the same position in their own step order (typically each scenario’s first fake reference) get the same address, because both count from zero independently. If two scenarios must not collide, key the value yourself (fold in ${run:id} or a captured id) rather than relying on ${fake:*} alone.

Secrets

Reference with ${secret:NAME}. Values come from PROEF_SECRET_<NAME> environment variables first, then the encrypted store (proef secret set NAME, or … | proef secret set NAME --stdin for scripts (never argv — ps shows it); proef secret list names, proef secret rm NAME removes — XChaCha20-Poly1305, key auto-created 0600 under ~/.config/proef/). Values never appear in artifacts, events, logs, or reports; events carry capture names only. The store file (.proef-secrets.json, mode 0600) holds ciphertext only and is gitignored by default; the key (~/.config/proef/keys/default.key) must never leave your machine.

Artifacts — the executed input

Every scenario emits <feature-path>--<scenario>.hurl — the feature’s suite-relative path with its extension dropped, then the scenario, both slugified (so two same-named features in different directories cannot collide) and capped at 120 bytes with a hash tail — the exact bytes the engine executes — plus .map.json (artifact lines ↔ feature lines, batch and step indices) and .vars (referenced globals as values, secrets as names). The header’s # replay: line is a complete stock-hurl command, including --secret NAME=<value> placeholders for you to fill. proef artifacts <dir> -o out/ --run-id ci emits the same set deterministically for hand-off.

Runs, records, CI

  • Exit codes: 0 pass · 1 test failure (including cancelled runs) · 2 input error · 3 system error.
  • .proef-runs/<run-id>/events.jsonl is the record — one JSON event per line (scenario/step lifecycle, per-attempt progress, failure detail); proef explain summarizes it; --junit and the GitHub Actions summary derive from the same run.
  • --format json prints one machine-readable summary object on stdout (the human report moves to stderr); --jobs N controls parallelism; proef.toml holds project defaults ([run] jobs, [http] timeout-ms).
  • proef fmt suite --check keeps hurl blocks canonically formatted; proef flows lists scenarios.

IDE integration — one test per scenario

The harness crate bridges proef into any nextest/libtest UI (rust-analyzer, RustRover, cargo nextest run): it lists scenarios via proef flows --format json and runs each as its own test via proef test --scenario <name>.

PROEF_HARNESS_SUITE=suite cargo nextest run -p proef-harness

Set PROEF_BIN=/path/to/proef when the binary is not on PATH. With PROEF_HARNESS_SUITE unset the harness exposes nothing, so a plain cargo test stays green. Each scenario appears as a separate test in the IDE’s runner — click-to-run one scenario without touching the terminal.

Unset is not the same as unreadable. A variable set to bytes that are not valid UTF-8 means you asked for something the harness cannot read, so it exposes a single failing proef::config trial naming the variable — it will not fall back to proef on PATH, and it will not report green having listed no tests.

Secrets in CI

Two working setups:

  • Values via env (simplest): set PROEF_SECRET_<NAME> from your CI’s secret storage — no store, no key, nothing on disk.
  • Committed ciphertext store: commit .proef-secrets.json (it holds ciphertext only) and supply the project key as the PROEF_KEY CI secret (base64 < ~/.config/proef/keys/default.key). A set-but-invalid PROEF_KEY is always an error, never a silent fallthrough.

Note: .proef-state.json (the persistent World) is plaintext and therefore gitignored — captures promoted with saveAs: global land there. A capture whose value equals a known secret is refused with a warning; it never persists (proef doctor also reports store/key health).

Environment variables

VariableRead byPurpose
PROEF_ENVproef test/flows/macros/artifacts/lspActive environment profile — the --env flag wins over it
PROEF_SECRET_<NAME>proef testSecret value override (beats the encrypted store)
PROEF_KEYproef test/secretBase64 project key override — decrypt a committed store without the key file
PROEF_CONFIG_DIRproef secretKey-file location (default: XDG config dir)
PROEF_HARNESS_SUITEnextest harnessSuite directory the harness lists and runs
PROEF_BINnextest harnessPath to the proef binary the harness invokes

Set-but-unreadable is an error, never silence. A variable whose value is not valid UTF-8 is something you asked for and proef cannot read, so it is reported rather than treated as unset: the commands above exit 2 naming the variable (proef doctor reports it as a failed check and exits 3, with its other environment findings), and the harness exposes a failing proef::config trial. Reading a malformed PROEF_KEY as absent used to mean decrypting with the wrong key and reporting tampering; a malformed PROEF_ENV meant running against the wrong environment. PROEF_CONFIG_DIR is a path and is read as raw bytes, so it has no such failure mode.

Suite-defined env vars (like PROEF_BASE_URL in the guides) are a convention of ${env:…} references in proef.toml, not built-ins — name yours freely.

Style

Keep prose at business level — no URLs, headers, or JSON in feature files; that’s what packs are for. Prefer several small macros composed with use: over one large one. Give steps name: labels — they anchor artifacts, events, and failure output.

Label a macro’s steps whenever it has more than one. A macro with three steps turns one feature sentence into three engine steps that share a step anchor exactly — same file, same line, same text — so on the console, in the HTML report, in JUnit, TAP and the job summary they arrive as the same sentence repeated, one row per step. The name: is the only thing that distinguishes them:

✓ tests/features/a.feature:9 — the workspace is provisioned › fixture warm-up probe (4ms)
✗ tests/features/a.feature:9 — the workspace is provisioned › provision the environment (7ms)

Without the labels those two lines differ only by their glyph.

proef in CI

Everything proef’s differentiators buy — typed exit codes, JUnit, sharding, the rerun overlay, the regression gate — pays off in CI, and each piece is documented on its own page. This page is the missing last mile: one paste-ready workflow, then the pieces it composes.

A complete GitHub Actions workflow

name: e2e
on: [push, pull_request]

jobs:
  e2e:
    runs-on: ubuntu-latest
    strategy:
      fail-fast: false
      matrix:
        shard: [1, 2, 3]
    steps:
      - uses: actions/checkout@v4
      - name: Install proef
        run: |
          curl -LsSf https://github.com/cargo-bins/cargo-binstall/releases/latest/download/cargo-binstall-x86_64-unknown-linux-musl.tgz | tar -xz -C /usr/local/bin
          cargo-binstall proef --no-confirm
      - name: Run the suite
        env:
          # Secrets reach proef only through the environment — never files,
          # never flags (values would land in the process listing).
          PROEF_SECRET_APITOKEN: ${{ secrets.API_TOKEN }}
        run: |
          proef test tests/features \
            --env staging \
            --shard ${{ matrix.shard }}/3 \
            --junit auto \
            --meta commit=${{ github.sha }}

The pieces:

  • --junit auto writes report.junit.xml into the run dir — under GITHUB_ACTIONS only, so local runs stay clean. Failures also reach the job summary and the PR’s changed-files gutter as ::error annotations automatically; no extra step.
  • --ctrf report.ctrf.json writes the same verdicts as CTRF JSON — the format that carries what JUnit XML cannot: retries and flakiness natively (a pass-after-retry lists its real failed attempts with messages), tags, and a file path per test. Point a CTRF consumer such as ctrf-io/github-test-reporter at it. The two files never disagree: a quarantined failure is skipped with a message in both (ADR-0019), because a dashboard reading “failed” beside exit 0 would contradict itself. Like --junit, the file is written even when [run] setup aborts the run — a job gating on it must never see it missing.
  • --shard I/N partitions by a stable hash of (file, scenario): adding a scenario never re-buckets the others, so shard timings stay comparable across commits. Every matrix job runs the same expression and the shards partition exactly. To balance by time instead of by count, see below.
  • --meta commit=… records provenance the run cannot harvest itself: proef never reads GITHUB_SHA or any CI variable (ADR-0020) — what the workflow hands over explicitly is what the record carries.
  • PROEF_SECRET_<NAME> supplies ${secret:name} values. They never appear in artifacts, events, logs, or reports — including base64/hex/ percent-encoded reflections.

Balancing the matrix by duration

A hash split balances by count, and a matrix finishes when its slowest shard does — so one long scenario can leave three runners idle. Every run that reaches its suite writes a small timings.json into its run directory; archive it once and hand it to the next run’s matrix:

      - name: Run the suite
        run: |
          proef test tests/features \
            --shard ${{ matrix.shard }}/3 --shard-weights timings.json
      # one job publishes the file the next run's matrix reads
      - uses: actions/upload-artifact@v4
        if: matrix.shard == 1
        with:
          name: proef-timings
          path: .proef-runs/*/timings.json

Every job must read the same file. The tempting shortcut — letting each job weight by its own local run history — is silently wrong: matrix jobs run on different machines with different (usually empty) runs-dirs, so each would compute a different assignment and scenarios would run twice or not at all while the suite still reported green.

A scenario the file does not mention falls back to the frozen hash, so a test added after the timings were captured still runs exactly once. A missing or malformed file is exit 2 — never a silent fall back to the unbalanced split. Full detail in CONFIG.md.

Gating on regressions between runs

proef diff compares two run records; --fail-on-regression makes it a gate. Download the base branch’s record (uploaded as an artifact by its own run) and compare:

proef test tests/features --junit auto        # today's run
proef diff path/to/base-events.jsonl          # vs the downloaded baseline
# exit 1 on a regression; new flakiness and perf deltas print either way

Continuing a cancelled run

A run stopped by --max-fail, a runner timeout, or a docker stop records what it never reached. SIGTERM and SIGHUP take the same graceful path as Ctrl-C: in-flight batches finish, the rest record as skipped, [run] teardown runs, the reports are written, and the record closes with run_finished + cancelled (exit 1) — so a timed-out job still leaves a complete record and its JUnit. Only a second signal (exit 130) or a SIGKILL truncates it.

proef test --rerun re-runs the last run’s failures and the scenarios it never got to — and its JUnit and HTML report cover the whole suite via the rerun overlay (rerun_of in the record), so one report stands for the composed result, never a false green.

Flakiness over history

Keep a few records ([run] keep-runs in proef.toml sets the rotation) and proef flaky renders verdicts over them — flapping, passes-only-on-retry, always-failing — from the same records CI already produced.

It also audits @quarantine itself, which nothing else can: a quarantined scenario failing every run is DISABLED (switched off — its failures gate nothing, so no job ever reports them), and one green throughout is recovered (the tag can come off). --format json carries quarantined and the verdict key, so a scheduled job can gate on either without parsing the table.

--by <key> splits the same history per run context — --by env, or any [meta]/--meta key such as --by runner. A scenario that flaps in one environment and is solid in another is not flaky but context-dependent, and a pooled history cannot tell those apart: it reports the one conclusion the merged view can never reach, naming the scenarios whose verdict changes with where they ran. A run that never set the key is its own (unset) bucket rather than being folded in with the runs that did.

Verdicts are keyed by the run’s input fingerprint — inputs.json beside each record, a hash of the feature sources, the loaded macros and fragments, and the resolved ${url:…}/${vars:…} scope — so a pack or config edit starts a fresh window instead of mixing runs of different inputs; --by splits within it. Three guards, each a [flaky] key with a flag twin (CONFIG.md): a scenario seen in fewer than min-samples runs (default 10) reads insufficient-data rather than earning a verdict — a fresh matrix needs ten records before the table says anything, and --min-samples 2 restores the pre-0.18 floor; a flagged scenario stays flagged until recovery-runs (5) trailing clean runs, so it cannot flip between adjacent runs; and a run in which more than outage-rate (0.8) of the suite failed is an environment incident, excluded wholesale, so one staging outage cannot mark the suite broken.

Editor setup — proef lsp

proef lsp starts a Language Server Protocol server that speaks generic LSP over stdio. Point any LSP-capable editor at proef lsp and get, live as you type:

  • Diagnostics — the same validation proef test --dry-run performs (unbound steps, unknown step kinds, malformed hurl blocks, unresolved ${…} references), published across the whole suite as you edit.
  • Go-to-definition — jump from a Gherkin step to the macro that binds it, or from a pack’s ref: to the # @proef annotation in the .hurl file (ADR-0018).
  • Completion — step completions offering the suite’s macro patterns; on a pack’s ref: line, the fragment names in scope; and inside a bind: table, the {{variables}} those fragments actually read, each labelled with the fragment that wants it. A variable a fragment supplies itself ([Options] variable:) is left out: it needs no bind:, and binding it would be refused as option_declared_twice.
  • Find-references — every step across the suite that a given macro binds.
  • Document symbols — a feature outlines to its scenarios (tagged ones showing their tags), a pack to its macros (detailed by the pattern each matches). Which vocabulary applies is decided by what discovery found in the file, not by its extension.
  • Hover — the macro a step binds, or a use: targets, with its pack, pattern and params; on a ref:, the fragment’s file and the variables still needing a bind:. Every fact comes from the same analysis the diagnostics do, so a hover can never contradict the squiggle on the same line.
  • Semantic tokens — the two variable tiers, told apart on screen. Inside a pack’s hurl: | block, ${…} (resolved at lower time, by proef, before any request exists) and {{…}} (resolved at run time, by hurl) look identical to every editor: a YAML highlighter sees a string, and a hurl highlighter never runs because the block is not a file. proef is the only party that knows which is which. ${…} is reported as a macro — a substitution performed before execution, which is what a macro is — and {{…}} as a variable; every mainstream theme colours those differently already, so nothing needs configuring. A $${ escape stays unhighlighted, because it is a literal ${ proef will not substitute.
  • Quick fixes — a misspelled name that already earned a “did you mean” becomes an applicable edit: use: and ref: targets, with: and bind: keys, step kinds, Examples placeholders, and data-table columns. A fix is offered only when it is certain — the suggested name is near enough, and the misspelling occurs exactly once as a whole token in that same file — so applying one is never a guess. Reach it from the squiggle or from the token itself; a use: error underlines the macro’s name key, which is often several lines above the word you typed.

The server analyzes the configured suite — proef.toml’s [run] suite if set, else the tests/ convention — resolved under the directory it is launched in (its working directory), discovering every .feature file and every packs/*.yaml / packs/*.yml macro pack beneath that root — the same resolution proef test uses, so the two never diverge. Launch your editor from the project root (or configure the server’s root/working directory to it).

When analysis fails

A panic inside analysis or a feature never ends the server. proef reports it once through window/showMessage — the channel an editor actually surfaces — and keeps serving; the next edit retries, and reports again only if the new state also fails. A server that died would show nothing, which reads as “proef has no opinion about this file” rather than as the failure it is.

Running proef lsp by hand also prints the panic to stderr, which is where the detail lives.

Naming the config: --config

proef lsp --config <path/to/proef.toml> names the config to read instead of searching for one, exactly as it does for every other subcommand (CONFIG.md). Reach for it when discovery cannot find the right file — most often a config that sits beside the suite rather than above it, which an upward search launched from the repository root can never reach.

For proef lsp the flag also outranks the workspace root the client announces: the flag names a file, and a named file is not a guess to be improved on. Without it, an editor rooted somewhere else loads a different config than the runner, and the diagnostics stop being trustworthy in exactly the layout the flag exists for — proef test --config … runs green while every ref: reads as unknown in the editor.

cmd = { "proef", "lsp", "--config", "/abs/path/to/proef.toml" },

Unlike the runner, an unreadable or absent file does not stop the server: it starts on defaults, because an editor offering less is better than one that will not boot. A relative path works, but prefer an absolute one — an editor’s working directory is rarely the one you assume.

--env <name> travels the same way (proef lsp --env staging, or PROEF_ENV): it selects the [env.<name>] profile the analysis resolves ${url:…}/${vars:…} against, so the editor reports the missing_config_var a staging run would hit and not the one the default profile would. Before 0.18 the flag was accepted for lsp and silently dropped.

File types served

KindPatternTypical editor filetype
Feature files*.featuregherkin / cucumber / feature
Macro packspacks/*.yaml, packs/*.ymlyaml

Neovim

Built-in LSP client (Neovim 0.8+), no plugin required. Add to your config and open a .feature file from inside the suite:

vim.api.nvim_create_autocmd("FileType", {
  pattern = { "cucumber", "gherkin", "yaml" },
  callback = function(args)
    vim.lsp.start({
      name = "proef",
      cmd = { "proef", "lsp" },
      -- The server scopes analysis to the configured suite under its launch
      -- directory; anchor root_dir at the nearest proef.toml (or the current
      -- file's directory) so the two agree.
      root_dir = vim.fs.dirname(
        vim.fs.find({ "proef.toml" }, { upward = true, path = args.file })[1]
      ) or vim.fs.dirname(args.file),
    })
  end,
})

With nvim-lspconfig you can instead register it as a custom server via vim.lsp.config/configs, using the same cmd = { "proef", "lsp" }.

Helix

Add a server and attach it to the languages you author. In ~/.config/helix/languages.toml:

[language-server.proef]
command = "proef"
args = ["lsp"]

[[language]]
name = "gherkin"
language-servers = ["proef"]

# Attach to pack YAML too, so pack diagnostics surface while editing macros.
[[language]]
name = "yaml"
language-servers = ["proef", "yaml-language-server"]

Run hx --health gherkin to confirm Helix found the proef binary.

Emacs (Eglot)

Eglot ships with Emacs 29+. Associate your feature-file major mode (for example feature-mode) with the server:

(with-eval-after-load 'eglot
  (add-to-list 'eglot-server-programs
               '(feature-mode . ("proef" "lsp"))))
;; M-x eglot in a .feature buffer opened from the suite root.

For lsp-mode, register a stdio client whose connection is (lsp-stdio-connection '("proef" "lsp")) for the same major mode.

v1 limitations

This is the first release of the language server. Known boundaries:

  • proef.toml config is a startup snapshot. ${url:…} / ${vars:…} values are read once when the server starts. After editing proef.toml (or switching PROEF_ENV), restart the server to pick up the change; until then those references analyze against the old (or, if the file could not be loaded at startup, an empty) scope and may warn.

  • Built-in macros have no jump target. The expect* family lives in a pack compiled into the binary, not a file on disk, so there is nothing for go-to-definition to open. Hover still answers — the macro is in the analysis like any other, and reports its pack as builtin:…, which is why the jump is unavailable. proef macros lists the whole family with the sentence each binds.

  • Completion ranking is best-effort. All of the suite’s macros are offered; ranking is a lightweight edit-distance heuristic. Full context-aware ranking is a follow-up.

  • External edits to closed files need a reopen. The server does not watch the filesystem; it re-reads a file’s bytes from its open editor buffer. If a .feature or pack file is changed outside the editor (or by another tool) while closed, reopen it so the server sees the new bytes.

  • No VS Code extension yet. v1 is a server-only generic-LSP binary. It works with any editor that speaks generic LSP (Neovim, Helix, Emacs, Sublime LSP, …); a VS Code wrapper is a possible follow-up.

    If one is built, its documentSelector must match on path ({ scheme: "file", pattern: "**/*.feature" }), not on a language id. The ecosystem is split — the two established Gherkin extensions register cucumber and feature respectively — so a selector naming either id attaches for some users and silently does nothing for the rest. The table under File types served is that split; a path selector is the only thing all of it has in common.

  • No Zed support. Zed binds language servers to languages it has a tree-sitter grammar for, and there is no Gherkin grammar in it. That grammar is a prerequisite, not a configuration step, so Zed waits on work outside this repository.

  • Overlay lookup can still miss if the suite root is reached through a symlink. The root is deliberately left uncanonicalized (canonicalizing would resolve symlinks and desync source names from the client’s document URIs), so if an editor resolves a symlinked suite root differently than the server’s raw working directory, the overlay lookup can miss and the LSP analyzes the saved on-disk bytes instead of the unsaved buffer. An editor’s percent-encoding choice no longer matters here — the overlay matches open buffers by decoded source name, not the raw URI.

Troubleshooting

First stop, always: proef doctor (native libraries, environment, secret store/key health) and proef test <suite> --dry-run (full validation with located diagnostics, no network). Every diagnostic code is indexed in DIAGNOSTICS.md.

Reading the output

Exit codes are a contract:

CodeMeaningTypical fix
0everything passed (warnings allowed)—
1at least one check failed: a test assertion, a cancelled run — or a --check-style gate (fmt --check, fragments --check, diff --fail-on-regression) that found what it gates onfix the system under test — or the expectation
2your input is at fault: packs, features, flags, filters, secrets, bad {{var}}/JSONPaththe diagnostic names the file and line
3the environment or proef is at fault: unreachable target, native libs, IO, output proef could not write (full disk, failing device)check the target, proef doctor, disk
130interrupted twice — the second signal (Ctrl-C, SIGTERM or SIGHUP) is a hard exit (128 + SIGINT), so cleanup and the record’s tail are skippedthe run record will read as incomplete; a single Ctrl-C — or a single SIGTERM/SIGHUP, e.g. a CI job timeout — cancels gracefully and still runs [run] teardown

Step glyphs:

GlyphStatusMeaning
✓passedran, all asserts held
✗failedran, an assert failed (details + reproduce: line at the end)
∅skippednot run: when: guard, an earlier failure, cancellation, or an authored @skip on the scenario (the tag spelling prints as the reason)
⚠warnedan optional: step failed, or a saveAs: global promotion was refused — the reason prints on the ↳ line

Under --console dotted the tree collapses to one glyph per scenario — . passed, F failed, s skipped, w warned (lowercase = non-gating) — and failures still print in full after the pool; --console quiet keeps only the run line and the summary. On a terminal the status vocabulary is colored; NO_COLOR (or a non-terminal stream) turns it off, and run.log never carries the paint either way. The record and the exit code are identical in every mode. Every run’s last line names its run id and wall-clock — the id is the reproduction key --shard, --shuffle and ${fake:…} all hang off.

Frequent situations

“missing secret value(s)” — the suite references ${secret:NAME} with no value available. Fix: proef secret set NAME (or export PROEF_SECRET_<NAME>=…). In CI, see AUTHORING — Secrets in CI.

An assertion fails on values that look identical — they differ only in whitespace, and proef says so:

Assert failure (f--s.hurl:9: actual: string <ok> expected: string <ok >)
  [whitespace only — actual: string·<ok>, expected: string·<ok·>]

The bracketed note appears only when the two values differ solely in whitespace, and repeats them with every whitespace character drawn: · for a space, \t / \r / \n for the usual escapes, and \u{a0} for the exotic ones — a non-breaking space pasted out of a browser being the classic. Usual causes: a trailing space in the expectation, a CRLF fixture leaking \r into a body, or a tab where the author typed spaces.

“unknown environment <name>” — --env <name> (or PROEF_ENV) names an environment proef.toml doesn’t define; the error lists the known ones. Fix the name or add the [env.<name>] section. Exit 2.

“<url|vars> variable <key> is not set” (resolve::missing_config_var) — a pack references ${url:key} / ${vars:key} that neither the base [url]/[vars] nor the active [env.<name>] defines. Add it, or select the environment that has it with --env. See CONFIG.md.

“no path given and no default suite found” — proef test got no path and there is no [run] suite in proef.toml nor a tests/ directory. Pass a path, set [run] suite, or create tests/. Exit 2.

Exit 3 with connection errors — the target is unreachable. Check the URL your ${url:base} resolves to for the active --env (a --dry-run prints nothing wrong because no request is sent), then the network. The dev fixture (cargo run -p xtask -- fixture) binds the default ${url:base} port (8787), so with no PROEF_BASE_URL set it becomes your local target automatically (it falls back to an ephemeral port, printing PROEF_BASE_URL, only if 8787 is busy).

“batch budget exceeded — scenario thread abandoned” — the watchdog killed a batch that outran its computed budget (timeouts × attempts + delays + repeats + margin, ADR-0007). Usually a huge retry:/delay: (or a hurl-side [Options] repeat:/ retry-interval:/max-time:) value — the pack lint caps literals (counts at 10 000, durations at one hour), and the computed budget itself clamps at four hours, so a lint-clean product that would otherwise run for days is abandoned at the ceiling. A {{var}}-driven retry:/delay:/[Options] repeat:/max-time: cannot be estimated at all — it resolves inside hurl at run time — so the batch falls back to the default budget ([http] timeout × 4, at least 60s) rather than to an estimate that assumes no retries. If a legitimately long templated retry is being abandoned, raise [http] timeout.

“global state file .proef-state.json is not valid JSON” — the persistent World is derived data (only saveAs: global promotions live there). Deleting .proef-state.json is safe; the next run recreates it.

Corrupt .proef-secrets.json — proef doctor names it; the next proef secret set moves the wreck to .proef-secrets.json.corrupt and starts fresh. Values are re-enterable; nothing else references the file.

“no scenarios matched the filters” — --tags/--scenario selected nothing; exit 2 by design so a typo’d filter can never produce a silent green CI run.

A failure line carries (via tests/hurl/admin.hurl#admin.search) — not an error. The step ran a named fragment (ADR-0018) rather than an inline hurl: block, and that is the third file involved: the request lives there, not in the feature or the pack. The spelling is the one ref: accepts, so it pastes straight back into a pack. Every failure sink carries it — console, explain, TAP, JUnit, the GitHub summary, the HTML report.

“ref: names no loaded fragment” (pack::unknown_ref) — either the name is a typo (the message suggests the closest) or no fragment files were loaded at all, which the message says plainly. The usual cause is a missing [run] fragments in proef.toml: without it nothing is scanned, so every ref: is unknown. Note the root resolves against the config file’s directory, not your working directory.

“reads name, which nothing supplies” (lower::unbound_placeholder) — a fragment’s variables are not implicit: what the .hurl file reads must be bound at pack, macro or step scope, captured by an earlier step, supplied by the fragment’s own [Options] variable:, or carried by a secret of that name — and inside a bind: value, an earlier-sorting sibling literal counts too (injected lines evaluate in name order). --dry-run reports it without a network. With proef lsp running, completing inside a bind: table offers what the pack’s ref:ed fragments still need — the union over every fragment the pack refs (ranked by the nearest ref:, each labelled with its owner), minus the names a fragment supplies itself.

“index is supplied twice” (pack::option_declared_twice) — the fragment sets the name in its own [Options] variable: and a bind: supplies it. Both reach the entry as variable: index= and hurl takes the last, which is the fragment’s — so the bound value would silently never be sent. Delete whichever is not authoritative: the bind: if the file’s own default is right, the variable: line if the pack should decide. The same rule already applies to retry:/delay:.

Go-to-definition on a ref: does nothing — check that [run] fragments is set and that the editor’s workspace root matches where proef.toml lives; the server resolves the corpus relative to that file’s directory.

“payload does not parse: contains no hurl entries” — the hurl: block is comments/blank lines only (a temporarily commented-out request). proef refuses it at load: a step that executes nothing must not report green.

A Then step shows ∅ not run (its request entry did not run) — the request it asserts on was skipped (guard or earlier failure); the asserts had nothing to attach to.

Windows/| head pipelines — closed pipes are tolerated everywhere; if a pipeline misbehaves, check the consumer, not proef’s exit code.

Building from source

proef-engine-hurl links native libraries. Debian/Ubuntu: apt install build-essential pkg-config libssl-dev libcurl4-openssl-dev libxml2-dev libclang-dev. macOS: Xcode CLT suffices. Verify with proef doctor — it reports the embedded hurl, parser/libxml2 linkage, and libcurl. Prebuilt binaries (brew/binstall/GitHub Releases) need none of this.

Digging deeper

Every run leaves .proef-runs/<run-id>/: events.jsonl (the machine record — EVENTS.md), run.log (console mirror), and artifacts/ with the exact executed .hurl files. proef explain summarizes the latest run, proef diff [base] [new] compares two of them (regressions, fixes, flakiness, perf), and proef report [run] writes a self-contained HTML page of a run; the reproduce: hurl --test … line under a failure replays the artifact with stock hurl, taking proef out of the loop entirely.

A cancelled run (Ctrl-C) is a complete record, not a truncated one — proef explain/proef report never banner it as incomplete, because its nonzero skipped count already says what didn’t get to run. proef diff --fail-on-regression disagrees on purpose: it still refuses to certify “no regressions” against a cancelled run, since a regression could be hiding among the scenarios it never reached. The three commands’ differing treatment of cancelled is a deliberate choice, not an inconsistency.

proef diff takes each side as a run id, a record directory, or an events .jsonl file — the stream is the record, so a file means the same thing under any name. That third form is the CI baseline flow: upload .proef-runs/<id>/events.jsonl as an artifact on your main branch, download it in the PR job as (say) baseline.jsonl, then

$ proef test tests/features --run-id pr
$ proef diff baseline.jsonl pr --fail-on-regression   # exit 1 on passed → failed

and a scenario that regressed against main fails the PR without a run record store shared between jobs.

proef diff keys steps by (text, ordinal) so that line shifts don’t lie and repeated steps stay distinct. The trade-off is positional: if a scenario loses an earlier duplicate of a step, every later instance shifts down one ordinal, and the comparison lines up two different runs’ steps. Timing and attempt counts for the shifted steps can then be attributed to the wrong one. Renaming or reordering steps between the two runs being compared is worth a second look at the numbers; adding or removing steps at the end is not affected.

proef.toml — the project configuration reference

proef.toml lives in the project root — found by searching up from the working directory, so any subdirectory works (a project is where its proef.toml is, not where your shell is) and is committed project config — it describes the suite, not your machine. Every key is optional; an absent file means all defaults. Unknown keys are rejected (deny_unknown_fields), so typos fail loudly instead of being ignored.

The file holds a few kinds of thing, kept in distinct sections: runner config (how proef behaves — [run], [http], [sla]), suite variables (data your tests reference — [url], [vars]), run metadata ([meta], recorded in the run head) and report tag links ([tag-links]), plus per-environment overrides ([env.<name>]). Secrets never live here (see below).

Precedence (highest wins): built-in defaults < proef.toml base tables < active [env.<name>] < command-line flags. For a suite variable that means [url]/[vars] < [env.<active>]; for a runner setting like jobs, the --jobs flag still wins over both.

Reference

# ── runner config (how the tool behaves) ────────────────────────
[run]
suite    = "tests"          # default suite path — `proef test` needs no argument
fragments = "tests/hurl"    # root of the hurl files packs may `ref:` (optional)
jobs     = 8                # parallel scenario workers
runs-dir = ".proef-runs"    # where run records land
keep-runs = 200             # past records kept there (0 = only the run in flight)
setup    = "tests/setup.feature"      # run once before the pool (optional)
teardown = "tests/teardown.feature"   # run once after the pool (optional)
exclusive-tags = "@serial"  # scenarios matching this run with the pool to themselves

[http]
timeout-ms      = 30000     # per-request timeout (batch-level default)
follow-location = false     # follow redirects
max-redirs      = 10        # redirect ceiling (only meaningful with follow-location)
user-agent      = "proef"   # sent with every request — how a server spots test traffic
cookie-store    = true      # false = run cookie-less (hurl's --no-cookie-store):
                            # prove a stateless API by refusing to carry a session
# TLS and proxy: the settings that differ *between environments*, which is why
# they belong here and not in a pack. Paths resolve against this file.
insecure        = false     # skip certificate verification (warns loudly when on)
cacert          = "certs/ca.pem"       # trust this CA instead of the system store
client-cert     = "certs/client.crt"   # mTLS: the certificate to present
client-key      = "certs/client.key"   # mTLS: its private key
proxy           = "http://proxy.corp:3128"
no-proxy        = "localhost,.internal"

[sla]                       # opt-in run-level latency budget (omit = no gate)
p95-ms = 250                # 95th-percentile per-step duration ceiling
max-ms = 1000               # slowest single-step ceiling

# ── suite variables (data your tests reference) ─────────────────
[url]
base = "http://127.0.0.1:8787"   # → ${url:base}; the dev fixture's default port

[vars]
apiVersion = "v1"                 # → ${vars:apiVersion}
username   = "dev@example.com"    # → ${vars:username}

# ── per-environment overrides (mirror the base tables) ──────────
[env.staging.url]
base = "https://staging.example.com"

[env.prod.url]
base = "https://api.example.com"
[env.prod.vars]
username = "release@example.com"
[env.prod.http]                   # an environment may override a runner setting too
timeout-ms = 60000
KeyDefaultNotes
[run] suite(unset)default path for proef test/flows/macros/fragments/artifacts (and what doctor and lsp inspect); falls back to the tests/ convention, then errors. An explicit path always wins
[run] jobsavailable parallelism--jobs flag wins; live threads never exceed the scenario count
[run] runs-dir.proef-runsrun records land here; rotation only ever deletes uuid-named dirs — see the --run-id note below
[run] keep-runs200how many past records runs-dir retains, besides the one being written; 0 keeps none but the run in flight
[run] setup(unset)feature run once before the pool (suite setup); its saveAs: global reaches every scenario; a failure aborts the run
[run] teardown(unset)feature run once after the pool (suite teardown), only if setup succeeded; its failure is a distinct exit 3
[run] exclusive-tags(unset)tag expression (same language as --tags) selecting scenarios that run with the pool to themselves; a malformed expression is exit 2
[http] timeout-ms30000per-entry [Options] in a hurl block override it; 0 is refused (exit 2) — libcurl reads zero as no timeout, the unbounded hang the default exists to prevent
[http] follow-locationfalseper-entry [Options] override it
[http] max-redirs(engine default)redirect ceiling; only reached when follow-location is on
[http] insecurefalseskip TLS certificate verification. The run prints a warning naming the profile whenever this is on — a green run that verified nothing is not the same result, and nothing else would say so
[http] cacert(system store)CA bundle to trust instead of the system store. Path — resolves against proef.toml
[http] client-cert(unset)client certificate for mTLS. Path. A combined PEM may carry the key too
[http] client-key(unset)private key for client-cert. Path. Setting it without client-cert is exit 2 — a key alone presents nothing, and the failure would otherwise surface at the server as an unexplained auth error
[http] proxy(unset)proxy URL for every request
[http] no-proxy(unset)comma-separated hosts bypassing proxy (curl’s NO_PROXY form)
[http] user-agent(engine default)User-Agent for every request. Run-wide only — hurl has no per-entry user-agent option, so one entry opts out with a User-Agent: header instead
[http] cookie-storetruefalse runs the suite cookie-less (hurl’s --no-cookie-store): no Set-Cookie is kept, none is replayed — how a stateless API is proven stateless. Run-wide only — hurl has no per-entry spelling for it at all
[sla] p95-ms(unset)95th-percentile per-step duration ceiling; unset = no gate
[sla] max-ms(unset)slowest single-step ceiling; unset = no gate
[flaky] min-samples10minimum observed runs before proef flaky classifies a scenario (else insufficient-data)
[flaky] recovery-runs5trailing clean runs that resolve a flagged scenario to healthy (hysteresis)
[flaky] outage-rate0.8a run failing over this share of suite scenarios is an environment outage, excluded from the history
[url] <key>(none)URL variables, referenced as ${url:<key>}
[vars] <key>(none)non-secret variables, referenced as ${vars:<key>}
[meta] <key>(none)run metadata recorded in the run head; [env.<name>.meta] overrides it, --meta flags win (ADR-0020)
[tag-links] "<glob>"(none)tag glob → URL template ({tag} substituted); matching tags link out from the HTML report and GitHub summary. Base table only
[env.<name>.<section>]inherits baseper-environment override of any base section (url/vars/http/run/sla/meta)

A custom --run-id is kept, but never rotated

--run-id accepts any single path component, and a run named that way writes a perfectly good record that explain, diff, flaky, report and --rerun all find — a directory is a record because it holds an events.jsonl, not because of the shape of its name (ADR-0021).

“The latest run” is resolved by time rather than by spelling, and where that time comes from differs by name: a generated id is a uuid-v7, which carries the moment proef minted it, while a custom id carries no time at all and orders by the directory’s modification time instead. Both are wall-clock, so the two compare directly — but a custom-id record that is later copied or restored sorts by that moment, not by when it ran.

What a custom id changes is rotation: proef deletes only uuid-named directories, so [run] keep-runs will never reclaim .proef-runs/pr-1234/. A caller minting a fresh custom id per build accumulates records without bound. That is deliberate and the asymmetry is the point — runs-dir may be ., and guessing at user-named directories is the worse failure by a wide margin. Reading a directory is safe; deleting one is not.

Use --run-id when you want a stable, externally-known name — a CI job archiving .proef-runs/pr-1234/ by path — and clean those up the same way you created them.

Environments

An [env.<name>] profile mirrors the base tables and deep-merges over them, key by key — so an environment lists only what changes; everything else inherits from the base (the Cloudflare-Wrangler / Cargo-[profile.*] model). Select the active environment at run time:

$ proef test                 # base [url]/[vars]; discovers the default suite
$ proef test --env prod      # [env.prod.*] merged over the base
$ PROEF_ENV=staging proef test   # same, via the environment variable

The rule is almost uniform: under [env.<name>], url.* / vars.* override variables and http.* / sla.* override runner settings wholesale — but [env.<name>.run] accepts only jobs. Every other [run] key (suite, runs-dir, setup, teardown, exclusive-tags, …) names what the suite is, not how hard to run it, and letting an environment swap those made two environments run different tests while claiming to be the same suite — so listing one there is a hard parse error (exit 2), never a silent no-op. A named-but-undefined --env is a user error (exit 2) listing the known environments. A ${url:key} / ${vars:key} referenced in a pack but defined in neither the base nor the active environment fails at lower time (proef::resolve::missing_config_var).

TLS, proxies and mTLS ([http])

These are the settings that describe an environment rather than a test, which is why they live here and not in a macro pack. Staging presents a self-signed certificate, CI traverses a corporate proxy, production speaks mTLS — and the same suite has to run against all three without a single pack changing. Put the differences in [env.<name>.http] and the shared parts in [http]:

[http]
user-agent = "proef-suite"          # every environment identifies test traffic

[env.staging.http]
insecure = true                     # staging's certificate is self-signed
proxy    = "http://proxy.corp:3128"

[env.prod.http]
cacert      = "certs/prod-ca.pem"   # a private CA, not the system store
client-cert = "certs/client.crt"    # mTLS
client-key  = "certs/client.key"

Three things worth knowing before you rely on this:

  • Paths resolve against proef.toml, not the working directory — the same one-path rule every other config path follows, so proef test from a subdirectory finds the same CA bundle it finds from the project root.
  • insecure = true warns on every run, naming the profile that set it. A suite that goes green without verifying a certificate has not proved what a green suite normally proves, and the run record deliberately carries no config, so the warning is the whole audit trail.
  • A per-entry [Options] block still wins — for all of these but one. These are batch-level defaults, so one request that must skip verification, or present a different certificate, says so in its own hurl block and overrides the table for itself alone. The exception is user-agent: hurl has no per-entry option for it, so [http] user-agent is run-wide and a single entry cannot opt out. Set a User-Agent: request header in that entry instead.

Credentials stay out of this table on purpose. There is no [http] user or netrc key: a password belongs in the secret store (proef secret set), where it is encrypted at rest and masked out of artifacts, events, logs and reports. A proxy URL is the one edge — if yours embeds credentials, it is plaintext in a file you probably commit, and it belongs in an environment variable your shell expands into the proxy’s own configuration instead.

Balancing a shard matrix by duration (--shard-weights)

--shard I/N assigns scenarios by a frozen hash of (file, scenario), so adding one test never re-buckets the others. What it cannot do is balance by time — and a CI matrix finishes when its slowest shard finishes, so a count-balanced split routinely leaves runners idle.

Every run that reaches its suite writes timings.json into its run directory (a run whose [run] setup aborted has no suite to weigh, and writes none). Archive that one file and hand it to the next run’s matrix:

# each matrix job, all reading the same archived file
- run: proef test --shard ${{ matrix.shard }}/4 --shard-weights timings.json

Every job must read the same file. The tempting alternative — “let each job weight by its own newest local record” — is silently wrong: matrix jobs run on different machines with different (usually empty) runs-dirs, so each would compute a different assignment and scenarios would run twice or not at all while the suite still reported green.

Two rules place scenarios, and they partition rather than compete:

  • a scenario the file mentions is placed by the balanced split (longest-processing-time-first: heaviest to the lightest shard, repeatedly);
  • a scenario it does not mention falls back to the frozen hash.

So a test added after the timings were captured still runs exactly once. That property is pinned by a test that runs a whole matrix and asserts set equality both ways.

What you give up is the thing hash mode was chosen for: a balanced split is not stable under insertion, so adding a slow scenario can move others. That is what balancing means, and it is why the flag is opt-in rather than the default. A missing or malformed weights file is exit 2, never a silent fall back to the unbalanced split you passed the flag to avoid.

SLA gate ([sla])

The optional [sla] table is a run-level latency budget: after a run, proef folds every step’s wall-clock duration into two aggregates and fails the run if either exceeds its ceiling. p95-ms caps the 95th-percentile step duration (the typical slow request); max-ms caps the single slowest step. Skipped steps (which never hit the network) are excluded from the population.

The gate is opt-in and off by default — with no [sla] table a run behaves exactly as before. A breach prints the offending metrics and the slowest steps on stderr and maps to exit 1 (a test failure), the same code as a failed assertion. It never introduces a new exit code and never downgrades a User/System fault (exit 2/3), so a breach can only turn an otherwise-green run red. Tighten or loosen per environment the usual way:

[sla]                       # baseline budget
p95-ms = 250
max-ms = 1000

[env.staging.sla]           # staging tolerates slower responses
p95-ms = 800

This is distinct from hurl’s per-request duration < <ms> assert (which fails one step inside a hurl: block): the [sla] gate is an aggregate budget over the whole run, while duration < is a targeted per-request check. Use the assert for a hard per-endpoint SLA and [sla] for a suite-wide budget.

Hurl fragments ([run] fragments)

Names the root directory holding the .hurl files a pack may ref: (ADR-0018), scanned recursively. Those files stay valid hurl — the # @proef <name> annotation is an ordinary comment — so the same file runs under hurl and under proef, and a corpus you already own is annotated once instead of transcribed into YAML where it drifts.

[run]
fragments = "tests/hurl"
# tests/hurl/admin.hurl — runs as-is under `hurl --variables-file …`
# @proef admin.search
GET {{base}}/api/v1/admin/search/{{index}}
Authorization: Bearer {{apiToken}}
[Query]
q: {{q}}
HTTP 200
# the pack supplies the fragment's variables; nothing is implicit
bind:
  base:     ${url:base}          # pack scope — every macro in the file
  apiToken: ${secret:apiToken}   # injected at run time, never into an artifact
macros:
  search:
    params: [q, index]
    defaults: { index: records }
    match: "the operator searches for {q}"
    bind: { q: "${q}", index: "${index}" }   # macro scope
    steps:
      - ref: admin.search

Relative paths resolve against this file’s directory, not the working directory, so the root may sit outside the suite — even outside the repo — and the config still works from any subdirectory. There is no convention fallback: with the key unset a ref: reports proef::pack::unknown_ref and says no fragment files were loaded, which beats guessing at a directory. Entries with no # @proef annotation are inert, and the files are not scanned at all until some pack names a fragment — so pointing at a corpus you did not write neither changes what runs nor costs anything to parse. When a pack does name one, the corpus is scanned once per command, not once per pack load: a run with [run] setup/teardown loads packs four times against the same files, and on a 200-file corpus rescanning each time was most of the run.

proef fmt never touches these files: it normalizes hurl blocks inside packs by locating them in YAML, and a corpus you do not own is not proef’s to rewrite.

The read is bounded. A single corpus file over 8 MiB is skipped with proef::pack::oversized_fragment_file, and the reader stops once the corpus as a whole passes 64 MiB (proef::pack::fragment_corpus_too_large). Both name the file, both leave every other file loading, and a ref: into a skipped file then reports unknown_ref beneath them.

These are not tuning knobs — they are the line past which the input is not a test corpus. A hurl file is human-authored text, so 8 MiB is on the order of 200,000 lines; the largest corpus anyone has reported is 15 files and 112 fragments. The bound exists because the root is one config line: pointed one directory too high it reads whatever is underneath, and a 279 MB file cost 601 MB of resident memory on proef flows — a command that never looks at a fragment. If you hit either limit, the fix is almost always to narrow [run] fragments to the directory that actually holds your hurl files.

Suite setup & teardown ([run] setup / [run] teardown)

[run] setup and [run] teardown each name a feature file run once around the whole suite (the model Playwright/Jest globalSetup use, ADR-0014). setup runs before the parallel pool and merges its saveAs: global promotions into the shared store before any scenario lowers, so it is the place to seed a fixture or provision shared state that every scenario then reads via ${global:…}. teardown runs once after the pool for cleanup.

[run]
setup    = "tests/setup.feature"
teardown = "tests/teardown.feature"

Failure semantics are deliberate: a setup failure aborts the run before the pool and maps to a user (2) or system (3) fault — a broken fixture is not a failing test, so it never becomes exit 1. Teardown runs only if setup succeeded (it still runs after ordinary scenario failures, for reliable cleanup), and a teardown failure is a distinct exit 3 cleanup fault — the suite’s own verdict stands, but the failure is never silently swallowed. Both features are excluded from the pool, so a setup/teardown feature inside the suite directory never also runs as an ordinary scenario. Auth is already covered by pre-set secrets, so setup is for seeding/provisioning, not obtaining a runtime token.

Run metadata ([meta], [env.<name>.meta], --meta)

Key/value pairs recorded in the run head and shown by the HTML report, the GitHub summary, explain and diff — never interpreted, never harvested. Static facts live in config; invocation facts ride the flag; the active env profile’s table deep-merges over the base, and flags win over both — the same base < env < flags chain as every other setting:

[meta]
team = "payments"

[env.staging.meta]
tier = "staging"
proef test --meta commit=$(git rev-parse HEAD) --meta build=$CI_JOB_URL

proef itself never reads git, the hostname, or CI variables — if you want the SHA recorded, your shell harvests it and proef records it (ADR-0020). The active --env profile name is recorded automatically as its own field: it is user-chosen input, and without it two records of one suite against different [env.<name>] merges read as regressions in diff.

Tag glob → URL template; a matching tag becomes a link in the HTML report’s by-tag table and the GitHub summary ({tag} is substituted). The pattern language is the same anchored glob --tags atoms use — one matcher for selection, exclusivity and links:

[tag-links]
"JIRA-*" = "https://jira.example/browse/{tag}"

Everything else ignores the table; a tag matching no pattern stays plain.

Scenarios that cannot share the pool ([run] exclusive-tags)

Some scenarios cannot run beside anything: one asserting absolute positions (items[0]) needs a store no concurrent scenario writes to. exclusive-tags is a tag expression — the same language --tags takes — selecting those. Atoms may glob (* any run, ? one character, anchored — @serial-* is the whole family; case-sensitive, unlike Robot Framework):

[run]
jobs = 8
exclusive-tags = "@serial"
  @serial
  Scenario: The report lists every record in order

A matching scenario waits for the pool to drain, runs alone, and the pool refills after it. Ordering is unchanged: scenarios still start in discovery order, so an exclusive one never loses its place — and that is also what the cost is made of. Queueing is strict FIFO, so while an exclusive scenario sits at the head no new scenario starts: the pool drains to empty, the exclusive one runs by itself, and only then does parallelism return to jobs width. Expect a throughput dip around each exclusive scenario roughly as long as the slowest scenario already running, plus the exclusive one’s own duration. Put another way, the cost is bounded and predictable rather than free.

Exclusivity is enforced against the dispatcher’s own bookkeeping. A scenario the watchdog abandons for blowing its batch budget (ADR-0007) leaves the active set immediately, but its detached thread keeps going until the request in flight returns — hurl cannot be cancelled mid-entry. In that window an exclusive scenario can start while an abandoned neighbour is still issuing requests. It takes a budget blowout in the same moment to happen, and it is the one caveat the isolation guarantee carries; a run that never trips the watchdog never meets it.

This is exclusion, not ordering. A scenario that must run before the others — installing a fixture the rest depend on — belongs in [run] setup, which already runs once before the pool exists.

A tag expression in config, not a reserved tag name, deliberately: with a bare convention, a scenario added months later lands untagged in the parallel pool and breaks isolation intermittently — which reads as flakiness rather than as a missing declaration. The expression keeps the rule in one reviewable place. A malformed one is a user error (exit 2), never a silently-ignored key.

Editor support (JSON Schema)

TOML language servers validate against JSON Schema, so one file buys completion, hover documentation and typo detection in proef.toml itself:

proef schema config > proef-config.schema.json

Then point your editor at it. Taplo and tombi (the engines behind VS Code’s Even Better TOML, and the usual Neovim/Helix TOML setups) both read a first-line directive:

#:schema ./proef-config.schema.json

[run]
suite = "tests"

The schema is generated from the same Rust model that parses the file, so it describes the keys as they are written — kebab-case (runs-dir, keep-runs, exclusive-tags), not the struct spelling — and inherits deny_unknown_fields, so an editor flags a typo’d key before you run anything. Each key’s hover text is its documentation from this page’s source.

Regenerate it after upgrading proef; there is no --add-to here (that installs the pack schema beside packs), because a project has one proef.toml and the #:schema line is a one-time edit to a file proef otherwise only ever reads.

What does not live here

  • Secrets — proef secret set / PROEF_SECRET_<NAME> env, on their own encrypted channel (${secret:…}); never in proef.toml (see AUTHORING — Secrets).
  • Per-machine values — proef.toml is committed and shared; anything that differs between teammates belongs in an environment variable (${env:NAME:-default}) or an environment profile you don’t commit.

Feature files reference ${url:…} / ${vars:…} / ${secret:…} and declare none of them — variables have exactly one home, this file, and test files stay pure prose. proef discovers the nearest proef.toml by searching up from the working directory (like cargo/git), so it is found from any subdirectory.

The search only goes up, so a discovered file must sit at or above the directory you run from — that is a requirement, not a convention. Putting proef.toml beside the suite (tests/proef/proef.toml) and running from the repository root does not work by discovery: nothing searches downward, so the config is simply not found and every ${url:…} reads as unset.

--config <path> names the file instead and removes the constraint entirely:

$ proef test --dry-run --config tests/proef/proef.toml

It is global — every subcommand accepts it, including proef lsp and proef test --watch. The editor loading a different config than the runner is the drift that makes diagnostics untrustworthy, and --watch watches the file the run actually resolved through, so editing it retriggers.

For proef lsp the flag also outranks the workspace root the client announces: the flag names a file, and a named file is not a guess to be improved on.

A named file that does not exist is a user error (exit 2) for the runner, not a fall back to defaults: discovery finding nothing means “no project here”, but a named path that is not there is a typo, and answering it with a silently unconfigured run is the thing worth refusing. proef lsp makes the opposite call and starts anyway — an editor offering less is better than one that will not boot, and unlike a run it produces no results to be wrong.

It applies to every subcommand in the same way, including the ones that read nothing from the file: proef fmt --config missing.toml exits 2 rather than formatting and reporting success. The flag names a file, and a named file that is not there is a typo whatever the command was going to do with it.

Which directory a path is relative to

Paths written in proef.toml resolve against the directory holding proef.toml. Paths typed on the command line resolve against the working directory. That is the whole rule, and it has no exceptions: suite, fragments, setup, teardown and runs-dir all follow it, as do the two files proef keeps beside them — .proef-state.json (the persistent World) and .proef-secrets.json (the secret store). An absolute value is taken as written, so a corpus or a record store may sit outside the project entirely.

The consequence worth stating plainly: a project is where its proef.toml is, not where your shell is. Running from a subdirectory finds the same config by walking up, and now resolves the same suite, writes the same run records, and reads the same World and the same secrets as running from the root. It is the convention Cargo (path/members are manifest-relative), tsconfig.json and pytest’s rootdir all follow.

With no proef.toml in scope there is nothing to be relative to, so paths stay relative to the working directory, unchanged.

A config that does not sit at the project root spells its paths from where it sits — a proef.toml in tests/proef/ beside a suite in tests/features/ writes suite = "../features".

How a path is spelled in what proef writes

The same rule, read backwards. Nothing proef records names the machine it ran on: a path that reaches an artifact, a sidecar, an event, a report or a diagnostic is spelled relative to the directory holding proef.toml.

That matters because ADR-0010 makes the emitted .hurl a contract — the same inputs must give the same bytes — and a run record is meant to travel from a laptop to CI and back. So all four of these produce the same record:

$ proef test                              # path derived from [run] suite
$ proef test suite                        # typed
$ proef test /home/you/proj/suite         # typed absolutely
$ cd sub && proef test                    # from a subdirectory

Two cases keep an absolute path, because no project-relative spelling of them exists: a suite or fragment corpus that genuinely sits outside the project (fragments = "/opt/shared-corpus"), and a run with no proef.toml in scope. A path typed relative is recorded exactly as typed — it is machine-independent already, and it is the spelling your terminal can open.

How long run records live

Each run writes one directory under runs-dir holding its artifacts, its events.jsonl and its run.log. [run] keep-runs bounds how many are kept: the default of 200 suits an archive, and a suite re-run on every save wants far fewer. The artifacts are the largest part and are byte-identical between runs of an unchanged suite — but they are what that run executed, and once the corpus moves on proef artifacts no longer reproduces them, so the cost is bounded rather than dropped.

Rotation only ever deletes directories named by a generated run id, because runs-dir may be . or otherwise shared with your own files. A record written under --run-id <name> is therefore never rotated: if CI mints a fresh name per build into a persistent directory, prune it yourself.

A starter file ships as proef.toml.example in the repository root.

Diagnostic codes — the greppable index

Every proef diagnostic carries a stable code (proef::<area>::<name>), printed above the rendered message. Codes are a contract: they never change meaning, and searching this file (or tests/errors/) for a code you hit is the fastest route to the cause. The corpus column marks codes with a seeded broken example under tests/errors/<area>__<name>/ — dry-running one shows the exact rendered output.

Severity is error (fails validation/exit 2) unless marked warning.

proef::feature::* — feature-file parsing and expansion

CodeMeaningCorpus
feature::parseThe file is not valid Gherkin (parser error, located)✓
feature::empty_fileThe file is empty — a Feature: header is required
feature::no_examplesA Scenario Outline has no Examples rows
feature::ragged_examplesAn Examples row’s cell count differs from the header
feature::bad_examples_headerDuplicate or empty Examples column name
feature::unknown_placeholderAn outline step uses <name> not present in the header✓
feature::empty_scenarioA scenario has no steps (background included)✓

proef::pack::* — macro-pack loading and validation

CodeMeaningCorpus
pack::yamlThe pack file is not valid YAML✓
pack::duplicate_macroTwo macros share a name✓
pack::empty_macroA macro has neither steps: nor expect:
pack::steps_and_expectA macro has both steps: and expect:
pack::empty_expectAn expect: item asserts nothing✓
pack::empty_stepA step has no payload and no use:
pack::multiple_payloadsA step has more than one payload key
pack::unknown_step_kindThe payload key names no registered engine✓
pack::invalid_hurlThe payload does not parse as hurl (incl. zero-entry payloads)✓
pack::payload_invalidA structured payload fails the engine’s validator✓
pack::bad_referenceA ${…} reference in the payload cannot resolve at probe time
pack::retry_not_finiteretry/repeat is -1, 0, or above the 10000 cap✓
pack::delay_unboundeddelay exceeds the 1-hour cap
pack::option_declared_twiceretry/delay set both in the block’s [Options] and as the step’s own key, or a variable both supplied by a fragment’s [Options] variable: and given by a bind:✓
pack::pattern_bracesUnbalanced {/} in a match: pattern
pack::pattern_empty_capture{} with no capture name
pack::pattern_no_anchorA pattern with no literal word (capture-only)✓
pack::pattern_unknown_captureA {capture} that is not a declared param✓
pack::pattern_duplicate_captureThe same {capture} written twice in one pattern✓
pack::adjacent_captures{a} {b} with no literal between captures✓
pack::default_not_paramA defaults: key that is not a declared param✓
pack::bad_save_targetsaveAs: target other than global
pack::unknown_refref: names no loaded fragment (suggests the closest)✓
pack::duplicate_fragmentTwo fragment files declare the same # @proef name
pack::bad_annotationA fragment file the engine’s parser could not read, an annotation it could not attach, or a name containing # (which could never be referenced)
pack::unreadable_fragment_fileA fragment file that could not be read at all (encoding, permissions) — its siblings still load, and it is silent until something ref:s the corpus
pack::oversized_fragment_fileA fragment file over 8 MiB — skipped unread (the size comes from the directory entry), siblings still load
pack::fragment_corpus_too_largeThe corpus as a whole passed 64 MiB — the read stops, naming the file it stopped at
pack::body_form_conflictA step is both ref: and a payload (or use:)✓
pack::bind_without_refbind: with no ref: to read it — on a step (an inline block takes ${…} instead), or on a macro whose steps have none (a use: target resolves its own)✓
pack::unread_bind_keya bind: key no fragment in that scope reads (did-you-mean over the readable names) — the finer half of bind_without_ref, and the one a typo produces
pack::unknown_useuse: names no known macro✓
pack::use_cycleuse: composition forms a cycle✓
pack::use_too_deepuse: nesting exceeds depth 32
pack::use_with_modifiersA use: step carries modifiers that belong on the target
pack::use_with_payloadA step is both use: and a payload
pack::with_without_usewith: on a step that has no use:
pack::missing_use_paramThe use: target requires a param with: does not supply✓
pack::unknown_with_keyA with: key the target macro does not declare✓

proef::bind::* — matching prose to macros

CodeMeaningCorpus
bind::unbound_stepNo macro pattern matches the sentence (suggests the closest, plus a paste-ready macro stub)✓
bind::ambiguous_stepMore than one macro matches (all candidates listed)✓
bind::missing_paramA required param has no capture, table value, or default✓
bind::bad_tableA data table has an unusable shape
bind::unknown_table_keyA table key that is not a declared param
bind::table_conflictA table value collides with a pattern capture✓
bind::docstring_unusedA docstring the macro never references (warning)

proef::lower::* — lowering to engine batches

CodeMeaningCorpus
lower::then_before_whenAn expect: step with no previous request to attach to✓
lower::bad_statusexpect: status: is not an HTTP status number
lower::kind_unroutedInternal safety net: a lowered step’s kind maps to no engine (registry drift — unreachable through the CLI, which is why it has no corpus case)
lower::expansion_too_deepMacro expansion exceeded depth 32 at run time
lower::unbound_placeholderA {{variable}} nothing supplies — read by the fragment, or inside a bind: value (as the engine’s parser reads it: a hurl function like {{newUuid}} is not a variable) — with no bind: in scope, no earlier capture, no fragment-own [Options] variable:, no earlier-sorting sibling literal, and no secret of that name
lower::bind_shadows_captureA literal bind: re-assigns a name an earlier step captured, so the bound value silently wins from that entry on (warning)
lower::multiline_bindA bind: value that resolves to a line break or control character (tab excepted) — a hurl [Options] variable: is a single-line scalar; the inline hurl: | form is what splices a multi-line body
lower::secret_in_composite_bindA bind: value mixes ${secret:…} into a larger string — bind the secret alone and put the surrounding text in the fragment
lower::dry_run_unknownA runtime-only global under --dry-run (warning)

proef::resolve::* — ${…} variable resolution

CodeMeaningCorpus
resolve::unknown_variable${name} found in no scope (suggests the closest)✓
resolve::missing_env${env:NAME} unset and no :-default given
resolve::missing_config_var${url:key} / ${vars:key} defined in neither proef.toml nor the active [env.<name>] (suggests the closest)✓
resolve::missing_global${global:key} absent from the World (strict mode)
resolve::unknown_namespace${ns:…} with an unknown namespace
resolve::unknown_run_field${run:…} other than ${run:id}
resolve::fake_unknown${fake:kind} names no generator (suggests the closest)
resolve::empty_referenceAn empty ${}
resolve::depth_exceededResolution passed depth 8 — a reference cycle

proef::emit::* — artifact emission

CodeMeaningCorpus
emit::invalid_artifactThe emitted artifact does not parse with the engine’s parser

proef::run::* — preparing a scenario to execute

CodeMeaningCorpus
run::asset_unstageableA file,…; asset could not be staged into the scenario’s asset root: absent beside the source that names it, named by a path proef will not follow, or claimed by two sources at once

proef::config::* — proef.toml loading

CodeMeaningCorpus
config::tomlThe file is not valid TOML for the config schema (located — the caret sits on toml’s own error span)
config::unreadableThe file (or an explicit --config path) cannot be read

proef::source::* — source access (LSP whole-suite analysis)

CodeMeaningCorpus
source::unreadableA discovered feature or pack source could not be read (surfaced by analyze_suite; the CLI treats an unreadable file as a system fault instead)

proef::tags::* — reserved-tag recognition

CodeMeaningCorpus
tags::reserved_tag_typoA tag looks like a reserved one (@quarantine, @skip) but is not exactly it, so it has no effect — a warning with the spelling it likely meant

Coverage note

The fragment-file codes (pack::duplicate_fragment, pack::bad_annotation, pack::unreadable_fragment_file, pack::unread_bind_key, lower::unbound_placeholder, lower::multiline_bind, lower::secret_in_composite_bind, run::asset_unstageable) are covered in crates/proef-cli/tests/fragments.rs rather than tests/errors/: they need a [run] fragments root, and the seeded corpus is deliberately config-independent.

The config::* codes are covered in crates/proef-cli/tests/cli.rs for the same reason: a broken proef.toml cannot live in the config-independent seeded corpus.

30 of the 77 codes carry a seeded corpus case today; the corpus guard asserts a minimum, not parity. When you add a diagnostic, add its code here and prefer seeding a tests/errors/<area>__<name>/ case alongside it. tags::reserved_tag_typo is a warning (a near-miss is not an error), so it cannot live in tests/errors/ — that corpus fails dry-run by design — and is covered by a unit test instead.

Every row here is a code some code path actually emits. That was not free: this file used to carry a pack::load row for a defensive case that never had a code (a non-diagnostic core failure while loading packs renders as a bare error:), so a reader who grepped for it found nothing and had no way to tell the index was wrong. A row nothing emits is worse than a missing row.

Event schema — the run record for machine consumers

.proef-runs/<run-id>/events.jsonl is the run record (ADR-0008): one JSON object per line, in emission order, no second format. proef explain, diff, flaky and the console tree derive from this stream; --junit and the GitHub summary are built from the same run’s in-memory outcome (identical content, different plumbing — the stream stays the only persisted format; the one exception is a --rerun’s JUnit, whose carried base scenarios are reconstructed from the base record — see rerun_of below). Consume it with jq, a log shipper, or anything line-oriented.

Two derived sidecars sit beside the record and are not records: timings.json (per-scenario durations, re-derivable from the stream, read back by --shard-weights) and inputs.json (the run’s input fingerprint — a hash of proef’s own inputs, proef flaky’s equivalence class, ADR-0020 amendment; absent on pre-0.18 records). Neither adds a variant or a field to the stream.

Wire shape: serde-tagged with "event", snake_case names. The first line is always run_started and declares schema (currently 1); the last is run_finished. The schema is additive-only: new variants and new optional fields may appear, existing fields never change meaning or vanish. Consumers must ignore unknown variants and unknown fields.

Truncated records. A record whose last line is not run_finished is a run that did not finish — a hard interrupt (second signal, exit 130), a SIGKILL, or a crash. (A single SIGTERM/SIGHUP is a graceful cancellation: the record closes normally with run_finished + cancelled.) Treat it as partial: the scenarios present did happen, but the totals were never written and no verdict was reached. explain, report and diff classify these and say “run incomplete” rather than reporting a total they cannot know. Consumers should do the same rather than inferring zero.

One exception to “fields never change meaning”, recorded rather than hidden. run_finished’s passed/failed/skipped counted every phase before 0.6.0; from 0.6.0 they are the main-suite verdict and exclude [run] setup/teardown (ADR-0014), so those totals agree with the exit code, --format json/TAP, and the summary line. One deliberate divergence: a failed phase additionally appears in JUnit as its own suite (a gated pipeline must see the failure), so JUnit’s per-case counts can exceed these totals on a run whose setup or teardown broke. A pre-0.6.0 record read by a current consumer reports the older meaning.

Variants

run_started — head of every stream. schema (u32) · run_id (string — uuid-v7 by default, but --run-id passes any name through verbatim, so consumers must treat it as opaque).

scenario_started — scenario (string) · file (string) · timestamp_ms (u64, run-relative ms — only present with injected timing) · worker (u64, 0-based worker index — only present with injected timing). The timing pair is stamped at the CLI sink on the worker thread (the sans-IO core leaves it absent), and powers the HTML report timeline (ADR-0015).

batch_started — a contiguous same-engine step batch was dispatched. scenario · engine (e.g. "hurl") · steps (count).

entry_running — live progress: one event per execution attempt of an artifact entry, retries included. scenario · engine · entry (0-based ordinal within the scenario’s artifact) · retry (0 = first attempt).

step_finished — scenario · engine · step ({file, line, text}, the authored feature anchor) · status (passed | failed | skipped | warned) · attempts (u32) · duration_ms (u64) · captures (capture names only — never values) · fragment (file.hurl#name of the named fragment the step ran, ADR-0018 — only present for a ref: step, so an inline hurl: block and every pre-fragment record omit the key entirely) · label (the pack step’s authored name:, resolved — only present when the step has one. One feature sentence commonly lowers to several engine steps, and they share a step anchor exactly; the label is what tells them apart, and it is the same text the emitter writes into the artifact’s entry comment) · detail (string, only present on failures/warnings/skips-with-reason) · attempt_details (array of strings — the messages from earlier, failed attempts of a step that ultimately passed; only present for a flaky pass, feeds JUnit <flakyFailure>) · reproduce_hint (string — the redacted curl of the failing request, the same line the console prints as reproduce:; only present on failed steps, absent in every pre-field stream, masked at the sink boundary like detail).

scenario_finished — scenario · file (feature path — with scenario, the run-wide identity: names are unique only within one file; absent in records that predate the field) · status · timestamp_ms (u64, run-relative end ms — only present with injected timing) · worker (u64, in the schema but never populated by proef’s own writer today — it is emitted from the main dispatcher thread, not the scenario’s worker, so consumers must accept the field without expecting it — ADR-0015).

env / metadata / shuffled — on run_started, all additive and absent when unset. env is the active --env/PROEF_ENV profile name; metadata is the explicit user-supplied map (--meta k=v, [meta], [env.<name>.meta] — proef never harvests: no git, no hostname, no CI env sniffing; ADR-0020); shuffled says the order was re-dealt (the permutation is seeded by run_id, so the pair reproduces it). Keys and values pass the sink-boundary mask like every text field.

rerun_of — on run_started, optional (additive; absent = not a rerun). The base record’s run_id a --rerun re-ran failures from. The rerun’s JUnit carries the base’s not-re-run suite scenarios as ordinary testcases, and report overlays the base for a whole-suite page — composition over records, the record files themselves never merge (ADR-0008). Totals and the exit code stay the rerun’s own (ADR-0014).

tags — on scenario_finished, optional (additive; absent when the scenario carries none and in every pre-field stream). The accumulated tags (feature → rule → scenario → examples, deduped, authored order, @ stripped). On the finished event only: the cancel-skip path emits no scenario_started, and per-tag skip counts need every scenario.

exclusive — on scenario_started, optional bool (absent = false, which is what every pre-field record meant). The scenario ran with the pool to itself ([run] exclusive-tags) — the bool the scheduler itself read, recorded for timeline post-mortems (R11-6).

reason — on scenario_finished, optional (additive; absent in pre-field records and on every non-skipped scenario). Why the scenario is skipped: an authored skip carries the pasteable tag spelling (@skip / @skip:reason), which always begins with @; a mechanical skip (cancellation) carries proef-fixed prose, which never does. --rerun keys its re-queue decision on exactly that split (ADR-0019).

phase — on scenario_started/scenario_finished, optional. "setup" or "teardown" when the scenario belongs to a [run] lifecycle phase, absent for an ordinary suite scenario. Consumers should use it rather than inferring phase membership from the feature path: run_finished’s totals exclude phases (ADR-0014), so a consumer that counts every scenario_finished will not match them. Added additively — records without it have no phases, which is what they had.

run_finished — tail. passed · failed · skipped — the main-suite verdict: scenario counts for the primary suite only (ADR-0014). [run] setup/teardown scenarios still appear as their own scenario_started/ scenario_finished events earlier in the stream, but are excluded from these totals, so they agree with the console summary: line, proef explain, proef report’s HTML headline, --format json, TAP, the SLA gate, and the exit code — JUnit agrees too except that a failed phase rides in as its own suite (see above) · cancelled (bool, only present when true).

Example stream

{"event":"run_started","schema":1,"run_id":"019f…"}
{"event":"scenario_started","scenario":"reference","file":"suite/case.feature","timestamp_ms":0,"worker":0}
{"event":"batch_started","scenario":"reference","engine":"hurl","steps":2}
{"event":"entry_running","scenario":"reference","engine":"hurl","entry":0,"retry":0}
{"event":"step_finished","scenario":"reference","engine":"hurl","step":{"file":"suite/case.feature","line":4,"text":"the cookie session is exercised"},"status":"passed","attempts":1,"duration_ms":12,"captures":[]}
{"event":"step_finished","scenario":"reference","engine":"hurl","step":{"file":"suite/case.feature","line":5,"text":"the response status is 200"},"status":"passed","attempts":1,"duration_ms":0,"captures":[]}
{"event":"scenario_finished","scenario":"reference","file":"suite/case.feature","status":"passed","timestamp_ms":12}
{"event":"run_finished","passed":1,"failed":0,"skipped":0}

Guarantees

  • Secrets never appear. Redaction applies once at the sink boundary before any reporter (property-tested); captures carries names only.
  • Order: events of one scenario are internally ordered; scenarios running in parallel interleave. Group by (scenario, file) before assuming sequence across the stream — scenario alone collides when two feature files reuse a name.
  • Flake-safe assertions: assert on attempts counts and normalized event order, never wall-clock (duration_ms is engine-measured and varies).

Recipes

jq -r 'select(.event=="step_finished" and .status=="failed") | "\(.step.file):\(.step.line) \(.detail)"' events.jsonl
jq -r 'select(.event=="run_finished")' events.jsonl          # the suite verdict (setup/teardown excluded)
jq -r 'select(.event=="entry_running") | .retry' events.jsonl | sort | uniq -c   # retry pressure
# which fragment files a failing run actually exercised (ADR-0018); `// empty`
# drops the inline steps, which carry no `fragment` key at all
jq -r 'select(.event=="step_finished") | .fragment // empty' events.jsonl | sort | uniq -c

proef — Product Requirements Document

Status: approved; US-1…US-12 all in service (multi-engine architectural only), plus post-M5 work: external config/environments (ADR-0012), the v0.6–v0.8 correctness series, v0.9.0, named hurl fragments (ADR-0018 — see the §3 amendment), the adoption response, the 0.11–0.14 hardening and CI-scale series, the RF capability audit and reserved tags (ADR-0019/0020), and the deep-improvement waves plus round 19. CLAUDE.md’s Status block is the running ledger; the goals and non-goals below are what has not changed. · Date: 2026-07-28, status refreshed 2026-08-31 · Owner: Emre Companion docs: ADRs for the why, TECH-SPEC for the how, IMPLEMENTATION-PLAN for the when.

1. Problem & context

End-to-end tests written as Gherkin business prose — bound to a compiled-in vocabulary of YAML macro packs — let non-programmers author real tests. Meanwhile the backend team maintains a corpus of hand-written Hurl files as its API-testing lingua franca. There is no tool that joins the two: Gherkin prose on top, Hurl-grade API execution underneath — and none whose engine seam can later admit further engines behind the same prose (architectural readiness only).

proef is that tool: a declarative, modular, multi-engine e2e runner (mostly Rust). One core parses/binds/lowers Gherkin; pluggable engines execute. The API engine embeds hurl itself (ADR-0001), so API semantics are hurl’s by construction, and every scenario also produces real .hurl artifacts the backend team can run and read (ADR-0010).

2. Goals

G1. Author API e2e tests as pure Gherkin prose in the 500-series style — no code, no URLs in prose, data tables and Scenario Outlines supported. G2. Execute with hurl’s exact semantics (asserts, captures, retries, templating) via the embedded engine; identical results to the stock hurl CLI on the same artifacts. G3. Emit canonical .hurl artifacts + sidecar maps for every scenario — debuggable with the tools the backend team already uses, replayable via hurl --variables-file. G4. Be modular: adding an engine is a new crate + a registry line, with zero changes to proef-core (the structural acceptance test, ADR-0002). G5. Track upstream hurl safely over time: exact pins, a zero-diff fork as patch vehicle, and an upgrade-canary CI job so new hurl releases are absorbed deliberately (ADR-0003). G6. First-class failure UX: every failure maps to the .feature line (and the artifact span), rendered with miette-style labeled diagnostics; stable exit codes 0/1/2/3. G7. CI-native: JUnit XML, GitHub job summary, JSONL run records, tag filtering, parallel scenarios, --dry-run validation gate, and (M5) a libtest-mimic harness so cargo nextest run and IDEs can drive scenarios.

3. Non-goals (v1)

Further engines (designed-for, not built — M6+); a desktop dashboard or server mode (precedent exists when needed); generating Gherkin, macros, or prose from hand-written hurl files (see the amendment below); OpenTelemetry export (semconv still immature — JSONL is the source of truth); dynamic plugin loading; Windows-static or musl-static binaries (dynamically linked like hurl’s own — consequence of ADR-0001); API mocking/contract testing; load testing.

Amendment (2026-08-11): the hurl non-goal is about generation, not direction

This non-goal previously read “importing/round-tripping hand-written hurl files into Gherkin (artifacts flow outward only)”. The clause and its parenthetical said two different things, and the parenthetical was the broader of the two: read literally it forbids hurl text being an input at all, which ADR-0018 (named hurl fragments) needs it not to.

What the non-goal protects is that proef never authors a test for you. A suite’s prose and its binding vocabulary are written by people, deliberately; a tool that derives them from an existing corpus produces scenarios nobody chose the words for, and the review that makes a feature file worth having never happens. That reasoning is untouched, and ADR-0016 (OpenAPI generation) stays declined on the same ground.

It does not extend to hurl text being an input source. A macro pack is already an input written in a non-Gherkin language; a .hurl file naming reusable fragments is another, and §1’s own framing — “the backend team maintains a corpus of hand-written Hurl files as its API-testing lingua franca. There is no tool that joins the two” — describes joining that corpus as the product’s purpose, not as a boundary. Nothing is generated: features and macros stay hand-authored, and a fragment is inert until a macro names it. The worklist reached the same reading independently for a different item (OPEN-FINDINGS M2: “the non-goal forecloses a direction of data flow, not the ability to check your own work”).

Recorded honestly: M3 asked that this charter be re-examined with a measured port cost, and that measurement does not exist. The narrowing is therefore argued from the non-goal’s own rationale rather than from evidence that pasting is expensive. If the measurement later shows pasting is cheap, that argues about priority — it does not restore a prohibition this amendment finds was never the point.

4. Users & personas

P1 — Test author (QA, PM, support engineer; not necessarily a programmer). Writes .feature files against the existing step vocabulary. Cares about: prose that reads naturally, --dry-run telling them exactly which step is wrong, failures pointing at their line, not internals. P2 — Pack maintainer (developer). Owns the macro packs: adds steps, params, asserts. Cares about: pack readability (raw hurl blocks, ADR-0004), schema autocomplete, load-time validation with did-you-mean hints, safe refactors (duplicate/cycle detection). P3 — Backend engineer (owns the hurl corpus). Consumes emitted artifacts; pastes between corpus and packs, or — since ADR-0018 — annotates a corpus file once and lets packs name its entries, so the same file stays runnable under stock hurl. Cares about: artifacts being idiomatic hurl, never containing secrets, runnable standalone; and that proef reading their corpus never edits or reformats it. P4 — CI pipeline (machine). Cares about: stable exit codes, JUnit/JSONL outputs, deterministic behavior, bounded runtime (cancellation/budgets, ADR-0007), the canary job.

5. User stories & acceptance criteria

US-1 (P1) As a test author I write a scenario in prose and run it against a live environment. AC: the four 500-series .feature files run via proef packs with prose unchanged except agreed wording fixes; proef test runs them green against the fixture; failures name feature file + line + step text. US-2 (P1) I validate without executing. AC: proef test --dry-run binds every step, expands every macro/outline, resolves ${…}, parses every generated artifact with hurl’s parser, checks every file,…; asset those artifacts read is present, and exits 2 with labeled diagnostics on any failure — no network I/O. US-3 (P1) I pass data per step. AC: inline {captures}, | key | value | data tables, and Scenario Outline <placeholders> all fill macro params; conflicts and missing required params are parse-time errors naming the line. US-4 (P1) Steps chain state. AC: a capture in one step (clientId) is usable in later steps of the scenario as {{clientId}}; saveAs: global persists across scenarios and runs (World, ADR-0005). US-5 (P1) Slow backends don’t flake. AC: a step with retry: polls until its asserts pass or the finite budget ends (maps to hurl [Options] retry); optional: steps warn instead of failing. US-6 (P2) I extend the vocabulary. AC: adding a macro with a match: pattern + raw hurl block to a pack makes the new prose step available; proef schema reflects it; load rejects ambiguous names, adjacent captures, literal-free patterns, infinite retries. US-7 (P3) I get artifacts. AC: proef artifacts writes per-scenario .hurl + .map.json (+ .vars when Worlds are referenced); stock hurl --test yields the same verdicts (spike-proven); no secret values appear in any artifact. US-8 (P4) CI consumes results. AC: exit codes 0/1/2/3 stable and integration-tested; --junit auto under GITHUB_ACTIONS; JSONL event log written per run; --tags filters. US-9 (P1/P4) Runs are observable. AC: console BDD tree with per-step timing/attempts; proef explain [run] summarizes the latest failures from run records; proef diff [base] [new] compares two runs for regressions, fixes, flakiness, and perf deltas; proef report [run] writes a self-contained HTML report of a run. US-10 (P2) Secrets stay secret. AC: proef secret set stores encrypted values; secret values never appear in artifacts, logs, reports, or events (property-tested). US-11 (P4) hurl upgrades are safe. AC: the canary job builds against the next hurl release and replays the suite; pins move only after it is green (runbook in IMPLEMENTATION-PLAN §7). US-12 (P1, M5) IDE/nextest integration. AC: the libtest-mimic harness lists one test per scenario and cargo nextest run executes and reports them. US-13 (P1) I can start from something that works. AC: proef init writes a minimal suite (proef.toml, one .feature, one matching pack) that passes --dry-run unchanged, installs the pack JSON Schema for editor completion, and never overwrites an existing file.

6. Functional requirements (condensed; TECH-SPEC is normative)

Authoring: full gherkin-crate grammar (Feature/Rule/Background/Scenario/Outline/ Examples/tables/docstrings/tags/i18n); tags filter runs; variables come from proef.toml (${url:}/${vars:}, ADR-0012), never the feature files. Packs: YAML skeleton with match: patterns ({name} captures), params/defaults/tags/description; step bodies as raw hurl: blocks (primary) or structured form (reserved for future engines); assert-only expect: macros merge into the previous request (Then-steps); use:/with: nesting with cycle/depth limits; schemars-derived JSON Schema; lint pass at load. Execution: scenario = unit of isolation and parallelism (--jobs); contiguous same-engine batching; engine-hurl via embedded run_entries with buffered I/O; per-entry [Options] override batch defaults (verified); World seeding/merge-back; cooperative cancellation + budgets. Artifacts: canonical emit, sidecars, vars files, # optional markers. Reporting: event spine → console/JUnit/JSONL/GitHub-summary reporters; run-record rotation. CLI: test ([path] --env --dry-run --tags --jobs --junit --format json|tap --watch --run-id --sarif --rerun --scenario[-file]; path optional — [run] suite then the tests/ convention), flows, macros (call counts + dead-macro report), artifacts, schema [--add-to], secret set|list, explain, diff [base] [new] --fail-on-regression, report [run] -o <file>, doctor. Config (proef.toml, ADR-0012): runner settings ([run] incl. setup/teardown suite lifecycle features, ADR-0014; [http]/[sla]) + suite variables ([url]/[vars], referenced ${url:key}/${vars:key}) + per-environment overrides ([env.<name>]); precedence defaults < base tables < active [env.<name>] (via --env/PROEF_ENV) < flags. Secrets via PROEF_SECRET_<NAME> env override → the encrypted store (proef secret set — hidden prompt, or --stdin for scripts), never in proef.toml.

7. Non-functional requirements

Correctness/fidelity: API semantics are hurl’s own binary-identical engine — no reimplementation drift is possible (ADR-0001/0010). Portability: Linux (glibc), macOS, Windows; documented build prereqs (libcurl/libxml2/libclang), doctor-checked; dynamically-linked dist binaries (hurl’s own model). Performance: startup overhead (parse+bind+lower for a 30-scenario suite) under ~1 s; execution dominated by the network; scenario-level parallelism with worker threads. Reliability: deterministic lowering (injected clock/run-id, sans-IO-lite core); bounded runtime under cancellation budgets; state writes atomic (temp+rename). Security: secret redaction invariants property-tested; artifacts guaranteed secret-free; encrypted-at-rest secret store; 0600 on sensitive outputs. Maintainability: exact pins + canary + thin-fork policy; CI gates: fmt, clippy -D warnings, nextest, doc -D warnings, deny, machete, zizmor, docs-check, public-api (audit nightly). Extensibility: new engine = new crate implementing EngineFactory/EngineSession + one registry line; step-kind schema contributed by the engine; zero core changes (the M6 acceptance test).

8. Success metrics

M-1: the four 50x API features run under proef with prose unchanged (US-1) — the port-fidelity bar. M-2: artifact parity — stock hurl CLI verdicts match proef verdicts on every artifact in the integration suite (already spike-proven; kept as a CI invariant until M4, then by construction). M-3: --dry-run catches 100% of the seeded pack/feature error corpus with line-accurate diagnostics. M-4: one hurl upstream release absorbed via the canary runbook with zero suite regressions. M-5: a future non-hurl engine lands (M6) with git diff --stat proef-core empty.

9. Release phasing

v0.1 = M0–M3 (authoring, validation, embedded execution, artifacts, console+JSONL); v0.2 = M4 (upstream tracking hardened, JUnit/GitHub reporters); v0.3 = M5 (breadth: multipart/form/docstring bodies, watch, explain, libtest-mimic harness); v1.0 = stability declaration of pack schema + CLI + event schema; M6 engines version independently. Detailed task breakdown: IMPLEMENTATION-PLAN.md.

proef — Technical Specification

Status: normative · Date: 2026-07-28, current through ADR-0020 (run metadata) · decisions referenced as ADR-XXXX. Post-M5 work — external config and environments (ADR-0012), the v0.6–v0.8 correctness series, v0.9.0, fragments, and the RF-audit series (reserved tags ADR-0019, run metadata ADR-0020) — is specified here too; “M0–M5” described this file’s scope only until those landed. Verified upstream facts cite hurl master @ 03fcb84c (2026-07-27) as file:line.

1. System overview

 .feature files          macro packs (YAML + raw hurl blocks)
      │                        │
      ▼                        ▼
 ┌─────────────────────────────────────────┐
 │ proef-core (pure: no IO, no clock, no   │
 │ rand — values injected)                 │
 │  parse ─ bind ─ lower ─┬─ emit          │
 │  (gherkin) (matcher)   │  (.hurl +      │
 │                        │   sidecars)    │
 │  dispatch: contiguous  │                │
 │  same-engine batches ──┼── events ──────┼──► reporter stack
 └────────────┬───────────┴────────────────┘    (console/JUnit/JSONL/GH)
              ▼ Box<dyn EngineSession>
   ┌──────────────────┐  ┌──────────────────┐
   │ proef-engine-hurl│  │ future non-hurl  │
   │ parse_hurl_file +│  │ engine (seam-    │
   │ run_entries      │  │ ready, none      │
   └──────────────────┘  │ scheduled)       │
                         └──────────────────┘
        World (typed vars + persistent global store) threads through every batch

Scenario = unit of isolation, parallelism, retry, and artifact emission. The orchestrator (in core, driven by cli) owns threads, the token, the World, and the event stream.

2. Workspace

proef/
  rust-toolchain.toml      # channel "1.97.1" (policy: latest stable at its x.y.1), rustfmt+clippy
  Cargo.toml               # virtual workspace, resolver="3", workspace.{package,dependencies,lints}
  deny.toml  .config/nextest.toml  proef.toml.example
  .github/workflows/ci.yml           # fmt, clippy -D warnings, nextest, doc -D warnings,
                                     # deny, audit, canary (ADR-0003)
  xtask/                   # automation as Rust (fixture, canary, docs-check, public-api); just aliases
  crates/
    proef-core/            # engine-agnostic: gherkin parse, packs, binding, lowering, IR,
      helpers/             #   emit, dispatch, World/state, events, errors, reporters
    proef-engine-hurl/     # EngineFactory/EngineSession impl over embedded hurl
    proef-cli/             # bin `proef`: clap, registry assembly, miette rendering
    proef-fixture/         # dev-only: in-process synchronous fixture API server (ADR-0011)
    proef-harness/         # libtest-mimic bridge: one Trial per scenario (US-12)
    proef-lsp/             # language server: SourceProvider + collect-all analyze_suite over core
  tests/                   # .feature corpus + fixtures
  docs/                    # this corpus

Dependency rules: engines depend on core; core depends on no engine; cli depends on both and is the only miette user (ADR-0009). Engines sit behind cargo features in cli (engine-hurl default-on; any future non-hurl engine would be added the same way — none scheduled). Only proef-engine-hurl carries native build prereqs; proef-core is pure Rust. Lints/conventions: a strict workspace lints table verbatim (clippy all=warn + curated pedantic slice), publish = false at the workspace root — overridden to true by the four crates that publish (proef, proef-core, proef-engine-hurl, proef-lsp), MIT OR Apache-2.0.

3. Core domain types (sketches; signatures normative, field lists indicative)

#![allow(unused)]
fn main() {
// world.rs — typed variable scope (ADR-0005)
pub enum Value { String(String), Number(f64|i64…), Bool(bool), Null }   // mirrors hurl Value subset
pub struct World { scenario: BTreeMap<String, Value>,
                   global: GlobalStore /* .proef-state.json, atomic temp+rename save */ }

// step.rs — lowered, engine-agnostic
pub struct StepRef { pub file: Arc<str>, pub line: usize, pub text: Arc<str> }  // feature anchor
pub struct LoweredStep { pub step: StepRef, pub kind: StepKindId, pub payload: StepPayload,
                         pub optional: bool, pub when: Option<Guard>,
                         // `file.hurl#name` for a `ref:` step, None for an inline block
                         // (ADR-0018) — qualified at lowering so a record stands alone
                         pub fragment: Option<String> }   // retry travels as baked [Options]
pub enum StepPayload { HurlEntries(String /* lowered hurl text */),
                       MergedAsserts { lines: usize /* expect: rows own the appended assert lines */ },
                       Structured(serde_json::Value) }
pub struct StepBatch { pub index: usize /* scenario-wide ordinal */, pub engine: EngineId,
                       pub steps: Vec<LoweredStep> }

// engine.rs — the seam (ADR-0002); see ADR text for EngineFactory/EngineSession
pub struct StepKindSpec { pub prefix: &'static str, pub schema: &'static str /* JSON-Schema frag */,
                          pub validate: Option<fn(&str) -> Result<(), PayloadProbeError>>,
                          // fragment files (ADR-0018): one Option, so extension and reader
                          // cannot disagree; discovery asks for the extension, never names one
                          pub fragments: Option<FragmentSupport>,
                          // raw [Options] keys → the core's budget policy, so option spellings
                          // live only in the engine that owns them (ADR-0007)
                          pub options: Option<fn(&str) -> Option<RawOption>>,
                          // which files a lowered payload sends, so the emitter records what an
                          // artifact reads without knowing the engine's body grammar (ADR-0002
                          // amendment, 2026-09-10 correction)
                          pub assets: Option<fn(&str) -> Vec<String>> }
pub struct FragmentSupport { pub ext: &'static str /* "hurl" */, pub scan: FragmentScanner,
                             // "which variables does one template value read", answered by the
                             // engine's parser — a hurl function ({{newUuid}}) is not a variable
                             pub template_reads: fn(&str) -> Vec<String> }
pub type FragmentScanner = fn(&str) -> Result<ScannedFile, FragmentScanError>;
// ScannedFile = { fragments: Vec<ScannedFragment>, unannotated: Vec<usize> } — the
// unannotated entry lines feed `proef fragments`' listing, not the pack loader
// Everything here is *read* from the entry — nothing is declared twice, so nothing can drift.
// A scanner reports ONLY annotated entries: an unannotated one is not a fragment, and a
// foreign corpus is mostly those, so building them only to be discarded is the bulk of a scan
pub struct ScannedFragment { pub name: String /* from `# @proef <name>` */, pub text: String,
                             pub line: usize, pub placeholders: Vec<String> /* reads */,
                             pub declared_options: Vec<String> /* ⊆ OPTION_FAMILIES */,
                             // `[Options] variable:` — what the entry supplies to itself.
                             // Answers its own placeholders (so the file runs standalone)
                             // and clashes name-to-name with a `bind:` of that name; kept
                             // apart from declared_options, which clashes family-to-family
                             pub supplied_variables: Vec<String> }
pub struct FragmentScanError { pub line: usize, pub column: usize, pub message: String }
// The vocabulary `declared_options` must use: matched by string equality against the pack's
// own option keys, so an engine-only spelling silences `option_declared_twice` rather than
// firing it. `MacroStep::declared_options()` derives the other half of that comparison.
pub const OPTION_FAMILIES: &[&str] = &["retry", "delay"];
// The one place `secret_bindings` (variable → secret) is joined with `secrets` (name → value).
// Yields borrows: an owned map would copy every secret value per scenario (ADR-0005)
pub fn secret_variables<'a>(bindings: &'a BTreeMap<String, String>,
                            secrets: &'a BTreeMap<String, String>)
                            -> impl Iterator<Item = (&'a str, &'a str)>;
pub struct DoctorCheck { pub name: &'static str, pub run: fn() -> DoctorResult }
pub struct BatchResult { pub steps: Vec<StepOutcome>, pub error: Option<EngineError> }
// `fragment` is carried here as well as on the event: JUnit, the job summary, the
// annotations, TAP and the console are built from RunSummary after the event stream has
// been written out, so they cannot read it back. Both are copies of one lowering-time
// source, so they cannot drift from each other.
pub struct StepOutcome { pub step: StepRef, pub status: Status, pub attempts: u32,
                         pub duration: Duration, pub detail: Option<String>,
                         pub attempt_details: Vec<String>, pub reproduce_hint: Option<String>,
                         pub fragment: Option<String> }

// events.rs — the spine (ADR-0008); serde, versioned
#[serde(tag = "event", rename_all = "snake_case")]
pub enum Event { RunStarted{..}, ScenarioStarted{..}, BatchStarted{..}, EntryRunning{..},
                 StepFinished{..}, ScenarioFinished{..}, RunFinished{.., cancelled} }
pub struct EventSink(Arc<dyn Fn(&Event) + Send + Sync>);   // borrowed events
}

4. Pipeline (all in proef-core; pure — inputs include injected run_id, now, env snapshot)

4.1 Load packs. The fragment corpus ([run] fragments) is read once per command into a pack::FragmentCorpus and scanned at most once, lazily — only when some pack actually carries a ref:, which is what makes “pointing at a corpus you did not write costs nothing” true of the scan (CONFIG.md). One proef test loads packs up to four times (the suite, then [run] setup/teardown, each validated and then run) against the same corpus, so the memo belongs with the corpus rather than the caller — a caller that scanned eagerly to share the result would trade the promise for the speed. Discover embedded helpers/ + project packs/; serde_norway with deny_unknown_fields; validation passes: (1) match: guard rails — must contain literal text, no adjacent captures, unclosed braces rejected; (2) params/defaults coverage; (3) duplicate macro names across packs → error (qualify pack.yaml#name); (4) use: cycle + depth ≤ 32; (5) unknown with: keys → “did you mean” (edit distance); (6) finite-retry lint — retry: requires a finite count (ADR-0007); (7) hurl blocks: lower a probe instantiation with placeholder params and parse_hurl_file it — syntax errors reported with block-relative spans mapped to pack file/line; (8) engine kinds: every step kind must be claimed by a registered engine’s StepKindSpec.

Fragments (ADR-0018) add five: (9) a ref: must name a loaded fragment → “did you mean” over the scanned names, or a pointer at [run] fragments when none loaded; (10) duplicate fragment name across files → error (qualify file.hurl#name, the same suffix matching pack.yaml#name uses); (11) a step is ref: xor a payload/use:; (12) an option family declared both in the fragment’s own [Options] and as the step’s YAML key → error (pass 6’s twinned-option rule, applied across the file boundary), and the same rule name-to-name for a variable both supplied by the fragment’s [Options] variable: and given by a bind: — the pair reaches one entry as two variable: k= lines, where hurl’s last-wins would drop the bound value into every later entry too; (13) a fragment file the engine’s FragmentSupport::scan could not read, or an annotation it could not attach, positioned in the .hurl file itself.

Fragments skip pass 7: they parse as authored, so the probe instantiation has nothing to guess. Whether every {{…}} a fragment reads is actually supplied — by a bind: in scope, by an earlier step’s [Captures], or by the fragment’s own [Options] variable: — is checked at lower time (§4.4, proef::lower::unbound_placeholder) — only lowering knows what the preceding steps captured, and a load-time half-check would be worse than one complete one. The self-supplied source is scoped to the fragment that declares it: a name another entry left in hurl’s shared set is the implicit inheritance this check exists to refuse, so only [Captures] carries a value forward.

4.2 Parse features. gherkin 0.16 (Feature::parse); tags from Feature/Rule/Scenario accumulate. Localized (# language:) features are supported and test-covered: the crate strips the dialect keywords, proef consumes the stripped step text, and a localized outline with Examples expands like any other (outline detection keys on Examples presence, which is dialect-independent). Caveat: a localized outline whose Examples block is omitted cannot be told apart from a plain scenario (the crate’s dialect keywords are private), so it surfaces as an unbound-step error rather than the crisp no_examples — a worse message, never a silent pass. # comment lines are plain gherkin comments (no # key: directive mechanism — variables live in proef.toml, ADR-0012).

4.3 Bind. For each step (keyword stripped): first macro whose match: pattern matches wins; ambiguity (2+ matches) is an error listing candidates. Captured {name} values: trimmed, surrounding quotes shed (quotes preserve inner spaces/commas). Data table rows | key | value | merge into args; key set by both capture and table → error. Defaults fill; missing required params → error. Unbound step → exit-2 error with closest-pattern suggestion (edit distance over pattern literals).

4.4 Lower. Outline expansion (parser does not do it): per Examples row, substitute <col> in scenario name, step text, docstrings, table cells; ragged rows / unknown placeholders → parse-time error with line. Background steps prepend to every scenario. Macro expansion: params bound, ${…} resolved recursively, depth ≤ 8 (captured args may contain ${…} — spike-verified necessity); $${ escapes; {{…}} passes through untouched. Assert-only macros (expect:) merge into the previous request entry — error if none (Then-before-When). Product: Vec<StepBatch> per scenario (contiguous same-engine runs; batch maximally — split only at optional: boundaries and engine changes, ADR-0010).

4.5 Emit. Canonical .hurl per scenario (stable formatting, snapshot-tested): header comment per entry # <file>:<line> — <step text>; # optional markers; sidecar <slug>.map.json (schema: { entries: [{hurl_lines: [a,b], feature: {file,line,text}, optional, captures: [names], batch: n, step: n}], schema: 1 }); <slug>.vars when ${global:} or ${secret:} referenced (secrets as names only). Artifact dirs: .proef-runs/<id>/ artifacts/ (per-run) and proef artifacts -o <dir> (stable hand-off).

4.6 Dispatch. Per scenario thread: check token → factory.open(ctx) lazily per engine on first batch → session.run_batch(batch, world, events, token) in order → merge outcomes/World → finish() all sessions (reverse open order) → emit events. optional: batch failure → warnings + continue; else fail-fast within the scenario.

5. proef-engine-hurl internals (verified seam facts inline)

Adapter. Per batch: seed VariableSet from World (insert; secrets via insert_secret) → parse_hurl_file(&batch_text) → run_entries(&file.entries, &batch_text, Some(&input), &runner_options, &variables, &mut stdout_buf, Some(&listener), &mut logger) with WriteMode::Buffered terms (upstream’s own threading mode: parallel/worker.rs:76,124-133) → map EntryResults to StepOutcomes via the sidecar (SourceInfo spans → feature lines) → merge HurlResult.variables back into the World (typed).

RunnerOptions mapping. Batch-level RunnerOptionsBuilder from config; per-entry [Options] override batch defaults by clone-then-override (runner/options.rs:43-58), variable= inserts persist for the rest of the call — verified semantics, relied upon.

HttpDefaults carries what [http]/[env.<name>.http] express: timeout-ms, follow-location, max-redirs, insecure, proxy/no-proxy, cacert, client-cert/client-key, user-agent, cookie-store. Each reaches the builder only when the project set it, so an unset key leaves hurl’s own default rather than this engine restating a constant that could drift from upstream on the next pin bump. Path-valued fields arrive already resolved — the CLI applies the one-path rule and core touches no filesystem (ADR-0012). Two keys are exceptions to the per-entry override rule above, verified against OptionKind rather than inferred: it has no UserAgent variant (run-wide; an entry opts out with a User-Agent: request header instead) and no cookie variant at all, so cookie-store is run-wide with no per-entry spelling whatsoever.

cookie-store = false maps to use_cookie_store(false) (hurl’s --no-cookie-store, 8.0.0). Verified: hurl enables curl’s cookie engine only under use_cookie_store (http/client.rs:390), which also means a cookie_input_file handed over with the store off is silently ignored — so the engine skips both halves of the §5 round-trip when the store is off. hurl’s own FIXME there (a handle once given cookie storage cannot have it removed) never reaches proef: the client is per-call (below), so a handle never transitions on → off.

Client lifetime (verified). run_entries constructs http::Client::new() internally per call (runner/hurl_file.rs:169) — fresh libcurl handle: connection cache and cookie jar do not survive across calls. Consequences implemented: batch maximally (§4.4); on forced splits, chain variables via HurlResult.variables (lossless) and, when cookies are in play, round-trip HurlResult.cookie_store → CookieStore::to_netscape() → temp file → next batch’s RunnerOptionsBuilder::cookie_input_file (http/cookie_store.rs:66-72, runner_options.rs:242-244) behind a SessionState struct. Upstream patch #1 (ADR-0003): run_entries(&mut Client) — two internal call sites (hurl_file.rs:124, worker.rs:124); adopt when accepted, delete SessionState cookie path.

Thread-safety (verified). No global mutable state in hurl (static mut/ lazy_static/OnceLock-mutable: zero hits); libcurl init is Once-guarded pre-main by the curl crate; sole FFI global write is libxml2’s error handler set idempotently per XPath eval — exercised concurrently by upstream itself. Scenario-per-thread is safe.

Cancellation & budgets (ADR-0007). No interrupt support exists upstream (verified); engine computes batch budget = Σ(timeout × (retry+1)) + intervals + margin, clamped to a four-hour ceiling (MAX_BATCH_BUDGET — lint-clean values still compose into an unbounded product, ADR-0007 amendment); watchdog abandons over-budget scenario threads; token checked between batches only.

Failure detail. Engine errors surface through hurl’s own DisplaySourceError::description into StepFinished.detail (additive event field), the console, JUnit, and the GitHub summary.

6. Pack schema v1 (normative field reference)

bind:                         # pack-scope fragment bindings (ADR-0018); macro and
  <var>: "${…}"               # step scope override, most specific winning
macros:
  <macroName>:                # unique across packs; qualify as pack.yaml#name on clash
    params: [a, b]            # required unless defaulted
    defaults: { b: "x" }      # optional params
    match: "…{a}…"            # step-definition pattern; omit → not Gherkin-reachable
    description: "…"          # docs + desktop palette later
    tags: [Domain]
    steps:                    # request steps (each lowers to ≥1 hurl entry)
      - name: "…"             # entry label (events/console)
        optional: true|false  # failure → warning (segments the batch)
        when: "${expr}"       # skip guard: skips when empty or literal false/0 after resolution
        retry: { count: N, interval_ms: M }   # finite only (lint); → [Options] retry
        saveAs: { captureName: global }        # promote capture(s) into the World
                              # (refused with a warning if the value equals a secret)
        bind: { <var>: "…" }  # step-scope fragment bindings (only with `ref:`)
        hurl: |               # PRIMARY form (ADR-0004): raw hurl, ${…} lowered first,
          …                   # {{…}} left for run time; validated by parse_hurl_file
        # OR structured payload (reserved for future non-hurl engines):
        # <kind>: { … }  (a future engine's structured payload)
      - ref: file.hurl#name   # ALTERNATE form (ADR-0018): one `# @proef <name>` entry
                              # of a scanned fragment file; `{{…}}` supplied by bind:,
                              # every one bound, captured by an earlier step, or set
                              # by the fragment's own [Options] variable: (which then
                              # may not also be bound — option_declared_twice)
      - use: pack.yaml#other  # composition, with:/inline args; cycle+depth checked
        with: { a: "${a}" }
    expect:                   # assert-only macro (Then-steps): merges into previous entry
      - status: "${status}"   # or raw hurl assert lines: hurl: |‐style fragment (M5)

JSON Schema is schemars-derived from these serde types plus engine-contributed StepKindSpec fragments; proef schema --add-to injects the editor modeline (a proven mechanism).

7. Gherkin mapping reference

One scenario = one flow/run-record/artifact set. Background: prepends. Rule: groups pass through (tags accumulate). Outline/Examples per §4.4. Data tables per §4.3. Docstrings: reserved for raw request bodies in generic steps (M5). Tags: @tag → flow tags, --tags filters by a boolean expression (and/or/not/parens, proef_core::tags); atoms may glob — * any run, ? one char, anchored, a metachar-free atom staying literal equality — and the one matcher serves --tags, [run] exclusive-tags and [tag-links]. @skip/@skip:reason and @quarantine are reserved tags carrying behavior, recognized at the CLI edge — core never reads tags (ADR-0019, ADR-0014). Step keywords: And/But resolve to the previous primary keyword (gherkin crate StepType); keyword itself is not matched against patterns.

8. Variables reference (ADR-0005)

Author-time (${…}, resolved in §4.4, recursive ≤ 8): ${param} · ${env:NAME} / ${env:NAME:-default} · ${url:key} / ${vars:key} (proef.toml [url]/[vars], base + active [env.<name>] deep-merged; injected — ADR-0012) · ${run:id} (uuid-v7-derived, injected) · ${global:key} (World read at lower time of the scenario) · ${secret:NAME} (encrypted store; emits {{secret_name}} + insert_secret) · ${fake:kind} (deterministic from run id and an occurrence index; the index is an incrementing counter scoped to one scenario — shared across every ${fake:…} resolve in it, so independent references never collide regardless of how many a step resolves; a step’s name: label resolves from a rewound copy of the counter so its Nth ${fake:…} reference reuses the payload’s/when:’s Nth occurrence by position, not generator kind — reproducing the payload’s own value exactly when the label’s references mirror the payload’s in kind and order, otherwise surfacing that occurrence’s own-kind value instead — then restores the real counter to the high-water mark the replay reached (never below it), so an extra fake the label alone introduces still reserves its slot and is never reissued; the counter resets to zero at the next scenario, so the same generator at the same position in two different scenarios currently coincides; port deterministic NL generators) · $${…} literal escape. Run-time ({{…}}): hurl captures and secret placeholders — resolved by the engine. Resolution order within a scope: step args

macro defaults.

9. Diagnostics (ADR-0009)

gherkin Span = 0-based byte offsets, end-exclusive → SourceSpan::new(start, end-start); parser appends a trailing newline when missing — attach the normalized source text to diagnostics (or clamp); never use LineCol.column (char-counted) in byte math. Pack YAML: serde_norway error locations; schema-path → YAML-location pass for lint findings. Engine failures render: feature line + step text + assert detail + artifact path:span (from sidecar). Every diagnostic carries a stable code (proef::pack::adjacent_captures, …) for greppability.

10. CLI reference (v1)

proef [--config PATH] [--env NAME] <command>   # global: the proef.toml to read, the [env.<name>] profile
proef init [dir]
proef test [file|dir] [--env NAME] [--dry-run] [--tags EXPR] [--jobs N] [--junit path|auto]
                      [--format json|tap] [--watch] [--scenario NAME] [--scenario-file FILE]
                      [--run-id ID] [--rerun] [--sarif PATH (with --dry-run)] [--max-fail N] [--shard I/N] [--shuffle] [--meta KEY=VALUE]... [--console MODE]
proef flows [file|dir] [--env NAME] [--format json]
proef macros [file|dir] [--env NAME] [--format json]
proef fragments [file|dir] [--env NAME] [--format json] [--check [--require-annotated]]
proef artifacts [file|dir] -o DIR [--env NAME] [--run-id ID]
proef schema [--add-to FILE…]  proef secret set|list|rm
proef explain [run-id] [--format json]          proef doctor [--format json]
proef diff [base] [new] [--fail-on-regression] [--format json]   # each side: run id, record dir, or events .jsonl
proef flaky [--format json] [--by KEY] [--min-samples N] [--recovery-runs N] [--outage-rate RATE]
                      # verdicts over the retained run history, keyed by the input fingerprint
proef report [run-id] [-o FILE]
proef fmt <file|dir> [--check]
proef lsp

A path-less test/flows/artifacts resolves [run] suite then the tests/ convention (else exit 2). Exit codes: 0 ok · 1 test failure · 2 user error · 3 system error (typed enum, assert_cmd-pinned). Ctrl-C, SIGTERM and SIGHUP all take the graceful cancel (ctrlc’s termination feature — a CI job timeout is a cancellation, not a kill); a second signal while a test/watch run is cancelling forces an immediate hard exit with code 130 (128+SIGINT, the shell convention; the handler carries no signal identity, so every second signal shares the code) — deliberately outside the graceful 0/1/2/3 taxonomy, so it is not an ExitCode variant. Config precedence: built-in defaults < proef.toml base tables < active [env.<name>] (selected by --env/PROEF_ENV) < flags; suite variables ${url:key}/${vars:key} resolve from [url]/[vars] deep-merged with the active env (ADR-0012). Secrets additionally resolve PROEF_SECRET_<NAME> env overrides before the store, and PROEF_KEY (base64) overrides the key file — CI decrypts a committed ciphertext store without the key ever touching disk. --dry-run = §4.1–4.5 including artifact parse-validation; no engine sessions, no network.

11. State & files

.proef-runs/<run-id>/ → events.jsonl (the record, ADR-0008), run.log (console tee), artifacts/*.hurl|.map.json|.vars (+ artifacts/assets/<slug>/, the staged file bodies), timings.json (per-scenario durations, read back by --shard-weights) and inputs.json (the input fingerprint, proef flaky’s equivalence class) — both derived sidecars, never a second record — report.html, report.junit.xml (when requested); [run] keep-runs-bounded rotation, default 200 (only uuid-named run records rotate; the in-flight run never does). .proef-state.json — persistent World: atomic temp+rename, 0600. proef.toml — project config: runner settings ([run] jobs/runs-dir/keep-runs/suite/fragments/setup/teardown/exclusive-tags, [http] timeouts, [sla] ceilings, [flaky] verdict thresholds) + suite variables ([url]/[vars]) + run metadata ([meta], ADR-0020) + report tag links ([tag-links]) + per-environment overrides ([env.<name>]); see docs/CONFIG.md, ADR-0012.

One path rule: a path written in proef.toml resolves against the directory holding proef.toml; a path typed on the command line resolves against the working directory. Absolute values are taken as written, and with no config in scope written paths stay relative to the working directory. This covers suite, fragments, setup, teardown, runs-dir, .proef-state.json and .proef-secrets.json — so a project is where its config is, not where the shell is, and running from a subdirectory reaches the same suite, records, World and secrets as running from the root.

And its naming dual: a path that reaches an artifact, sidecar, event, report or diagnostic is spelled relative to that same directory (front::SourceNaming). Resolution makes a written path absolute; naming makes it project-relative again, so nothing proef records names the machine that produced it — the four ways to point at one suite (derived, typed, typed absolute, from a subdirectory) emit one artifact byte-for-byte, which is what ADR-0010’s same-bytes contract means across machines. A path that arrives relative is recorded exactly as it arrived; one that lies outside the project keeps its absolute name, there being no project-relative spelling of it. DiskSourceProvider (proef lsp) is deliberately outside this: it keys document identity on absolute names.

12. Parallelism & cancellation

Scenario-per-OS-thread, --jobs bounded (default: available_parallelism, capped by scenario count); events funneled through the sink (the console reporter buffers per scenario and replays contiguously); one CancellationToken per run, child per scenario; Ctrl-C graceful / second Ctrl-C hard (ADR-0007). Global-World writes serialize through the store lock; scenario ordering within a file is preserved for artifact naming, not execution order — except that a scenario matching [run] exclusive-tags waits for the pool to drain and runs alone, which constrains concurrency rather than reordering the queue.

13. Security

Secrets: encrypted at rest (chacha20poly1305), surfaced only via insert_secret; redaction invariants (never in artifacts/events/reports/logs) property-tested; a saveAs: global capture whose value equals a known secret is refused (warned) — .proef-state.json is plaintext and never receives secret-derived material; sensitive files 0600; proef doctor reports store/key health. context_dir confines file bodies (hurl’s own sandbox option) to the scenario’s staged asset root — a directory holding only what proef put there, so the sandbox is narrower than the suite it used to be. Assets are staged from beside the source that referenced them (feature for inline, fragment for ref:), which is hurl’s own per-file rule; references must be plain relative paths, and one name may come from only one source per scenario. No telemetry.

14. Dependencies (exact at M0; policy: latest stable at adoption, Renovate-managed)

Engine: hurl =8.0.1, hurl_core =8.0.1 (--locked; ADR-0003). Core: gherkin 0.16, serde 1, serde_json 1, serde_norway 0.9, schemars 1, thiserror 2, tokio-util 0.7 (default-features = false; CancellationToken only). CLI: clap 4 (derive, env), miette 7 (fancy), uuid 1 (v7), notify =8.2.0, ctrlc, chacha20poly1305 rpassword base64, quick-junit, toml. LSP: lsp-server 0.7, gen-lsp-types 0.11 aliased as lsp-types with its url feature (proef-lsp’s stdio transport, wired into proef lsp; ADR-0017 amendment). Fixture/harness (dev): tiny_http (ADR-0011 — axum conflicts with the tokio-runtime ban), libtest-mimic. Engine runtime: tempfile (Netscape cookie round-trip between batches, §5). Dev: insta assert_cmd predicates proptest tempfile quick-xml + cargo-fuzz targets; openssl-sys rides as the engine’s vendored-openssl feature carrier. Synthetic data (${fake:*}) is a dependency-free SplitMix64/FNV implementation in-core — the fake crate was not needed. Datetime uses jiff, never chrono, in our own code (hurl’s internal chrono is its business) — currently jiff is a dev-only dependency of the fixture’s /health identity; the sans-IO core still reads no clock (injected timestamps only). Banned: serde_yaml/serde_yml, chrono (ours), reqwest (superseded), maybe-async, async-trait (v1). Build prereqs (doctor-checked): Debian build-essential pkg-config libssl-dev libcurl4-openssl-dev libxml2-dev libclang-dev; macOS: Xcode CLT.

15. Conventions

Toolchain pinned to latest stable adopted at its x.y.1 point release — ~3-4 weeks after x.y.0, so a new minor waits out its first patch (1.97.1 at writing; tools and third-party crates track latest immediately, and exact pins like hurl outrank everything). rust-toolchain.toml cites this section as its authority, so the policy is stated here rather than only in RELEASING.md, where the correction first landed (R18-2): an unwritten policy contradicting the written one is a docs defect, and fixing it in two files while the normative spec still said “always latest stable” left the contradiction in the source of truth. Edition 2024; resolver 3; workspace lints (the legacy suite table); CI gates (PR): cargo fmt --check, clippy --all-targets --all-features -D warnings, cargo nextest run, cargo test --doc, RUSTDOCFLAGS="-D warnings" cargo doc, cargo deny check, cargo machete, zizmor, xtask docs-check, proef doctor, fuzz smoke + xtask public-api (nightly rustdoc); cargo audit runs on the nightly schedule. Automation in xtask (+ just aliases); no shell scripts for logic.

proef — Implementation Plan

Status: M0–M5 delivered, and everything after them — external config and environments (ADR-0012), the v0.6–v0.8 correctness series, v0.9.0, and named hurl fragments (ADR-0018). M6 (future engines) remains unscheduled. CLAUDE.md’s Status block is the running ledger; the milestone sections below describe the plan as it was executed, not the current frontier. · Date: 2026-07-28, status refreshed 2026-08-12. Normative design: TECH-SPEC; decisions: ADRs. Sizes are t-shirt (S ≈ days, M ≈ small weeks, L ≈ multi-week) — deliberately not fake-precise.

0. Guiding principles

The porting rule: when the prior spike proved a feature, understand the feature and implement it the cleanest way in this architecture — no copy-paste, no compatibility with spike code. Every milestone lands green on all CI gates. Core stays pure (no IO/clock/rand — TECH-SPEC §4). Nothing merges without its tests (TESTING-STRATEGY). The spike (research/) is evidence, not a starting codebase.

Global definition of done (every milestone)

fmt, clippy -D warnings, nextest, doctests, rustdoc -D warnings, deny, audit all green · new behavior has unit + (where applicable) snapshot/property tests · no unwrap/ expect in library code paths (CLI main excepted) · public items documented · CHANGELOG entry · exit codes unbroken (assert_cmd suite).


M0 — Foundations (size M)

Objective: compilable, gated, empty-but-shaped workspace with the seam in place.

Tasks:

  1. Init repo proef/; virtual workspace (resolver 3), workspace.package + workspace.dependencies + workspace.lints; rust-toolchain.toml pinned to current stable (1.97.1 at writing); deny.toml, nextest config, README.
  2. Crates: proef-core, proef-engine-hurl (empty adapter), proef-cli (clap skeleton), xtask (+ justfile aliases). Reserve crates.io names (0.0.0 placeholders, publish = false locally thereafter).
  3. Core scaffolding: ExitCode enum + CoreError/EngineError taxonomy (ADR-0009); Event enum v1 + EventSink (ADR-0008); World/Value types + GlobalStore (atomic temp+rename) (ADR-0005); EngineFactory/EngineSession/StepKindSpec/ DoctorCheck traits (ADR-0002); CancellationToken plumbing (ADR-0007).
  4. CI: gates workflow + a stub canary job (builds proef-engine-hurl against hurl =8.0.1 — becomes the real canary in M4); Renovate config (pins grouped; hurl excluded from auto-bump).
  5. proef-cli: doctor (native-lib checks from engine doctor() — first proof the capability hook works), --version, exit-code integration tests.

Acceptance: workspace builds on Linux+macOS CI; proef doctor reports libcurl/ libxml2 status via the engine-contributed check; assert_cmd pins exit codes; all gates green. Proves: ADR-0002 seam compiles and registry assembly works.

M1 — Front end: packs, binding, Gherkin, lowering (size L)

Objective: .feature + packs → validated, lowered scenarios; --dry-run without artifacts.

Tasks:

  1. Pack model (serde_norway, deny_unknown_fields) incl. hurl: raw blocks, expect:, use:/with:, when:, retry:, saveAs:; schemars derivation + engine StepKindSpec fragment merge; proef schema [--add-to].
  2. Validation passes 1–8 (TECH-SPEC §4.1) with miette diagnostics + stable codes; finite-retry lint; probe-instantiation parse of hurl blocks via hurl_core.
  3. Matcher: {name} tokenizer + leftmost matcher + guard rails (cucumber-expression semantics; property tests: no-panic, quote round-trip, adjacent-capture rejection).
  4. Gherkin: parse, tags, Background, Rule pass-through, outline expansion, data-table merge; binding with ambiguity detection + closest-pattern suggestions.
  5. Lowering: macro expansion (cycle/depth), recursive ${…} resolver (depth 8, $${ escape; property + fuzz targets), Then-merge (expect: → previous entry), batch segmentation (maximal; splits at optional:/engine change).
  6. proef flows, proef test --dry-run (no emit yet), corpus test over tests/.

Acceptance: the four 500-series features (+ seeded error corpus) dry-run with line-accurate diagnostics; property/fuzz targets in CI (fuzz smoke = N seconds, full = nightly). Proves: PRD US-2/3/6 front-end half.

M2 — IR, emitter, artifacts (size M)

Objective: lowered scenarios → canonical .hurl + sidecars; --dry-run complete.

Tasks:

  1. Canonical emitter (stable formatting rules) + # optional markers + per-entry feature-ref comments; line-map construction.
  2. Sidecar <slug>.map.json (schema v1) + <slug>.vars; proef artifacts -o DIR; per-run layout under .proef-runs/<id>/artifacts/.
  3. --dry-run gains artifact parse-validation (parse_hurl_file on every emitted file — the real parser as the validator).
  4. insta snapshot suite: features+packs → artifacts + sidecars (golden corpus).

Acceptance: spike parity — the 500-series features emit artifacts that stock hurl parses (checked in CI via the canary toolchain image); snapshots reviewed. Proves: ADR-0010 emit half; US-7 static half.

M3 — engine-hurl: embedded execution (size L)

Objective: proef test runs scenarios end to end via embedded hurl.

Tasks:

  1. Adapter: VariableSet seeding (World + secrets via insert_secret), parse_hurl_file + run_entries with Buffered terms + EventListener → events; EntryResult → StepOutcome mapping via sidecar spans.
  2. RunnerOptions mapping (config → builder; per-entry [Options] override relied on as verified); HurlResult.variables merge-back; saveAs: global promotion; typed Value bridging.
  3. Segmentation runtime: optional: warn-and-continue; SessionState cookie round-trip (Netscape temp file) for split scenarios; variables chaining.
  4. Parallelism: scenario threads + --jobs; budgets + watchdog + token checks; Ctrl-C graceful/hard paths.
  5. Reporters v1: console BDD tree (attempts/timings), JSONL event record, run-record rotation; --output json.
  6. Fixture server (tiny_http, ADR-0011) + integration suite: success/4xx/retry-delay/auth/ malformed-JSON/optional/World-chaining/cancellation-budget cases.

Acceptance: the 500-series features run green against the fixture with prose unchanged (US-1); failure demo maps to feature line + artifact span; exit codes correct under pass/fail/user-error/system-error; cancellation bounded-time test passes. Proves: ADR-0001/0005/0007 runtime; US-1/4/5/9.

M4 — Upstream tracking hardened + CI reporters (size S–M)

Objective: riding upstream is a runbook, not a risk; CI outputs complete.

Tasks:

  1. Canary job real: build+test against next hurl release (scheduled + on release); failure opens an issue with the diff of HurlResult behavior.
  2. Thin-fork rehearsal: apply a scratch one-commit patch via [patch."crates-io"] from the fork tag, build, revert — documents the mechanics; draft upstream PR #1 (run_entries(&mut Client)) from the verified two-call-site change.
  3. JUnit (quick-junit) + GitHub job summary reporters (--junit auto under GITHUB_ACTIONS); reporter-stack decorators (Normalize/Summarize) formalized.

Acceptance: canary catches an artificially-pinned older/newer hurl mismatch in rehearsal; JUnit consumed by CI UI. Proves: ADR-0003; US-8/11; M-4 metric path.

M5 — Breadth + integrations (size M)

Objective: the conveniences that make proef the daily tool.

Tasks: multipart/form/docstring bodies through packs; expect: raw-hurl assert fragments; more [Options] exposure (delays, location); --watch (notify 8.2.0); proef explain; proef secret set|list (encrypted store port); ${fake:*} NL generators; libtest-mimic harness (one Trial per scenario; nextest contract: --list --format terse, --exact --nocapture) + docs for IDE use; proef fmt for pack hurl blocks.

Acceptance: US-10/12 green; nextest runs the suite; watch-mode demo. Proves: ADR-0008 harness leg; PRD v0.3 scope.

M6 — More engines (future; sized when scheduled)

A future non-hurl engine — its step vocabulary carried as structured step payloads rather than raw hurl. Structural acceptance test: git diff --stat proef-core is empty. Mixed-engine 500-series suites become runnable. (Note 2026-07-29: a driver bringing its own async runtime would conflict with the tokio-runtime ban — pick a sync driver at sizing time. The core’s structured-payload paths are already exercised by tests; no core work is expected.)


5. Sequencing & parallelization

M0 → M1 → M2 → M3 strictly ordered (each consumes the previous stage’s types). Within M1, tasks 1–3 parallelize with 4; M2.4 can start as soon as M2.1 emits. M4.3 reporters can start during M3.5. Docs (GETTING-STARTED + AUTHORING, delivered 2026-07-29) draft during M2–3 and finalize at M5 (deliberately after the schema stabilizes).

6. Risk register

RiskLIMitigation / trigger
hurl minor release breaks the seamMMexact pins; canary (M4); thin-fork shim; pinned seam integration test
run_entries #[doc(hidden)] churnMMsame as above + upstream PR #1 conversation opens a stability dialogue
Native build prereqs trip up a machineMLdoctor first-run UX; README one-liner; CI images pre-baked
Segmented scenarios lose connections/cookiesLLbatch-maximally; SessionState; upstream patch #1 erases it
Runaway scenario under retriesLMfinite-retry lint; budgets + watchdog (ADR-0007); CI timeout
Pack schema churn post-v1MMschema: 1 field; additive-only until v1.0; snapshots catch drift
Event-schema consumers breakLMversioned events; JSONL replay tests
gherkin-crate stagnation returnsLMactive again (0.16); parser is replaceable behind core’s parse stage

7. Runbook — absorbing a new hurl release

  1. Canary red/green report arrives (scheduled job). 2. Read upstream CHANGELOG diff.
  2. Bump pins on a branch (=X.Y.Z, --locked), run full suite + snapshots. 4. If breakage: fix adapter; if upstream regression or removed seam: add minimal patch on the fork branch, consume via [patch], open upstream PR, note in ADR-0003 log. 5. Merge; tag; update TECH-SPEC §14 versions. 6. If the fork carries patches: rebase them onto the new release tag; drop any that merged upstream.

8. Day-one checklist

git init proef && cd proef → commit toolchain+workspace manifests (M0.1) → cargo new the four crates → copy workspace lints/deny/nextest configs → wire CI → first green pipeline → open M0 tracking issue with this plan’s task list. Suggested first PR sequence: M0.1+M0.2 together, then M0.3 split by module, then M0.4/M0.5.

proef — Testing Strategy

Status: normative · Date: 2026-07-28 · tools: cargo-nextest, insta, proptest, cargo-fuzz, assert_cmd, tiny_http fixture. Everything below is device-free and CI-green with no external network.

1. The layers

Unit (every crate): matcher tokenization/matching edge cases; resolver escapes and depth cap; Then-merge rules; batching/segmentation boundaries; sidecar math; World error→exit-code mapping.

Property (proptest): matcher — arbitrary patterns/text never panic, valid pattern+generated text round-trips captures; resolver — $${…} escape round-trip, resolution is idempotent once fully resolved, depth cap always terminates; secret-mask invariant — for arbitrary events/reports containing a known secret value, rendered output never contains it, nor any of its derived encoded forms (base64 both alphabets ± padding, hex both cases, percent-encoding, JSON-string escape — ADR-0005 as amended), with a companion property pinning that text free of the secret and its forms passes through untouched; World — snapshot/restore is an involution; fragment scanner (proef-engine-hurl) — over generated hurl files, every reported line lies inside the file, entries are accounted for exactly once, starts are ordered and distinct, and no fragment’s text runs into the entry after it. That last one is the entry-boundary arithmetic’s whole job, and it is asserted because a draft without it passed while the boundary was deliberately broken. The scanner is proptested rather than fuzzed on purpose: it needs hurl_core, and cargo dependencies are package-level, so putting it in fuzz/ would compile hurl for every target there and drag native libraries into a job that has none.

Fuzz (cargo-fuzz, nightly job + PR smoke): fuzz_match_pattern (pattern×text), fuzz_resolve (template strings), fuzz_tag_expr (tag expressions), fuzz_pack_load (YAML bytes → loader must error, never panic), and fuzz_fragment_binding (a pack against a real corpus: ref: resolution, unread bind: keys, a bind: colliding with a variable the fragment supplies itself). Parser-adjacent hand-written code is exactly where fuzzing pays.

Both loops take their target list from cargo fuzz list, never a list written into a workflow. The names used to be spelled out in ci.yml and nightly.yml, so a target ran nowhere until both were edited and nothing failed to say so.

fuzz_fragment_binding is structure-aware — it builds a well-formed pack and corpus from the input rather than hoping the fuzzer discovers one. That is a measured choice: a byte-oriented version never resolved a single ref: in 1.45 million runs, because reaching those rules means finding valid YAML and a matching corpus name at once. When adding a target that needs structure, verify it reaches the code by probe — panic on the condition under test, run briefly, confirm it fires — because a target that compiles and finds nothing reads exactly like a target that compiles and finds no bugs.

fuzz/ is its own workspace (the root Cargo.toml excludes it, since fuzzing needs nightly), so no root-workspace command compiles it: a changed proef-core signature breaks the targets while every root-workspace gate stays green, leaving the fuzz jobs as the only signal. cargo check --manifest-path fuzz/Cargo.toml --all-targets runs on the pinned stable toolchain in seconds, so the gates job carries it — earlier than the fuzz smoke, and on both gate platforms.

Snapshot (insta): emitter — golden corpus of (features + packs) → artifacts + sidecars, byte-stable (the canonical-format compatibility surface, ADR-0010); diagnostics — rendered miette output for the seeded error corpus (every validation pass in TECH-SPEC §4.1 has at least one golden failure); proef schema output; event-stream JSONL for a reference run (with injected clock/run-id — core purity makes this deterministic).

Integration (fixture server): synchronous tiny_http dev crate (proef-fixture) modeled on the spike’s fixture — not axum as originally written: axum’s tokio requirement conflicts with the workspace’s no-async-runtime ban (ADR-0006/0007 + deny.toml), and a sync fixture keeps that invariant binary-wide (errata 2026-07-28, M3). Endpoints, extended: bearer-auth endpoints, search, create (201/422 paths), delayed push-visibility (exercises retry for real), cookie-setting endpoints (exercises SessionState round-trip), slow endpoint (exercises budgets/watchdog), malformed-JSON endpoint. Suite covers: green path (the four 500-series features), capture chaining, World/global across scenarios, optional: warn-and-continue, cancellation (token cancel mid-run completes within budget, reports written), parallel --jobs determinism (event Normalize), artifact↔execution same-bytes assertion (hash the emitted file and the text handed to parse_hurl_file).

Fragments (ADR-0018): crates/proef-cli/tests/fragments.rs builds a self-contained project per test — its own proef.toml, corpus and pack in a temp dir — because the reference corpus under tests/ is config-independent by design (several tests run it from a temp cwd with settings passed by environment variable and no proef.toml in scope), so anything needing [run] fragments cannot live there. That is also why four diagnostic codes are covered here rather than in tests/errors/ (DIAGNOSTICS.md says which).

The headline case runs one file under both runners: proef test against the fixture, then stock hurl invoked on the same bytes with an equivalent variables file, asserting the corpus comes back byte-identical. The engine is embedded, so a hurl binary is not a build requirement — that half skips with a printed note when none is on PATH rather than being faked. Provenance is asserted at both ends: the JSONL record for the event-driven readers, and --junit for the RunSummary-driven ones, since those are fed by a second copy that a green suite would not otherwise exercise.

LSP over stdio (crates/proef-cli/tests/lsp_stdio.rs): the real binary, spoken to as an editor does. This is the only place proef.toml → DiskSourceProvider → document URI is exercised end to end: the proef-lsp unit tests inject absolute source names through a fake provider, so a config-layer change can break every go-to-definition while they stay green — which has happened. These tests canonicalize their temp root, because on macOS a tempdir is /var/… whose real path is /private/var/…, and without that any cwd-relative path logic silently no-ops and the test passes without reaching the behaviour.

Corpus: --dry-run over every .feature in tests/ — the suite’s own features are the regression corpus.

Documentation (xtask docs-check + crates/proef-cli/tests/docs.rs): the docs make claims a machine can settle, so they are settled mechanically rather than by review. docs-check reads files, and does six things: every workspace crate appears in TECH-SPEC §2 and CLAUDE.md; every ADR file appears in the decision log; every diagnostic code the workspace emits has a DIAGNOSTICS.md row and vice versa; every release names each kind of change once (repeats accumulate by appending, which is how a changelog gets written); every relative link resolves; and every fenced toml/yaml example parses with the product’s own parsers, so the check means “proef would accept this”, not “some parser would”. tests/docs.rs needs a built binary and therefore lives with assert_cmd: it asks clap whether every documented command and long flag exists.

The split is a rule, not an accident, and each half states it in its own header — a check that reads files belongs in docs-check even when a test would be easier to write, because the doc-only CI step is the fast one and a file-reading check placed in the test suite silently stops running there.

Both were written against defects that had already shipped — an ADR whose first example could not load, and a row marked shipped that named a --html flag which never existed. In each the surrounding prose was correct, which is precisely what a careful reader does not catch. Two scoping rules keep them honest rather than noisy: only the indexed corpus is linted (docs/superpowers/ is a dated archive, and editing history to satisfy a checker is the wrong direction), and command detection is restricted to code spans and fenced blocks — prose says “proef discovers packs”, and treating that as an invocation produced sixty false positives against four real ones. Names the docs discuss as proposals are listed explicitly in tests/docs.rs, so adding one is a decision rather than the check going quietly soft.

CLI (assert_cmd): exit codes 0/1/2/3 pinned per command and failure class; --format json schema-checked; --junit well-formed (quick-junit round-parse).

Canary (M4): scheduled + on-release job builds against the next hurl version and replays the integration suite; red = issue with behavior diff, pins never auto-move (runbook: IMPLEMENTATION-PLAN §7).

2. What is deliberately NOT tested here

Hurl’s own HTTP semantics (asserts, filters, templating execution) — that is upstream’s test surface; proef tests the adapter contract (options mapping, variable bridging, span mapping, segmentation) against the fixture instead of re-verifying hurl. This is a direct consequence of ADR-0001 and the reason the differential-oracle harness from the research phase was retired.

3. CI matrix & gates

Linux (ubuntu-latest, prereqs pre-baked) + macOS on every PR; Windows weekly (vcpkg libs) while the port stabilizes, then per-PR (port green 2026-07-28: VCPKG_ROOT export, hurl’s crates.io-missing icon supplied in CI, /-normalized path identifiers). Gates: fmt, clippy -D warnings, nextest (all crates), doctests, rustdoc -D warnings, deny, cargo-machete, zizmor (workflow static analysis), xtask docs-check, proef doctor smoke, public-api snapshot, fuzz smoke (30 s/target), corpus dry-run, CLI suite, and — on a pull request — the changelog-entry check (§8). The complexity ratios run as their own step, alone, for the reason §7 gives. Snapshot tests (insta) run inside nextest — a drifted snapshot fails there, no separate step. Nightly: full fuzz (10 min/target), canary, cargo-audit (advisories against unchanged code — deny covers PRs). Coverage: measurable on demand, not gated in CI (P13’s local half). just cover runs cargo llvm-cov nextest over the workspace (just cover-html for a browsable report, just cover-lcov for a CI service’s lcov); the number today is ~90% line coverage of the unit + integration suites (doctests excluded — nextest does not run them). When a CI coverage job lands it must be a ratchet, not a fixed threshold — the 2026 norm and the only kind that suits a pre-1.0 codebase: fail a PR only if coverage drops, never on an arbitrary floor, and keep it informational (a PR comment) rather than a hard merge gate. A fixed percentage gate is explicitly the wrong shape here; it punishes honest additions of hard-to-cover error paths and invites coverage theater. The xtask binary’s low number is expected — it is automation exercised by running it, not by unit tests.

4. Test data management

tests/features/ — the real suite (also corpus input). tests/errors/ — seeded broken features/packs, one file per diagnostic code, name = expected code (golden snapshots). Insta snapshots live next to their suites (crates/proef-cli/tests/snapshots/), reviewed via cargo insta review. Fixture data is generated in-process (no committed binary blobs beyond one JPEG for multipart, M5).

5. Determinism rules (make flakes structural, not cultural)

Core purity (no IO/clock/rand — TECH-SPEC §4) means every non-integration layer is bit-deterministic by construction. Integration layer: fixture delays are token-driven (visibility timestamps), not sleep-raced; retry tests assert attempt counts, wall time only as generous upper bounds; parallel tests assert on Normalized event order, never raw interleaving. Any test needing “now” receives it as a parameter.

Retry-until-green is the anti-pattern, and that is why proef ships no scenario @retry. A scenario-level retry is the headline feature of several runners and is deliberately absent here: re-running a test until it passes hides precisely the defects worth finding. A bug that fails one run in four survives three retries 99.6% of the time (1 − 0.25⁴), so the suite reports green while the product is broken for a quarter of its users. proef’s shape is detect-then-quarantine: proef flaky returns a verdict over run history, @quarantine stops a known flapper gating the build while keeping it visible in every sink (ADR-0019), and per-step retry: covers the case that is genuinely polling — a resource that becomes visible on the Nth attempt — rather than rerolling a verdict. proef’s own suite is held to the same rule: a red test here is reproduced and filed, never re-run until it cooperates and then forgotten.

6. Every diagnostic code is named by a test

DIAGNOSTICS.md calls codes “a contract: they never change meaning”. A contract with nothing holding it to it is a wish — 23 of 75 were in that state when the rule was written: reachable in production, documented, and exercised by nothing at all, not even an assertion on their message text.

source_guards.rs enforces it. A code counts as covered when either a seeded tests/errors/<area>__<name>/ directory exists (the corpus driver dry-runs it, so the rendered diagnostic is exercised end to end) or the literal code string appears in a test. Naming the code, not matching the prose — the wording is expected to improve, while the code is the part that promises not to change.

Two codes are exempted by name with recorded reasons (source::unreadable, config::unreadable need a file the process may stat but not read, which CI runners do not reproduce because they run as root). The guard checks its own exemption list too: an exemption that outlives its code silently excuses nothing.

Reaching a defensive guard is worth the effort rather than a reason to skip it. lower::kind_unrouted fires only when the engine registry and pack validation disagree, so its test makes them disagree; lower::expansion_too_deep sits behind pack validation’s identical limit, so its test bypasses validation with load_collecting — the only way to hand lowering a graph validation would have stopped, and therefore the only way to prove the second line of defence still works.

7. Complexity claims are asserted as ratios, never as benchmarks

A published performance claim is a claim like any other, and this project has now watched four separate ones decay in prose. The guard for a shape claim — “linear in the macro count”, “~2× per doubling” — is a ratio between two input sizes, not a stopwatch against a threshold:

  • A ratio tests what was actually promised. The claim is a shape; a shape is a ratio.
  • The separation is wide enough to be safe. validation_cost_stays_linear_in_the_macro_count observes ~2.05× against a bound of 3.0; restoring the pre-#138 quadratic shape measures 4.01×. Take the minimum of several interleaved samples — scheduler noise only ever adds, so the fastest observation is the closest to the work actually done — and assert the smaller load was slow enough to time at all, or the ratio is meaningless.

A timing test runs alone, or it does not run. These are #[ignore]d and have their own CI step and just perf; nothing else shares the machine. The first version of this section claimed the opposite — that a ratio “survives a shared runner” because load inflates both sides and cancels — and shipped a test that failed on its second full-suite run. Measurement: 2.05× alone, 3.09× under nextest’s full parallelism. The larger input has the larger working set, so memory-bandwidth contention penalises it more; the ratio drifts rather than cancelling, and interleaving cannot fix a systematic effect. nextest’s test-groups bound concurrency within a group, which does not isolate one from the rest of the suite — so #[ignore] plus a dedicated invocation is the only mechanism that actually delivers isolation.

Benchmark frameworks were considered and are deliberately absent. iai-callgrind is the right tool for gating in CI, because instruction counts ignore runner noise entirely — but it needs valgrind, making it a gate the maintainer cannot reproduce on macOS. criterion and divan measure wall time, which is the same noise regime as the ratio test while also adding a dependency tree to a workspace that audits every edge.

8. A change that lands records itself

RELEASING.md states that every landed change adds an [Unreleased] line in the commit series that lands it. Nothing enforced that, and the rule was broken exactly once — by the series that added the guard for the changelog’s shape. A rule whose only enforcement is a sentence in another document is a rule with a known decay rate, which is the same finding this suite keeps re-deriving.

So a pull-request job asks one question: did any crates/** or xtask/** .rs file change without docs/CHANGELOG.md changing too? If so it fails, naming the files and quoting the rule. [no changelog] in the PR title waives it.

It was sized before it was written, because a gate with a high false-positive rate trains people to reach for the waiver and is then worse than nothing. Across the 21 source-touching merges preceding it the rule would have fired once — on the one commit that actually broke it. Pure-test and pure-performance changes all carried an entry already, so “source changed” tracks “worth recording” closely here. That is a measurement of this repository’s habits rather than a general law, and the waiver exists for where it stops holding.

The check reads a diff rather than files, so it is neither a docs-check task nor a tests/docs.rs test — it lives in the workflow, which is the only place the base commit is known. The PR title reaches it through env, never interpolated into the shell body: a title is attacker-controlled text, and ${{ … }} inside run: is a template injection that zizmor flags.

Runbook — thin-fork patching (ADR-0003 tier 2)

Rehearsed 2026-07-28 (M4). The steady state carries zero diff; this runbook is the mechanics for the moment a release breaks the seam or a small change is needed before upstream accepts it.

The drill (as rehearsed, local-path variant)

  1. Obtain the pinned source (fork tag in the real flow; the vendored registry copy suffices for a drill):

    cp -R ~/.cargo/registry/src/index.crates.io-*/hurl-8.0.1 /tmp/hurl-fork
    # apply the minimal patch (one commit on the fork branch in the real flow)
    
  2. Wire the override at the workspace root (Cargo.toml):

    [patch.crates-io]
    hurl = { path = "/tmp/hurl-fork" }               # drill
    # hurl = { git = "https://github.com/<org>/hurl", tag = "8.0.1-proef.1" }  # real
    
  3. Verify cargo resolves the fork, then build and run the full suite:

    cargo tree -p proef-engine-hurl | grep "hurl v"   # must show the path/git source
    cargo build -p proef-engine-hurl && cargo nextest run
    
  4. Revert: remove the [patch.crates-io] block, git checkout Cargo.lock, confirm cargo tree shows the registry source again.

Rehearsal result: override resolved, engine + suite built green against the patched path, revert restored registry resolution. Elapsed ≈ one hurl rebuild.

The real flow (when a patch is actually needed)

  1. Fork Orange-OpenSource/hurl; branch proef-patches-<version> off the release tag; apply the minimal diff (one commit per logical patch).
  2. Tag X.Y.Z-proef.N; consume via the git [patch.crates-io] form above.
  3. Open the corresponding upstream PR immediately (tier 3 — the fork’s diff must trend back to zero); note the patch in ADR-0003’s log.
  4. On each upstream release: rebase the branch onto the new tag, drop merged patches, re-run the canary, move the pins per IMPLEMENTATION-PLAN §7.

Pin-bump checklist — hurl 8.1 watch items (recorded 2026-08-16)

Behavior changes already on hurl master that the canary cannot see — compile-and-test stays green while a guarantee shifts. Work each of these when the pins move to 8.1 (IMPLEMENTATION-PLAN §7); each names the fixture to add.

The structural reason these need a list at all: the fragment scanner matches OptionKind narrowly (if let OptionKind::Variable — one arm, by design, so a semver-allowed variant addition is not a compile break). New [Options] in a fragment file are therefore silently accepted, not flagged. That is the right default for options with local effect, and exactly wrong for the first item below.

  • variables-file: (upstream #2021) — check the sandbox before accepting it. Upstream opens the named file with a raw File::open against process CWD, no ContextDir confinement. A fragment corpus is foreign by design, so a corpus file saying variables-file: ../../secrets.env would read a file outside the project on proef’s behalf. On bump: decide refuse-or-flag (option_declared_twice’s family machinery fits), and add a fixture — a .hurl with an escaping variables-file: must not silently read the target.
  • Cross-host cookie strip (upstream #5118, landed) — a redirect across hosts stops forwarding cookies. Fixture: a fixture-server redirect pair asserting which cookies arrive, so the behavior flip shows up as a diff in our suite rather than as a user’s broken auth flow.
  • --file-root resolution change (upstream #2830, watch) — multipart asset paths may move from CWD-relative to hurl-file-relative. This is proef’s multipart seam: artifacts are emitted to a different directory than the pack they came from, so relative asset paths are exactly the bytes that would change meaning. Fixture: a multipart scenario whose asset path only resolves under one of the two rules.
  • Value::Duration (upstream #3519, watch) — a captured duration re-rendered into a template may change its string form. Fixture: capture a duration-typed value, splice it into a later request, snapshot the bytes.
  • New [Options] variants generally (no-header, http2-prior-knowledge, fail-with-body, no-jsonpath-coercion, …): the scanner accepts them silently (above). On bump, sweep the new variants once and sort each into “local effect, fine” or “needs the variables-file: treatment”.

Currently drafted patches

  • docs/upstream/0001-run-entries-reusable-client.patch — run_entries accepts &mut Client (verified two-call-site change; erases per-segment connection + cookie costs, deletes proef’s SessionState cookie round-trip once adopted). Applies cleanly to 8.0.1 and compiles (verified in the M4 drill). PR text: docs/upstream/0001-PR-DESCRIPTION.md.

ADR-0001 — Embed hurl’s crates in-process as the API engine

Status: Accepted · Date: 2026-07-28

Context

The API engine must execute HTTP tests with Hurl semantics. Constraint history: the original “entirely in Rust” rule (which forbade C-linked deps and favored a pure-Rust reimplementation) was relaxed to “mostly Rust — C system libraries acceptable within margin”, and the product intent was fixed as “a wrapper / Gherkin adapter over hurl, riding upstream hurl as it improves”. Verified facts: hurl’s runner::run_entries is the seam hurl’s own parallel workers use (buffered stdio, per-entry results with spans, captures, libcurl timings); hurl links libcurl + OpenSSL + libxml2 and needs libclang at build; hurl_core alone also links libxml2; the crates break API in minor releases (issue #3846); the format itself is spec’d and stable. A working spike cross-ran generated artifacts under a prototype pure-Rust runner and the stock hurl CLI: the first cross-check caught a real semantic divergence (contains = element equality on collections, not substring) — demonstrating the standing cost of reimplementation.

Decision

proef-engine-hurl embeds the hurl crates, pinned exactly: generate .hurl text from the IR, hurl_core::parser::parse_hurl_file, then hurl::runner::run_entries with WriteMode::Buffered terms, an EventListener for progress, and VariableSet in/out. No pure-Rust HTTP reimplementation ships; no hurl subprocess is required at runtime.

Consequences

Positive: hurl semantics by construction (zero drift class); full assert/filter/XPath/ cookie/redirect/HTTP-2 surface available to packs on day one; rich in-process results with source spans mapped back to .feature lines; ~a third less engine code than the reimplementation plan. Negative: build prereqs on every machine (libcurl-dev libxml2-dev libclang pkg-config; macOS ships the libs) — mitigated by docs + proef doctor; dynamically-linked release binaries (no musl static); crate API instability — mitigated by ADR-0003; run_entries is #[doc(hidden)] (semi-blessed seam) — covered by a pinned integration test.

Alternatives considered

Pure-Rust native engine + hurl-CLI differential oracle (the original recommendation under the zero-C rule): zero C deps and static binaries, but permanent semantic-drift liability and a differential harness to maintain; recorded in research/ as the path back if static distribution ever becomes a requirement. Transpile + hurl CLI subprocess: full fidelity, least code, but an external binary dependency, no in-process events, and clunky cross-scenario variable threading; its remnant lives on as the upgrade-canary and proef artifacts hand-off. Divergent fork of hurl: rejected — upstream keeps the format low-level by philosophy (issue #2090) and a divergent fork defeats riding upstream improvements (see ADR-0003 for the thin-fork nuance).

ADR-0002 — Multi-engine core: factory/session seam, step-kind routing, batching

Status: Accepted · Date: 2026-07-28 (amended 2026-09-01 — the core’s entry grammar is a named closed set; see the Amendment below, and its 2026-09-10 correction)

Context

Requirement: all tests are Gherkin; the parser dispatches to pluggable engines — API (hurl) now, a future non-hurl engine possible later behind the same seam — a factory/session seam (multiple engine implementations behind one trait, a step’s kind-prefix routes it to its engine, one shared variable scope). Balanced-architecture stance: deliberate seams for a future engine, no gold-plating. Ecosystem survey: probe-rs’s ProbeFactory/DebugProbe split is the closest production analog; sqlx registers compiled-in drivers explicitly; dispatch cost at batch granularity is noise, so dyn vs enum is decided on coupling, not performance (enum_dispatch would couple core to every engine crate — wrong direction).

Decision

Two traits in proef-core; engines implement both; the CLI assembles the registry.

#![allow(unused)]
fn main() {
pub trait EngineFactory: Send + Sync {
    fn id(&self) -> &'static str;
    fn step_kinds(&self) -> &'static [StepKindSpec];  // pack namespace + schema fragment
    fn doctor(&self) -> Vec<DoctorCheck>;
    fn open(&self, ctx: &ScenarioCtx) -> Result<Box<dyn EngineSession>, EngineError>;
}
pub trait EngineSession: Send {
    fn run_batch(&mut self, batch: &StepBatch, world: &mut World,
                 events: &EventSink, cancel: &CancellationToken) -> BatchResult;
    fn finish(&mut self) -> Result<(), EngineError>;
}
}

Routing: a macro step’s kind names its engine (http: → engine-hurl; other kind prefixes reserved for a future non-hurl engine). A lowered scenario is an ordered heterogeneous step list; the core dispatches contiguous same-engine batches in order. The World is the interop bus between batches and engines. Sessions are per-scenario, opened lazily, torn down in finish (+ Drop backstop); engines may hold sessions concurrently within a scenario. Registry: Vec<Box<dyn EngineFactory>> in proef-cli, engines optionally behind cargo features (one feature per engine). Lifecycle is enforced by ownership shape (only a session runs batches), not typestate generics (which would break dyn).

Consequences

Adding an engine = one crate + one registry line; pack schema and doctor extend via step_kinds()/doctor() without core edits — the acceptance test: a future non-hurl engine lands with zero proef-core diff. Core stays free of engine-specific types. Costs accepted: Box<dyn> indirection (irrelevant at batch granularity); two traits instead of one (justified: lifecycle safety + capability discovery). Engines own their artifacts (hurl files / screenshots / HAR).

Alternatives considered

Single Engine trait with runtime lifecycle state (v3 draft) — weaker lifecycle guarantees; enum dispatch — inverts the dependency direction; dynamic loading (dlopen/ WASM) — rejected as over-architecture, compiled-in covers every stated future; typestate generics — fights dyn, ownership shape gives most of the safety.

Errata

2026-07-28 (M1/M5): The routing example above names the API step kind http:; ADR-0004’s examples and TECH-SPEC §6’s normative pack schema use hurl: (the raw-block key doubles as the routing kind). The implementation follows the tech spec: the step kind and the engine id are hurl, so “a step’s kind names its engine” holds verbatim. Read http: in the Decision above as hurl:. Other kind prefixes remain reserved for a future non-hurl engine as written.

Amendment — the core’s entry grammar is a named closed set

2026-09-01 · Accepted. “Core stays free of engine-specific types” is true and stays true. “Core stays free of engine-specific syntax” was never true, and the worklist carried the gap for two rounds without resolving it. This amendment states the real boundary and makes it enforceable.

Why the core knows any hurl at all

The core performs text surgery on entries: bake_entry_options splices an [Options] block into each entry after its header block, and an expect: macro merges asserts into the previous request entry (ADR-0004). Both operations have to find an entry boundary in text the engine will later parse. That is structural, not incidental — the surgery is what the pack format is built on — so a boundary recogniser has to live somewhere, and pushing it behind the seam would move the literals without making the algorithm engine-independent.

The measurement

Not the “~290 lines, all in lower.rs” the worklist recorded — that figure counted #[cfg(test)] fixtures, where a core test exercising the pipeline necessarily writes some engine’s payload. The vocabulary is thirteen distinct literals across four files:

GroupTokensWhere
written — the core generates this hurl[Options], [Asserts], HTTP *, variable:, retry:, retry-interval:, delay:lower.rs
recognised — read to find an entry boundary``` (body fence), HTTP / HTTP / HTTP/lower.rs, emit.rs, pack/validate.rs
quoted — a hurl snippet shown to an authorGET ${url:base}/PATH, HTTP 200bind.rs

The four boundary recognisers (is_method_line, is_section_header, is_response_line, is_header_line) are already one canonical pub(crate) set shared by three of those files. That half is done.

The third group is the one this measurement nearly missed, and it is worth naming why. bind.rs renders a did-you-mean help string for an author whose sentence bound no macro, and that string contains a small hurl example. It generates nothing and parses nothing, but it is engine syntax living in the core, and it drifts like any other copy. The guard’s first version could not see it — the literal spans lines, and a per-line scan discards a run that never closes — while this amendment claimed the set was closed. Multi-line literals are where a larger piece of engine syntax would naturally be written, so the blind spot sat exactly where the risk is highest. The guard now lexes whole files.

The same row cost a second correction (2026-09-02). Lexing whole files surfaced HTTP 200; the GET ${url:base}/PATH line directly above it in the same literal stayed invisible for another round, because the guard classified four shapes — fence, response line, section header, option line — and a method line was not among them, though this section names it as one of the four recognisers. A guard is closed only over the shapes it can classify, so the two claims have to be checked against each other rather than assumed to agree. The classifier now knows method lines, which is what added the row above. In the same pass the scan stopped truncating at a file’s first #[cfg(test)] mod and began excising every test module instead: production code placed after one was silently unscanned, and html.rs and pack/validate.rs already carry a second test module.

proef’s own pack keys (macros:, match:, secret:, steps:, use:) are shaped like option lines and are excluded by name rather than listed as sanctioned rows: an inventory that is a third exceptions stops reading as a closed set.

The asymmetry this exposes

StepKindSpec::options exists, in its own words, as “the seam that keeps option spellings out of proef-core” — added because matching "retry-interval:" as a literal meant “one rule lived at two altitudes.” It covers recognising options. The core still writes retry:, retry-interval:, delay: and variable: as literals, so the same rule still lives at two altitudes, in the other direction.

Correction (2026-09-10) — the set was fourteen, and the fourteenth was unclassifiable

The measurement above says thirteen literals across four files. It was fourteen. The one it missed is "file," in emit.rs, where file_refs_in found the assets an artifact reads by scanning for that literal and a closing ; — hurl’s body grammar, in proef-core, for the entire life of asset staging.

It went unrecorded for the same reason the method line did, and this section had already written the rule that predicts it: a guard is closed only over the shapes it can classify. engine_grammar_kind knew fences, HTTP, [Section] headers, method lines and key: value options. A body constructor is none of those — no colon, no brackets, no uppercase — so the literal was never classified, never entered the inventory, and was never reported missing from it. The set was not measured and found closed; it was measured through a classifier that could not see this member.

That is the third decay of this section’s own claim: once by an order of magnitude in the count, once by a multi-line literal, and now by a shape. Each time the count was wrong in the direction of the guard’s blind spot, which is the only direction it can be wrong in.

Resolved by the second remedy, not the first. Decision 2 below sends the author to one of two options: widen the sanctioned set on the record, or put the syntax behind the seam. This is the first time the second was taken. The scan is now StepKindSpec::assets, a fourth engine-contributed hook beside validate, fragments and options — so the recognised group loses its body-reference member entirely rather than gaining a sanctioned row.

Moving it also fixed the reading. A text scan cannot tell a real file,…; body from the same six characters inside a JSON or assertion body; the engine reads its own AST and can. It also has to avoid hurl’s shared visit_filename hook, which carries the [Options] file paths (output, cacert, client-cert, client-key, netrc-file, unix-socket) alongside real bodies — output: names a file the run writes, and staging it would demand a source that cannot exist. Only the two body positions are read. None of that distinction is expressible in core, which is the argument for the seam stated as a capability rather than as a rule.

The classifier gained a body arm in the same change, so the blind spot is closed independently of the literal that exposed it: a bare lowercase keyword followed by a comma (file,, hex,, base64,) is now classified, and reintroducing one into core fails the guard with body "file," in emit.rs.

Decision

  1. The set above is the sanctioned core entry grammar. It is closed: a token outside it, or an existing token appearing in another core module, is a defect against this ADR.
  2. It is pinned by crates/proef-cli/tests/source_guards.rs (hurl_grammar_in_core_is_the_closed_set_the_adr_names), which lexes every production literal in proef-core and fails on growth, on relocation to another core module, and on shrinkage — then sends the author back here. A claim of this shape decays the moment it is only prose. This one already had, twice: once by an order of magnitude in the count, and once in this very section, which asserted a closed set while the guard behind it could not read a multi-line literal.
  3. Migrating the written group behind the seam (an emitter beside StepKindSpec::options) is deferred, not rejected. It buys nothing today: hurl is the only engine and no other is scheduled, so the migration would add a fn pointer, a trait obligation and a public-API break to relocate seven literals that exactly one implementation will ever supply. Trigger: a second engine being scheduled. That is also when ADR-0002’s acceptance test — a new engine lands with zero proef-core diff — first has anything to say about them; until then it is unfalsifiable here either way.

Consequences

The acceptance test is narrowed on the record: a second engine lands with zero proef-core diff except the written group, which is a known, enumerated, guarded debt with a named trigger rather than an open question. Anyone reaching for new hurl syntax in the core hits a failing test that names both remedies.

ADR-0003 — Upstream tracking: exact pins, thin zero-diff fork, upgrade canary

Status: Accepted · Date: 2026-07-28

Context

Requirement (stated): “use a hurl fork with as few changes as possible — I want to keep using hurl as it improves over time; a wrapper/Gherkin adapter over hurl.” Verified: hurl ships breaking crate-API changes in minor releases (#3846; maintainers advise cargo install --locked); the release cadence is multiple per year; run_entries is #[doc(hidden)]. One concrete patch need is already identified (ADR-0010 / TECH-SPEC §5): run_entries creates its HTTP client internally, so accepting &mut Client would erase per-segment connection costs — a verified two-call-site change.

Decision

Three-tier policy, in order of preference. (1) Steady state: depend on published crates with exact pins (hurl = "=8.0.1", hurl_core = "=8.0.1"), build --locked; the GitHub fork exists but carries zero diff. (2) Patch vehicle: when a release breaks the seam or a small change is needed, carry a minimal-diff branch on the fork, consumed via Cargo [patch."crates-io"] (or a git-tag dep), rebased onto each upstream release. (3) Upstream everything: every patch is PR’d upstream so the fork’s diff trends back to zero. An upgrade-canary CI job (weekly + on upstream release) builds against the next hurl version and replays the full suite; pins move only after it is green, via the runbook in IMPLEMENTATION-PLAN §7.

Consequences

“Keep using hurl as it improves” becomes a scheduled chore, not a gamble; the wrapper never becomes a divergent fork; breakage is discovered pre-pin-bump. Costs: upgrade PRs are deliberate work per release; the fork must be rebased when (and only when) it carries a patch; MSRV follows upstream (hurl master already 1.97.1 — neutralized by the project’s always-latest-stable toolchain rule).

Alternatives considered

Track master via git dependency — unvetted breakage flows in continuously. Vendor the hurl source into the repo — a divergent fork in disguise; loses provenance and cadence. Caret/tilde version ranges — semver is demonstrably not honored for the library surface; exact pins are the only safe mode.

ADR-0004 — Pack format: YAML skeleton + embedded raw Hurl blocks

Status: Accepted · Date: 2026-07-28

Context

Requirement (stated): macro packs must be human-readable — “is there a better alternative than macro YAML?” Analysis (architecture review §8) evaluated YAML+schema, KDL, TOML, Pkl/CUE/Dhall, RON/JSON5, Rhai/Lua scripting, a custom DSL, and Karate-style Gherkin-native macros against readability/writability, comments, multiline bodies, templating interplay, editor tooling, serde support, and team familiarity (existing packs are YAML with schemars-driven autocomplete). Key insight: the unreadable part of packs was never YAML itself — it was HTTP-as-YAML-trees, while hurl’s own plaintext format is the human-readable HTTP DSL, and the backend team already reads/writes it fluently.

Decision

Packs stay YAML (serde_norway; schemars JSON Schema; comments; block scalars) but only as the thin binding skeleton: macro name, match: pattern, params, defaults, tags, description, composition (use:/with:), step modifiers (optional:, when:, retry:, saveAs:). The HTTP payload of a hurl step is a raw Hurl block:

steps:
  - name: Resolve the record name to its id
    hurl: |
      GET ${url:base}/api/v1/admin/search/records
      Authorization: Bearer ${secret:apiToken}
      [Query]
      q: ${name}
      HTTP 200
      [Captures]
      recordId: jsonpath "$[0].id"

Blocks are validated at pack load by parse_hurl_file after ${…} lowering — real hurl syntax errors with real spans. Structured step trees are reserved for a future non-hurl engine, which would have no native text DSL. Assert-only macros use expect: (merged into the previous request entry — the Then-step rule).

Consequences

Pack bodies are literally hurl: copy-paste flows both ways with the backend corpus; no bespoke assert/capture schema to maintain for the API engine; the emitter for hurl steps approaches the identity function. Costs: autocomplete inside the block is plain-text (mitigated: load-time parse errors are immediate; editors have hurl highlighting; a proef fmt pass can normalize blocks); one lowering pass must run before parse (already required for ${…}).

Alternatives considered

KDL — pleasant syntax but no schema/LSP story comparable to YAML, zero team familiarity; recorded as the fallback if YAML friction materializes. TOML — wrong shape for nested step lists (kept for proef.toml config). Pkl/CUE/Dhall — second language + toolchain, over-architecture at this size. Rhai/Lua — packs become programs; kills static validation and --dry-run guarantees. Custom DSL — a parser/LSP/formatter to own forever. Karate-style callable feature files — collapses the macro/test distinction; a typed params/defaults/validation model is strictly stronger.

Amendment (2026-07-30): the top-level key is macros:

The pack root key was renamed templates: → macros: to end a three-way naming split (the YAML key said templates, the docs and internal model said macro, the file/dir said pack). The entry is now uniformly a macro; a pack is a file of macros. Pure rename — format, schema, and semantics are unchanged (error-corpus snapshots regenerated, diff verified as templates:→macros: only). No templates: alias is kept — one canonical spelling (golden rule: one way to do one thing).

ADR-0005 — Two-tier variables (${…} / {{…}}), the World, and secrets

Status: Accepted · Date: 2026-07-28

Context

The author-time variable system (${var}, ${env:NAME:-default}, ${run:id}, ${global:key}, ${fake:*} seeded from the run id, $${} escape, recursive expansion) is proven with test authors. Hurl has its own runtime templating ({{name}}) fed by captures. The spike validated running both tiers side by side — including the bug it surfaced: step-captured args can themselves contain ${…} and need recursive (depth-capped) resolution.

Decision

Two explicit tiers. ${…} is author time: params, env (with defaults), run id, World reads (${global:key}), fake data, secrets references — resolved during lowering, recursively with a depth cap of 8, and baked into artifacts (except secrets). {{…}} is run time: hurl-native templates for captures, left verbatim in artifacts so the embedded engine and the stock CLI resolve them identically. World: one typed variable scope per scenario plus a persistent global store (.proef-state.json, atomic temp+rename). Engine bridging: seed hurl’s VariableSet from the World before each batch; merge HurlResult.variables back after; saveAs: global promotes a capture into the persistent store. Secrets: ${secret:NAME} resolves from the PROEF_SECRET_<NAME> environment override, else the encrypted store .proef-secrets.json (chacha20poly1305 + rpassword); values are injected via VariableSet::insert_secret (hurl redacts them in logs/reports); artifacts carry {{secret_name}} placeholders, never values; our reporters additionally redact by value (property-tested invariant).

Amendment (2026-08-16): “redact by value” includes each secret’s common encoded forms — base64 (both alphabets, with/without padding), hex (both cases), RFC 3986 percent-encoding, and the JSON-string escape — derived inside Redactions::new so every sink is covered by construction. Demonstrated live before the amendment: a server reflecting a bearer token base64-encoded put a trivially-decodable string into an assert-failure detail, the raw needle never fired, and the encoded credential reached the console and events.jsonl. The needle set covers the reversible transforms that occur at HTTP boundaries; a secret reflected hashed or re-encrypted matches no needle list, and the ADR does not claim otherwise. Over-redaction is the accepted failure direction.

Consequences

Artifacts are runnable by both toolchains with identical meaning; authors keep a familiar mental model unchanged; secrets are structurally absent from every persisted output. Cost: two syntaxes coexist in packs — mitigated by the strict rule of thumb (“$ = before the run, {{ = during the run”) documented in the pack authoring guide.

Alternatives considered

Single-tier (resolve everything at author time) — breaks capture chaining and makes artifacts non-parametric. Single-tier (everything hurl {{}}) — loses env defaults, fakes, and World reads. String-only World — kept typed here (hurl Value model) because captures cross engines; stringly-typed round-trips would lose numbers/bools at engine boundaries.

Errata

2026-09-07 (0.18): the invariant’s reach was found unenforced on one of its two paths. The event stream masks through Redactions::apply_event, an exhaustive destructure; the CI sinks that render from RunSummary (JUnit, CTRF, TAP, timings.json, the GitHub summary and annotations) masked failure detail, but five of them bypassed the masker for the identity fields (scenario, file, tags, the skip reason) — no live leak, since secrets lower to {{name}} and the engine pre-redacts details, but a boundary held by convention. Closed per sink (#171), then made structural (#178): Redactions::apply_outcome destructures ScenarioOutcome/StepOutcome without .., so a new text field fails to compile until it is masked, and each sink redacts one outcome after matching @quarantine on the raw identity. Per-sink rather than a wholesale apply_summary, because RunSummary also feeds exit_code_excluding, whose quarantine matching needs the unredacted identity.

2026-07-29 (v0.3.1): the “secrets reach no sink” invariant now explicitly covers the persistent World: a saveAs: global capture whose value equals a known secret is refused (the owning step warns) — .proef-state.json is plaintext at rest and must never receive secret-derived material. PROEF_KEY (base64) may supply the project key via the environment for CI use of a committed ciphertext store.

2026-07-28 (post-M5 hardening; API removed 2026-07-29 per YAGNI): “snapshot/restore across scenario retries” originally described a mechanism whose trigger was never specified anywhere in the corpus — no CLI flag, tag, or pack directive schedules a scenario-level retry (US-5’s step-level retry: is implemented and is the flake tool in practice). Decision: scenario-level retries are deferred indefinitely, and the unused GlobalStore::snapshot/restore API has been removed (YAGNI: no dead promised-behavior code). Whoever implements scenario retries specifies the trigger surface and its mechanism in a superseding ADR. Also note: the scenario merge-back is write-set-only (World tracks its saveAs promotions) — merging a whole snapshot back would lose concurrent scenarios’ updates.

ADR-0006 — Engine traits are sync + dyn; no async machinery in v1

Status: Accepted · Date: 2026-07-28

Context

All planned engines are blocking at the edge: hurl is synchronous libcurl, and a future non-hurl engine’s driver can be driven blocking — even one that also offers an async API. Verified ecosystem facts (mid-2026, stable 1.97): async fn in traits is stable for static dispatch but still not dyn-compatible (AFIDT is nightly-only, no timeline; RTN unstable); maybe-async’s sync/async toggle is a non-additive cargo feature with documented ecosystem breakage. Calls across the seam are coarse (batch-level), so async buys no throughput inside the runner itself.

Decision

EngineFactory and EngineSession are synchronous traits, used as Box<dyn …>. Parallelism is scenario-per-OS-thread. If a future host (server mode, desktop UI) is async, it wraps engine calls in spawn_blocking at its edge; the core never learns about executors (reinforced by the sans-IO-lite rule: core does no IO at all). maybe-async is explicitly banned. Traits meant to be implemented externally (Engine*) are not sealed; internal traits that must stay evolvable are sealed.

Consequences

Simple, dyn-compatible seam today; no async runtime in the dependency tree (tokio-util’s CancellationToken is runtime-independent — ADR-0007); a future async migration is additive (an AsyncEngineSession adapter or AFIDT adoption when stable), not a rewrite. Cost: a long-running async host pays one thread per in-flight scenario — acceptable at e2e-suite scale.

Alternatives considered

Async-first trait via async-trait — boxes every call, forces an executor decision on every consumer, and models nothing real while engines block. maybe-async dual API — non-additive feature hazard. Callback/actor-per-engine threading model — more machinery than batch dispatch needs; revisit only if an engine genuinely multiplexes (e.g. a driver with concurrent event streams), and then inside that engine crate, invisible to the seam.

ADR-0007 — Cancellation: cooperative at batch boundaries, with budgets

Status: Accepted · Date: 2026-07-28

Context

Source-verified: hurl has no cancellation mechanism anywhere — no signal handling in the workspace, no abort check in the entry loop; delay/retry-interval are uninterruptible thread::sleeps; retry/repeat accept Count::Infinite; the only bounds are libcurl’s per-request timeouts (default 300 s), which do not bound total entry time under retries. The standard here is structural cancellation (CancellationToken threaded everywhere). Verified: tokio_util::sync::CancellationToken is runtime-agnostic (tokio sync primitives are documented runtime-independent; default-features = false — no tokio runtime enters the tree); is_cancelled() polling and child_token() work from plain threads.

Decision

Cancellation is cooperative at batch boundaries: the orchestrator checks a per-run CancellationToken (child token per scenario) before opening sessions and before each batch; EngineSession::run_batch receives the token so engines may honor it at finer grain when they can (a future non-hurl engine might; engine-hurl cannot mid-run_entries). Stuck-batch policy, layered: (1) the pack lint rejects infinite retries (retry: must carry a finite count) and unbounded repeat; (2) engine-hurl clamps per-request timeouts and computes a batch budget = Σ(entry timeout × (retries+1)) + retry intervals + margin; (3) a watchdog marks a scenario thread abandoned when its budget expires — the runner records a System failure with full context and detaches the thread (process exit reaps it) rather than blocking the run on an unjoinable thread. Ctrl-C: first signal cancels the token (graceful: finish current batches, run teardowns, write reports); second signal hard-exits.

Consequences

Bounded, explainable runs; no dependency on hurl gaining cancellation; a clean seam for engines that can do better. Costs: a cancel can wait out one in-flight batch (bounded by its budget); abandoned threads leak until process exit (accepted: the process is short-lived by design). If interrupt support ever lands upstream, adopting it is an engine-internal change (candidate for the ADR-0003 patch pipeline).

One consequence reaches a later feature, so it is recorded here rather than discovered again. [run] exclusive-tags promises a matching scenario the pool to itself, and the dispatcher enforces that against its own active set. An abandoned scenario leaves that set the moment the watchdog fires, while its detached thread keeps issuing requests until the entry in flight returns — the whole reason abandonment exists. So an exclusive scenario can start while an abandoned neighbour is still talking to the target: exactly the interference the key promises away, in the one window this ADR knowingly leaves open. It takes a budget blowout in the same moment, and a run that never trips the watchdog never meets it, so this is documented rather than engineered away — closing it means a bounded grace on the exclusive fill gate while detached threads report alive, which buys a rare guarantee with a per-exclusive-scenario delay every run pays. Revisit if the isolation guarantee ever has to be absolute.

Amendment (2026-09-07) — the budget family is closed over its inputs, and bounded as a product

The 0.18 survey found three holes in the value-cap regime, one of them the exact shape this ADR exists to prevent:

  • max-time: was read by the budget calculator and invisible to the lint. entry_timeout has always taken a literal max-time: as the entry’s timeout, while the option recogniser did not know the key — so [Options] max-time: 100000h was lint-clean and produced a multi-year batch budget the watchdog dutifully honoured. It now carries the duration cap like every budget input, and a test pins the rule the hole broke: every option the budget reads must be one the lint can see (every_budget_input_carries_a_value_rule).
  • retry-interval: multiplied into the budget with no value cap. It was in the recogniser for double-declaration purposes only; it now carries the duration cap.
  • Individually capped values compose into an unbounded product. retry: 10_000 (at the count cap) times a 30 s timeout is ~83 hours, lint-clean; saturated arithmetic reaches Duration::MAX, whose Instant + budget addition panics — contained by the dispatcher’s catch_unwind, but reported as a phantom “scenario thread panicked” system fault from a user-authored value. Two closures: the computed batch budget clamps to an absolute ceiling of four hours (MAX_BATCH_BUDGET — generous for any batch of API calls with finite retries, and a truthful watchdog abandonment for a runaway product), and the dispatcher’s deadline arithmetic uses checked_add with a far-future fallback, so an engine that ever hands core an unclamped budget degrades to “no deadline” rather than a panic.

In the same family: [http] timeout-ms = 0 was accepted and means no timeout to libcurl — the exact unbounded hang the default defends against, opted into by a value that reads like “immediately”. Refused as a user error now.

Alternatives considered

Killing scenario threads — unsound in Rust (no safe thread kill). Running each batch in a subprocess for killability — reintroduces the subprocess architecture ADR-0001 rejected, per-batch. Relying on timeouts alone — unbounded under retry loops (verified), and no graceful-report path on Ctrl-C.

ADR-0008 — Serde event spine, decorator reporters, libtest-mimic harness

Status: Accepted · Date: 2026-07-28

Context

a live-event seam (EventSink(Arc<dyn Fn(RunEvent)>), domain events decoupled from wire format) is proven; its run record is a separate reporter path. The two best in-domain designs both converge on “typed event stream + composable consumers”: cargo-nextest (runner emits structured events; reporter fans out to human/JUnit/machine outputs) and cucumber-rs (decorator Writer stack: Normalize → Summarize → leaves, with Tee, marker traits for ordering guarantees). Verified: nextest officially supports libtest-mimic custom harnesses (documented CLI contract; libtest-mimic 0.8.x); libtest’s JSON format itself is still unstable nightly territory; OTel test semconv is “development” maturity.

Decision

One serde-able event enum in proef-core is the spine: RunStarted, ScenarioStarted, BatchStarted, StepFinished (engine id, StepRef feature/line, status, attempts, duration, capture names), ScenarioFinished, RunFinished. The JSONL run record is the appended event stream — explain, history, and any future UI replay from disk; no second record format. Reporters are event consumers composed decorator-style: Normalize (repairs interleaving from parallel scenarios) → Summarize → leaves: console BDD tree, JUnit XML (via quick-junit), GitHub job summary, JSONL appender. Secret values never enter events (capture names only — redaction invariant, ADR-0005). M5: a libtest-mimic harness binary exposes one Trial per scenario, making cargo nextest run and IDE test UIs drive proef with zero custom protocol work; libtest JSON remains an output adapter, never the native schema. OTel export: deferred; if added, a thin optional reporter mapping to test.* semconv names.

Consequences

Single source of truth for live progress and persistence; reporters are ~a page each; new outputs are additive leaves. Replayability makes run records diffable and testable (insta snapshots over event streams). Cost: event schema becomes a compatibility surface — versioned with a schema field from day one.

Alternatives considered

a split design (live events + separate JSON record) — two sources of truth to keep consistent. tracing as the event bus — wrong tool: tracing is operator telemetry, not a typed result stream (kept for diagnostics). Cucumber’s writers verbatim — async trait

  • World coupling we don’t need; the decorator shape is what’s adopted.

Errata

2026-07-29: the variant set has grown additively since acceptance: EntryRunning (live per-attempt engine progress) joined the six original variants, RunFinished gained a cancelled flag, and StepFinished gained a detail failure field — all serialized only when present, so pre-existing streams parse unchanged. The additive-only rule held; this note keeps the variant inventory honest.

2026-08-24 (RF wave 2): scenario_finished gained reason (why a scenario is skipped — authored spellings start with @, mechanical prose never does; ADR-0019) and tags (the accumulated tag set, finished-event only because the cancel-skip path emits no start); scenario_started gained exclusive (the scheduler’s own bool, R11-6); step_finished gained reproduce_hint (the failing request’s redacted curl). All serialized only when present; EVENT_SCHEMA_VERSION stays 1. run_started further gained env/metadata/shuffled (ADR-0020) and rerun_of (the E2 rerun overlay) — same additive discipline.

2026-09-06 (0.18): the record is now protected the way the console already was. events.jsonl was handed a bare File, so a disk filling mid-run truncated the record while the run exited by its verdict; the record’s writer now latches its first failure and the exit funnel turns it into a system error (exit 3) through the same fold as the JUnit/CTRF and GitHub-summary write failures (escalate_environment_failures). A single SIGTERM/SIGHUP is a cancellation — the record closes normally with run_finished + cancelled — so only a second signal, a SIGKILL, or a crash leaves a truncated record (EVENTS.md). The sidecars that sit beside the record (timings.json, inputs.json) are derived aids, never a second record: the stream stays the only persisted format, and EVENT_SCHEMA_VERSION stays 1.

ADR-0009 — Error taxonomy by fault, stable exit codes, miette at the edge

Status: Accepted · Date: 2026-07-28

Context

The error model categorizes by who is at fault — User / TestFailure / System(anyhow) — with a total mapping to its stable exit-code scheme (0 ok · 1 test failure · 2 user error · 3 system error), integration-tested via assert_cmd. 2026 ecosystem consensus (blessed.rs et al.): thiserror in libraries, anyhow at the application edge; miette’s Diagnostic adds codes/help/labeled source spans for user-facing errors. Survey of backend traits (tower/sqlx/rustls/probe-rs): behind dyn, a unified error enum with boxed sources beats associated type Error. gherkin 0.16 spans are byte offsets (verified), directly convertible to miette SourceSpan (with an EOF-trailing-newline clamp; LineCol.column is char-counted — never mixed into byte math). snafu/error-stack: adopted by some large codebases, unnecessary at this crate count.

Decision

A fault-category model, extended for engines: proef-core defines

#![allow(unused)]
fn main() {
pub enum CoreError { User(..), TestFailure(..), System(..) }        // → exit 2 / 1 / 3
pub struct EngineError { pub class: EngineErrorClass,               // Infra | AssertFailed | Setup
                         pub message: String,
                         pub source: Option<Box<dyn Error + Send + Sync>> }
}

AssertFailed folds into TestFailure; Infra/Setup into System. thiserror 2 everywhere; no anyhow in library crates (only inside System’s boxed source at the edge). miette lives only in proef-cli: parse/bind/validation errors wrap into Diagnostics with labeled spans into .feature files (gherkin byte spans) and pack YAML (serde_norway locations); engine failures render the feature line + artifact span from the sidecar. Exit codes are a typed enum, pinned by CLI integration tests.

Consequences

Every failure has a fault category, a stable exit code, and a source-located rendering; engine crates stay miette-free (usable headless); the explain command reuses the same classification. Cost: two error layers (core vs engine) — justified by the seam: engines can’t know exit codes, core can’t know engine internals.

Alternatives considered

Associated type Error on engine traits — erased behind dyn anyway. anyhow everywhere — loses matchable categories that exit codes require. snafu — per-crate context ergonomics we don’t yet need; revisit if the workspace grows past ~10 crates.

Amendment — 2026-08-04 (exit code 130 documented, not a variant)

test and watch hard-exit with code 130 (128+SIGINT) on a second Ctrl-C while a run is already cancelling (ADR-0007) — the shell’s own convention for a signal-terminated process. This is a sanctioned OS-signal escape hatch, not a graceful outcome the fault-category model classifies, so it is intentionally not an ExitCode variant; ExitCode stays the total 0/1/2/3 mapping above.

Amendment — 2026-09-06 (every signal, and undelivered output)

Ctrl-C, SIGTERM and SIGHUP all take the graceful cancel (ctrlc’s termination feature), so a CI job timeout or docker stop is a cancelled run — exit 1 with a complete record — not a kill; the 130 hard exit above fires on a second signal of any of the three, and the handler carries no signal identity, so one code covers them all. And output proef could not deliver never looks like success: a failed write of the run record, of a JUnit or CTRF file, or of the GitHub step summary re-classifies the exit to System (3) through one fold, escalate_environment_failures, beside the stdout latch this taxonomy already covered.

ADR-0010 — Artifacts as contract: emitted .hurl is the executed input

Status: Accepted · Date: 2026-07-28

Context

The backend team’s hurl corpus makes .hurl the interop format. The spike proved generated artifacts run identically under the prototype engine and the stock CLI — and that keeping two implementations honest requires differential testing. Embedding hurl (ADR-0001) enables something stronger: the artifact and the executed input can be the same bytes. Verified seam facts that shape the mechanics: run_entries creates its HTTP client per call (fresh connections; cookie jar seedable only via Netscape-format file; variables chain losslessly via HurlResult.variables); per-entry [Options] override batch-level RunnerOptions defaults (clone-then-override, verified).

Decision

For every scenario, the emitter produces canonical .hurl text; that exact text is what parse_hurl_file + run_entries execute — drift between artifact and execution is structurally impossible. Alongside each artifact: a sidecar map (<slug>.map.json: entry ↔ feature file/line/step text, optional flags, capture names, batch boundaries) and, when World/global values are referenced, a generated <slug>.vars file so the backend team can replay with hurl --variables-file. optional: entries carry an # optional marker comment (no hurl equivalent — the runner segments around them). Execution batches maximally: one run_entries call per scenario unless optional: boundaries or interleaved other-engine steps force a split; when a split occurs, variables chain via HurlResult.variables and cookies (if used) round-trip via a Netscape temp file behind a SessionState struct. Queued upstream patch #1 (per ADR-0003): run_entries accepting &mut http::Client — verified two-call-site change — which erases per-segment connection/cookie costs entirely. Artifacts never contain secret values (ADR-0005). Artifact output: per-run under .proef-runs/<id>/artifacts/ plus proef artifacts for a stable CI hand-off directory.

Consequences

Debugging = opening the artifact with tools the team already knows; the backend corpus and proef packs stay mutually copy-paste-able; --dry-run validation includes parsing the real artifact with the real parser. Cost: canonical formatting is a compatibility surface (snapshot-tested); segmented scenarios pay reconnection costs until patch #1 lands upstream.

Alternatives considered

Internal-only IR execution with optional export — loses the same-bytes guarantee and demotes artifacts to lossy exports. Differential-oracle architecture (v1 plan) — superseded by ADR-0001; preserved in research/ as the static-distribution fallback.

ADR-0011 — Fixture server is synchronous tiny_http, not axum

Status: Accepted · Date: 2026-07-28 (decision made at M3; recorded as an ADR during the post-M5 hardening pass — it had been noted only as inline errata in TECH-SPEC §14 and TESTING-STRATEGY §2)

Context

TESTING-STRATEGY as originally written named axum for the integration fixture server (proef-fixture). Axum requires a tokio runtime. The workspace bans the tokio runtime outright (ADR-0006/ADR-0007: engines and core are sync; only tokio-util with default-features = false enters the tree, for CancellationToken). A dev-dependency would not leak into shipped binaries, but it would put a full multithreaded async runtime into every cargo nextest process, contradict the “no async machinery” line we enforce with cargo deny bans, and normalize exactly the dependency the ADRs exclude.

The fixture’s needs are modest: a handful of JSON endpoints, per-env state isolation, deterministic token-driven delayed visibility, cookies, one deliberately slow route and one deliberately malformed one — all exercised by at most a dozen concurrent scenarios.

Decision

proef-fixture is built on tiny_http (synchronous, dependency-light) with a plain std::thread accept loop and an Arc<Mutex<_>> state map keyed by the X-Proef-Env header. It starts on an ephemeral port (Server::http("127.0.0.1:0")), reports its base URL, and shuts down via an AtomicBool + recv_timeout poll. Portability note: address introspection uses ListenAddr::to_ip() — the Unix variant of ListenAddr exists only on unix targets, so matching on it breaks the Windows build.

Consequences

  • The workspace stays runtime-free end to end; the deny-list ban on tokio stands without a dev-dependency exception.
  • The fixture is one file, debuggable with a thread dump, and starts in microseconds — each integration test spawns its own isolated instance.
  • No HTTP/2, no TLS, no streaming in the fixture — acceptable: the engine’s HTTP behavior is hurl/libcurl’s concern, not the fixture’s; the fixture only scripts responses.

Alternatives considered

  • axum (as originally specced): rejected — drags in the banned tokio runtime; every capability the fixture needs is available synchronously.
  • hyper in blocking mode / raw std::net: more code for no additional fidelity.
  • Out-of-process fixture binary: slower startup, port coordination, and a lifetime-management problem tests would have to solve; in-process tiny_http gives free isolation per test.

Amendment — 2026-07-31 (dev-loop CLI binds the advertised default port)

Fixture::start() stays ephemeral (127.0.0.1:0) — the integration suite spawns a dozen concurrent instances and each needs its own port. But the shipped proef.toml advertises base = http://127.0.0.1:8787, so a first-time cargo run -p xtask -- fixture on a random port left the default unreachable and forced a PROEF_BASE_URL export. The dev-loop CLI (xtask fixture, one instance at a time) now calls the new Fixture::start_on(port) to bind 8787 by default (override: ... -- fixture <port>), falling back to an ephemeral port — with the PROEF_BASE_URL line printed — only when 8787 is busy. The library API and the per-test isolation described above are unchanged; only the human entry point picks a stable, documented port.

ADR-0012 — Project configuration & environments in proef.toml

Status: Accepted · Date: 2026-07-30

Context

Test files must stay pure prose — no URLs, no environment data, no variable definitions (operator requirement: “test files for testing, not for variable definitions”). Before this, the only non-secret variable source was the in-feature # baseURL: directive plus ${env:…}, which put configuration inside the .feature files. Suites also had to be given an explicit path on every proef test invocation.

Decision

proef.toml (already the config file — ADR-0004 kept TOML for it) gains variable-bearing sections and per-environment override profiles, modeled on the Cloudflare-Wrangler wrangler.toml [env.<name>] pattern (and Cargo [profile.*]):

  • [url] and [vars] — non-secret variables, referenced in packs as ${url:<key>} and ${vars:<key>} (new lower-time resolver namespaces, within the ADR-0005 ${…} tier).
  • [env.<name>.<section>] — per-environment overrides that deep-merge over the base tables, key by key: [env.prod.url] / .vars override variables, .http / .run override runner settings. Unlisted keys inherit the base.
  • --env <name> / PROEF_ENV selects the active environment.
  • [run] suite — the default suite path, so proef test needs no path argument; the tests/ directory is the zero-config fallback convention.

Core stays sans-IO: the CLI loads proef.toml, validates the --env name, deep-merges, and injects the resolved scope as LowerCtx::config_vars (keyed "<namespace>:<key>"). The resolver reads that injected map — it never touches a file. Secrets remain on their own encrypted channel (${secret:…}) and never appear in proef.toml.

Consequences

  • Feature files become pure prose; URLs / credentials / env data live in one external file — the single variable-definition mechanism (the legacy # key: directive was later removed; see the 2026-07-31 amendment).
  • A referenced-but-undefined ${url:…} / ${vars:…} is a user error at lower time (proef::resolve::missing_config_var) — the same strictness as ${env:…}.
  • Deep-merge (not Wrangler’s non-inheritable vars) means an environment lists only deltas; this is the deliberate divergence from Wrangler’s known footgun.
  • proef test / flows / artifacts accept an optional path plus a --env flag.

Alternatives considered

  • Per-environment files (proef.<env>.toml, Spring / dotenv style) — cleaner git diffs per env, but more files; the single-file [env.<name>] model keeps one mental model and reuses the existing loader. Recorded as the fallback if one env grows large.
  • Wrangler-exact non-inheritable vars — rejected: forces re-listing every var per env (the documented Wrangler footgun); deep-merge is the more ergonomic default.
  • A ${baseURL} magic bare name — rejected: namespaced ${url:base} is collision-free and consistent with ${secret:…} / ${env:…} (one way to do one thing).

Amendment (2026-07-31) — the # key: directive mechanism is removed

The original decision kept the in-feature # key: value directive (e.g. # baseURL:) working “for one-off per-file overrides.” That left two ways to define a variable — a directive inside a .feature file, and [url]/[vars] in proef.toml — violating one-way-to-do-one-thing (operator: “variables should be defined in 1 way, not multiple ways”). The directive mechanism is therefore removed:

  • FeatureFile::directives, collect_directives, lower::resolve_directives, and the ResolveCtx::directives scope are deleted. ${…} plain-name resolution is now args > defaults only; feature files carry no variable definitions.
  • Config becomes the single variable source. Config values may themselves embed ${env:NAME:-default} (resolved recursively), so the env-override + default that # baseURL: ${env:PROEF_BASE_URL:-…} provided is preserved as base = "${env:PROEF_BASE_URL:-…}" under [url].
  • proef.toml is now discovered by walking up from the working directory (like cargo/git), so config is found from any subdirectory (needed once the value moved out of the self-contained feature file — e.g. the libtest-mimic harness invokes proef from its own crate dir).

# comment lines before Feature: remain valid gherkin comments; they are simply no longer parsed as directives.

ADR-0013 — Typed macro parameters

Status: Proposed — recommendation: defer (the shape below is recorded for when a real need appears) · Date: 2026-08-02

Context

A macro declares params and binds {capture} values from prose, data-table rows, defaults, and use:…with:. Every arg is an untyped string today: the matcher only checks that a capture names a declared param (matcher.rs), never that the value looks like what the step expects. A mistyped value (the record abc is fetched where a UUID was meant) is caught only at hurl run time — if at all — not by --dry-run.

Cucumber Expressions solve this with typed parameters ({int}, {uuid}, custom types) validated before execution. proef wants the same shift-left check, expressed within its architecture. Round-2 code validation (IMPROVEMENT-PLAN §12, item N1) established three hard constraints:

  1. Placement must be the declaration site, not the pattern. Three of the four arg sources are not captures (data-table rows, defaults, with:), and use:-only macros have no pattern at all — so an inline {name:type} form could not type them, and a type buried in the match: string is invisible to proef schema (schemars).
  2. The two-tier variable rule caps the value. An arg may legitimately be ${…} / {{…}} that only resolves at lower/run time (ADR-0005). Type-checking such an arg at bind time would false-positive, so the check must skip any raw arg containing ${ or {{ — it is a best-effort lint over literal args, not a runtime type system.
  3. One-canonical-way forces a single params spelling. params is a YAML sequence today ([q, index]); adding a typed spelling alongside it would be two ways to declare params — the golden-rule violation that killed the templates: alias and the # key: directive.

Decision

  • params becomes a name → type map: params: {q: uuid, index: int, note: any}. A missing/any type means “declared, unchecked” (today’s behaviour). Parsed with a custom Deserialize (not an untagged enum — CLAUDE.md bans those near the arbitrary_precision footgun). This is a breaking pack-shape change — every existing pack migrates params: [q, index] → params: {q: any, index: any}, with no alias (consistent with the templates:→macros: and # key: removals).
  • The type set is a closed registry — any, int, number, uuid, email, iso8601, word — modelled on the fake::GENERATORS registry (fake.rs), validated at load so an unknown type is a pack error. No user-defined types in v1 (revisit if a real need appears).
  • Validation lands at three sites: bound captures + data-table rows at bind time (proef::bind::param_type_mismatch), and static defaults / with: values at load time (next to the existing default_not_param / unknown_with_key checks). All checks are skipped when the raw arg contains ${ or {{ (constraint 2). Datetime types parse via jiff (sans-IO — literal parse reads no clock).
  • proef schema reflects the typed params shape (constraint 1 satisfied).

Consequences

  • Breaking: every pack’s params migrates to the map form. A one-time mechanical change, called out in CHANGELOG and GETTING-STARTED; the error-corpus gains a bind__param_type_mismatch case.
  • Best-effort, not a guarantee: literal args are checked; ${…}/{{…}} args are not. This must be documented so authors do not read it as a type system — its value is catching typos in literal prose at --dry-run, not enforcing runtime types.
  • --dry-run gains a real new class of caught mistake; proef schema autocomplete gets richer.
  • The honest cost/benefit: a breaking migration for a best-effort literal-args lint. Recorded so the trade is deliberate, not incidental.

Best-practice basis & recommendation

Research into the field (2025–2026) is decisive on form and honest about worth:

  • Form (if built): the industry norm is declaration-site typing referenced by name, never a full type spec inline. Cucumber Expressions put only a bareword type-name in the pattern (indexing a registry); SpecFlow/Reqnroll type by the binding method’s return type; Bruno (v4, 2024) added declaration-site typed variables with a string default. A name → type map is the canonical shape (JSON Schema properties, OpenAPI). The inline {name:type} form (behave/pytest-bdd) only works because their “declaration” is code, not a schema — with a schema it loses tooling. So the decision above (single-shape map, closed vocabulary) is the correct form, and the list-or-map union is the worst option on every axis (permanent double surface, weaker autocomplete, bifurcated examples — the exact ambiguity “one canonical way” forbids). ESLint’s flat-config hard break (v9→v10, bounded deprecation then cut) is the precedent for a single-owner format choosing a clean break over an indefinite dual-shape.
  • Worth (candid): a best-effort lint over literal values is the mypy/TypeScript gradual-typing bargain — genuinely useful for shallow typos, provided it degrades to “unchecked,” never “pass,” on the parts it can’t see. But an API test runner’s args are deferred-heavy (base URLs, captured ids, run-time tokens — all ${…}/{{…}}), the exact population the lint cannot check, which shrinks the realized benefit. proef’s own corpus bears this out: today’s params are overwhelmingly string-ish, so the high-value types (uuid/int/iso8601) would rarely fire. Karate — the nearest-neighbour API tool — never types inputs at all; it types responses.

Recommendation: defer. The clean, best-practice-aligned move for proef right now is not to add a low-current-value lint behind a breaking format change; it is to record the correct shape (above) and adopt it if a pack corpus emerges where structured literal args are common. Building it earlier is defensible only in the disciplined form above — never as a list-or-map union. (If the operator wants it now regardless, implement the single-shape map with the honest-degradation discipline; the migration is mechanical.) (Cucumber Expressions, Reqnroll conversions, Bruno typed variables, OpenAPI 3.1, ESLint 10 removal, Karate schema validation)

Alternatives considered

  • Inline {name:type} (Cucumber-Expression style) — rejected: invisible to proef schema, mis-binds in the tokenizer ({q:uuid} becomes a capture literally named q:uuid), and covers only 1 of the 4 arg sources.
  • params accepts either a sequence (untyped, non-breaking) or a map (typed), via custom Deserialize — the non-breaking alternative. Rejected under one-canonical-way (two shapes for one field), but it is the fallback if a breaking migration is judged too costly for a best-effort lint. (This is the open fork for the operator.)
  • Keep params untyped (do not do N1) — the status quo; the shift-left gap stays. A legitimate choice given the cost/benefit above.
  • User-defined type registry — deferred; a closed set is simpler and covers the common shapes.

ADR-0014 — Suite-level setup & teardown

Status: Accepted · Date: 2026-08-02

Context

Gherkin Background runs steps once per scenario, and a persistent global store threads saveAs: global captures across scenarios and runs (ADR-0005). What is missing is a once-per-suite phase: authenticate once and share the token, seed a fixture before the run and tear it down after. Cucumber solves this with @BeforeAll/@AfterAll hooks; Karate with callSingle. Round-2 code validation (IMPROVEMENT-PLAN §12, item N3) fixed the constraints:

  • Tags never reach the sans-IO core runner. ScenarioSpec/ScenarioOutcome carry no tags; @quarantine is computed at the CLI edge (exec.rs) as a non_gating set. So a lifecycle phase must be orchestrated at the CLI edge, not inside the runner.
  • Only saveAs: global promotions cross scenarios, and each pooled scenario snapshots the global store at prepare time — so a setup phase must finish and merge its globals before the parallel pool starts, or early scenarios miss the seeded value.
  • An assert-failed setup must not be masked. A setup step that fails an assertion classifies as a test failure (fault: None), which the worst-wins fold would let through as exit 1 without aborting the pool — every scenario would then run against un-seeded state and cascade confusing failures.

Decision

  • Two new proef.toml keys — [run] setup and [run] teardown — each naming a feature file. setup runs (its scenarios, in order) once before the suite pool; teardown once after. State reaches the suite through the existing saveAs: global store; setup completes and merges before the pool is built.
  • Orchestrated in execute() around the parallel pool (CLI edge), never in the sans-IO core. The setup/teardown feature is excluded from build_specs so it never also runs as an ordinary scenario, and it is invisible to --tags/--scenario/--rerun.
  • Failure semantics (exit-code contract, ADR-0009) — validated against Playwright, Jest, k6, and pytest (see basis below):
    • A setup failure of any kind aborts the run before the pool launches and is never masked. It maps to a user (2) or system (3) fault, not a test failure (exit 1) — a broken fixture is not a failing test, the same distinction Playwright draws between a “clear setup error” and a “cryptic test failure.”
    • Teardown runs only when setup succeeded (gated on setup-success, not pool-success): k6 and pytest-yield both skip teardown when setup threw, because tearing down un-created state is itself a fault source. The setup feature is responsible for cleaning its own partial state on failure.
    • Teardown does run after the pool even when scenarios failed (Playwright/pytest/k6 all do), so cleanup is reliable.
    • A teardown failure is loudly reported and yields a distinct non-zero signal (a system/cleanup fault, exit 3) — never a silently-green suite (no mainstream tool masks a teardown failure). It stays distinct from a test failure (exit 1): a cleanup hiccup does not mean the API under test is broken, but it is not hidden.
  • --dry-run is unaffected: it validates only (never calls the runner), and the setup/teardown features are validated like any other feature but never executed.

Consequences

  • A genuine once-per-suite lifecycle, expressed declaratively (Hurl prose in a feature), with no code hooks — the glue-code path proef exists to avoid.
  • Exactly one setup and one teardown mechanism — a [run] construct, not a second Background concept and not a tag. The setup feature is authored like any suite feature and reuses the whole pipeline (bind → lower → emit → execute).
  • Touches the pinned exit-code tests: a failing setup gates the run; new assert_cmd cases cover the short-circuit and the “setup fault is never masked” invariant.
  • Product-neutral: the setup steps are ordinary prose bound to macros; nothing about the mechanism assumes a particular backend.
  • Auth-once boundary (explicit): the pattern that motivates suite-setup elsewhere — “authenticate once, share the token” (Playwright storageState, Karate callSingle) — is largely already covered in proef by a pre-set secret (proef secret set / PROEF_SECRET_*) resolved per scenario. Sharing a runtime-obtained secret through setup would collide with the invariant that saveAs: global refuses secret-valued captures (ADR-0005). Promoting a runtime capture into the secret channel is therefore out of scope for this ADR: the global store carries only non-secret setup state (seeded ids, fixture config). Recorded so the boundary is explicit rather than discovered later.

Amendment — 2026-08-10 (cancellation, and what --dry-run validates)

This ADR was specific about a failing setup and a failing teardown, and silent on the operator interrupting the run. That silence read as considered when it was not, and the behaviour it left was the opposite of this ADR’s premise: on Ctrl-C the teardown phase ran with the already-cancelled token, so every teardown scenario resolved Skipped, phase_failed ignored a phase that only skipped, and cleanup silently never happened.

Decision — cleanup outlives the interrupt. Teardown runs on its own, independent token, never the run’s. On Ctrl-C the pool stops at its batch boundary, the operator is told cleanup is running, and teardown completes — so an interrupted run does not strand whatever setup created.

Note the word: independent, not “child”. A child token cancels when its parent does, which is precisely the behaviour being fixed; child_token() would have re-implemented the bug.

ADR-0007’s responsive interrupt is preserved by the escape hatch that already existed: a second Ctrl-C hard-exits (130) out of teardown as out of anything else, and the announcement says so. A hung teardown is bounded by the same batch budgets and watchdog as any other phase, so this needs no timeout of its own.

This is the standard graceful-shutdown shape — first signal begins bounded cleanup, second forces exit — rather than the test-runner norm, which is worse: Jest does not call globalTeardown on Ctrl-C (#6029) and Go does not run t.Cleanup on SIGINT (#41891), both long-standing complaints rather than settled design.

Corollary — a phase that only skipped is a failure. A skipped phase carries no fault, so the worst-wins fold passed it silently; that is the shape that hid cancelled cleanup. A setup that completes no scenario now aborts the run (the suite would otherwise execute against state setup never created), and a teardown that completes no scenario is reported and fails the run. The setup abort is also what keeps teardown gated on setup-success as this ADR requires: the early return is the gate, so teardown never dismantles what was never built.

--dry-run now does what this ADR already claimed. The Decision above says the phase features are “validated like any other feature but never executed”. They were validated by nothing — --dry-run never read the keys — so a broken [run] teardown surfaced only after a full suite had run, while the identical mistake in [run] setup failed in milliseconds. Both phases are now validated by one loader shared with execute, which also pre-flights teardown before the pool, so the same mistake costs the same either way and a bad path is a user error (2) rather than a blanket system fault (3).

Best-practice basis

Config-key-names-a-file is the dominant model (Playwright globalSetup/globalTeardown, Jest globalSetup, Vitest) — the [run] table already supplies the “once per run” qualifier, so [run] setup reads like globalSetup without repeating global. Passing state out-of-band through a serialized shared store (not live memory) is universal — Playwright’s storageState file, Jest’s env-var workaround, Karate’s callSingle cache, k6’s returned data — and proef’s global store is the direct analog. Setup-completes-before- workers and setup-failure-aborts are unanimous. The refinements above (teardown skipped on setup failure; teardown failure is non-zero, not silent) come straight from k6/pytest and the Playwright/pytest “teardown is not silently swallowed” norm. (Playwright global setup, Jest config, k6 lifecycle, pytest #2508, Karate callSingle)

Alternatives considered

  • @setup / @teardown tags — rejected. A tag means “filter/select”; overloading it with “change execution phase” is a second meaning, and tagged setup scenarios would entangle with --tags/--scenario/--rerun/flows/name-dedup with undefined ordering between two @setup scenarios. A [run] construct is single, explicit, and ordered.
  • Code hooks (Before/After functions) — rejected. Arbitrary-code hooks are exactly the imperative glue the declarative-macro model avoids, and they break the sans-IO boundary. The legitimate need (setup/teardown IO) is met with Hurl steps instead.
  • Per-scenario Background only (status quo) — does not cover once-per-suite work; authenticating in every scenario’s Background is wasteful and cannot seed shared state that must exist before the first scenario.
  • A dedicated [setup]/[teardown] top-level table — rejected as heavier than needed; these are run-orchestration knobs, so they belong under [run] beside suite/jobs.

ADR-0015 — Injected observability timestamps (run-level timeline)

Status: Accepted · Date: 2026-08-02

Context

The HTML report has a per-scenario timing waterfall (IMPROVEMENT-PLAN §12, N6a) derived purely from each step’s duration_ms. It cannot show cross-worker occupancy — which scenarios ran concurrently, on which of the --jobs workers — because the sans-IO core reads no clock and the event stream carries no wall-clock timestamp or worker identity.

The core’s purity is deliberate (deterministic snapshots/properties). ADR-0012 established the escape valve: values the core must not compute (config) are injected at the CLI edge. run_id is already such a value — the core carries it (RunStarted.run_id) but the CLI generates it. The same pattern extends to timing.

Decision

  • Add two optional, additive fields to the ScenarioStarted and ScenarioFinished events: timestamp_ms: Option<u64> (milliseconds since the run began) and worker: Option<u64> (0-based worker index). Both are #[serde(default, skip_serializing_if = "Option::is_none")], so old records parse unchanged and single-threaded/None cases serialize identically (ADR-0008 additive-only). They are kept off RunStarted, whose exact wire bytes are pinned.
  • The core leaves them None — it carries fields it never fills, exactly as it carries a run_id it never generates (sans-IO preserved). The CLI wraps the event sink in a stamping closure: EventSink::new(move |ev| inner.emit(&stamp(ev))). stamp runs on the worker thread (emit is synchronous, called from the scenario worker), so it reads the run-start Instant and maps thread::current().id() → a stable 0-based index there, at the edge. The core never sees a clock or a thread id.
  • Errata (2026-08-11): only ScenarioStarted carries a worker. As implemented, ScenarioFinished is emitted from the main dispatcher thread, not the worker that ran the scenario, so stamping a thread index there would name the wrong one; it carries the end timestamp and worker: None. The worker identity comes from ScenarioStarted, which is emitted on the worker thread, and the timeline pairs the two. EVENTS.md has always described it this way — the bullet above did not.
  • The HTML report gains a run-level timeline: a lane per worker, each scenario a bar from its start to its finish timestamp — the Gantt/occupancy view. It renders only when the stamps are present; a record without them falls back to the N6a per-scenario waterfall alone. The view stays a pure function of the (now richer) event record.

Consequences

  • Core stays sans-IO; the injection is the established run_id/config pattern.
  • The JSONL record gains observability fields, additively (ADR-0008 — still the one record; no sidecar timing file).
  • Snapshot honesty: old records parse unchanged, but every event in a new run now carries a (non-deterministic) timestamp_ms, so the reference_event_stream snapshot changes — it needs a new insta filter ("timestamp_ms":\d+ → 0) plus a deliberate cargo insta review. This is allowed (the same deliberate-acceptance policy as emitter changes), and is called out rather than glossed. The HTML snapshot likewise regenerates.
  • Enables the cross-worker timeline; the derived HTML view remains pure over the record.

Alternatives considered

  • Stamp inside the core — rejected: reading a clock in proef-core breaks the sans-IO invariant that makes snapshots and property tests deterministic.
  • A second sidecar timing file — rejected: ADR-0008 makes the JSONL event stream the record; a parallel timing artifact is a second record format.
  • Derive occupancy from duration_ms alone (no injected fields) — impossible: without an absolute clock there is no way to know which scenarios overlapped in wall-clock time; only the sequential per-scenario waterfall (N6a) is derivable, which is why it shipped first.
  • Wall-clock (unix ms) instead of run-relative — rejected: run-relative starts the timeline at 0, is cleaner to render, and avoids putting absolute wall-clock in the record.

ADR-0016 — OpenAPI → suite generator (scope decision)

Status: Proposed — recommendation: defer; the oracle/drift mode is permanently rejected regardless · Date: 2026-08-02

Context

A recurring market expectation of a “serious” API-testing tool is to generate tests from an OpenAPI spec (Schemathesis, Dredd, Step CI). proef has none, and the Round-2 scope stress-test (IMPROVEMENT-PLAN §12.5-A) flagged it as the single biggest capability a reviewer would call missing — while also being the one that sits on proef’s permanent charter line.

PRD §3 names, as a permanent non-goal, “API mocking/contract testing”, and the operative expansion (IMPROVEMENT-PLAN §3) spells it: “contract testing (OpenAPI drift / Pact / Schemathesis).” Yet the validation established that a generate-then-freeze framing — read the spec once at the CLI edge, emit concrete editable .feature + macro packs, then execute them through the normal deterministic pipeline — technically clears the sans-IO and determinism objections (ADR-0012 is exact precedent for IO-at-the-edge). So the question is genuinely open and needs an ADR to settle the boundary rather than let it erode.

Decision

The bright line (normative, decided either way): an OpenAPI spec may be a one-shot seed — read once to emit prose + packs the author then owns, edits, and maintains — but it may never become a recurring oracle: re-read on every run, used to drift-check the live API against the spec, or gated on a generated-vs-committed diff. That is OpenAPI-drift contract testing (PRD §3), and it is permanently rejected. Concretely, a generator, if it ever exists, must have no --check/--verify/--diff mode and must never be consulted after the initial emission.

On the narrow scaffolder itself: defer (recommendation). A one-shot proef generate --openapi spec.yaml -o suite/ that scaffolds editable prose + packs (the shipped bind::unbound_step stub-gen at suite scale) is defensible in-charter under the bright line, but is not worth building now for the reasons below. If a concrete need emerges, it may be adopted only as: CLI-only (never proef-core), deterministic given a seed, output fully owned by the author after emission, and bound by the line above.

Consequences

  • Deferring records the boundary so the omission is an intentional, documented call — and so the oracle mode is now explicitly foreclosed, not merely absent.
  • Output-quality tension (the strongest argument against). OpenAPI describes an API in endpoint terms (POST /records/{id}/notes → 201); proef’s value is business prose (“a member posts a note to a record”). A generator produces the former, which the author must rewrite into the latter — so it saves little over hand-authoring and risks a corpus of mechanical prose that undercuts proef’s prose-first premise.
  • Dependency + direction cost. A robust OpenAPI 3.0/3.1 parser (the dialect split, $ref, oneOf/allOf, discriminators) is a heavy, churny supply-chain surface to pin and clear through cargo-deny/cargo-audit. It also introduces a new inward-ingestion direction — proef has only ever flowed outward (ADR-0010: artifacts are emitted, never imported).
  • One-canonical pressure. The first emission is cargo new-style and harmless; the risk is re-generation, which would give packs a second maintenance path (regenerate vs hand-edit). The one-shot-seed discipline is not machine-enforced, so it must be a documented rule.
  • Already partly covered. The #9 stub-gen convenience emits a paste-ready match:+hurl: macro skeleton for an unbound step today — the same idea at step scale, without any spec dependency.

Alternatives considered

  • Full OpenAPI-drift checker (Schemathesis/Dredd-style: spec re-consulted per run to catch divergence) — permanently rejected. It is verbatim the PRD §3 non-goal, non-deterministic against a live API, and the reason the bright line exists.
  • Property/fuzz-case generation from the spec (Schemathesis-style negative cases) — out. Runtime fuzzing breaks the sans-IO/deterministic-artifact invariants; frozen generated fuzz cases are a variant of the one-shot scaffolder and share its deferral.
  • Build the narrow scaffolder now — declined on cost/value (output quality, dependency weight, inward direction) despite being technically in-charter under the bright line. Buildable later without a new ADR, provided it obeys this one.
  • Status quo (no generator) — the recommended near-term state; authors write prose against the macro vocabulary, aided by stub-gen for missing steps.

Best-practice basis

The generate-then-freeze vs re-consult-the-oracle distinction, the sans-IO-clearance via the ADR-0012 IO-at-the-edge precedent, and the “one --check flag from the non-goal” risk are from the Round-2 scope validation (IMPROVEMENT-PLAN §12.5-A). Schemathesis and Dredd are the reference generators; both re-consult the spec as an oracle, which is exactly what this ADR forecloses. (Schemathesis, Dredd, PRD §3, IMPROVEMENT-PLAN §12.5-A.)

ADR-0017 — proef lsp language server

Status: Accepted · Date: 2026-08-03 (implemented 2026-08-03) Design spec: docs/superpowers/specs/2026-08-03-proef-lsp-design.md

Context

Feature/pack authoring has no editor support — no live diagnostics, no jump-to-macro, no step completion. Weak IDE support is a documented gap in Karate and the BDD field (IMPROVEMENT-PLAN §10), and it is the one substantive unbuilt item on the Round-1 roadmap (§5 item #11, graded ✅ FITS). proef is unusually well-placed for it: its analysis is already headless and sans-IO — front::run yields the same Diag objects (stable code, byte span, severity, help) regardless of driver — so an LSP is a second front-end over existing analysis, not new analysis.

Decision

Build a proef-lsp crate (surfaced as the proef lsp subcommand) as a server-only, generic-LSP stdio binary delivering the full v1 feature set — diagnostics, go-to-definition, completion, and find-references. Key choices:

  • Sync lsp-server + lsp-types (rust-analyzer-family), not an async stack — honouring ADR-0006’s tokio ban.
  • Whole-suite model, recomputed wholesale on change (debounced), not an incremental (salsa-style) index. Mature LSPs index the whole project because cross-file features are core; they carry incremental machinery only because their scale is large. proef’s suite is tens of small files and the pipeline is milliseconds, so wholesale recompute buys the cross-file capabilities (live cross-file re-validation, find-references) at per-document simplicity. Incremental is YAGNI until a real perf ceiling.
  • Second front-end over sans-IO core. The one enabling refactor: an injectable source provider so file discovery + reading go through a trait (disk for the CLI, overlay-then- disk for the LSP), and a collect-all mode (accumulate every diagnostic instead of fail-fast). Both keep proef-core sans-IO — the IO is injected, the ADR-0012 pattern.
  • Server-only v1; a VS Code extension is deferred (off proef’s pure-Rust brand; a thin wrapper can follow).

Consequences

  • A new crate + two deps, pinned as shipped: lsp-server 0.7.9 and lsp-types 0.97.0 (MIT/Apache — clean under cargo-deny); the provider trait moves the proef-core public-api snapshot (a deliberate, reviewed change).
  • lsp-types 0.97 models document URIs as its own Uri type (RFC-3986), not url::Url. The 0.97 line dropped the url dependency, so the converter and every handler key documents on Uri (parsed/compared as an RFC-3986 string), never url::Url — a change from the pre-0.97 API that would silently fail to compile against the old assumption. Superseded — see the amendment below.
  • The front-end refactor (injectable provider + collect-all) touches front.rs and its callers; the CLI path must stay behaviourally identical, guarded by the existing integration
    • snapshot suites.
  • Highest bug-risk surface is the byte↔UTF-16 + source-normalization (BOM, trailing newline) converter; mitigated by property tests and by reusing tests/errors/ (real spans/ranges).
  • A genuine competitive differentiator, and the natural depth move after the v0.4.0 breadth work — but a multi-week (L) effort; sequenced so each step (handshake → provider/collect-all/ converter → diagnostics → definition → completion → references) is independently testable.

Amendment — the types crate is gen-lsp-types, and Uri is url::Url again

lsp-types stopped receiving releases after 0.97; gen-lsp-types is the maintained successor, generated from the LSP metamodel. proef depends on it under the original name — lsp-types = { version = "0.11", package = "gen-lsp-types", features = ["url"] } — which is rust-analyzer’s own aliasing pattern and leaves every lsp_types:: path in the crate untouched. The decision above is unchanged: still sync, still the rust-analyzer family, still server-only.

Two consequences change:

  • Uri is url::Url. The generated crate gates its URI type behind features (url, fluent-uri, or a bare String newtype). Choosing url costs nothing — the embedded hurl engine already pulls url into the workspace graph — and buys back from_file_path/to_file_path, the native-path bridge the 0.97 Uri had no equivalent for and that documents.rs therefore hand-rolled (drive-letter prefixes, segment joining, percent-encoding). That bridge is deleted; the wrapper that remains exists only to pin the pipeline’s source-name identity rule. The consequence the original ADR recorded is retired with it, and the swap removes three crates (lsp-types, fluent-uri, serde_repr) while adding none.
  • Methods are enums, not string constants. Request::METHOD is now an LspRequestMethod<'static> whose From<&str> falls back to Custom, so dispatch compares enum values and an unrecognised method lands in a variant rather than matching nothing.

One behaviour moved: what counts as a malformed document URI. fluent-uri rejected a raw space; url percent-encodes it. The malformed-params test therefore asserts on a schemeless URI, which url genuinely rejects — the guarantee under test (a bad URI is answered with InvalidParams, never a dead server) is unchanged.

Alternatives considered

  • Incremental (salsa) index — rejected for v1: it is the complexity mature LSPs accept for large scale, which a test suite does not have. Buys nothing at proef’s scale; can be added later behind the same interface.
  • Per-document analysis (no cross-file model) — rejected: it would leave stale diagnostics when a pack changes and cannot do find-references; the wholesale-recompute model gives the cross-file behaviour without the index cost.
  • Async LSP stack (tokio + tower-lsp) — rejected: violates ADR-0006; lsp-server is the sync, rust-analyzer-proven alternative.
  • Ship a VS Code extension in v1 — deferred: a TypeScript/npm deliverable off the pure-Rust brand and a separate release surface; generic-LSP config covers Neovim/Helix/Emacs/Sublime today, and the extension is a thin follow-up.
  • Diagnostics-only or diagnostics+go-to-def MVP — considered; the operator chose the full feature set for v1 (the complete authoring experience), sequenced internally so value still lands incrementally.

Amendment — 2026-08-04 (go-to-definition gaps closed)

v1 go-to-definition resolved only feature step → macro, landing on the macro’s name key; two narrower targets were cut and recorded only in a source comment. Both are now implemented: a use: reference inside a pack jumps to the macro it names, and either path lands on the macro’s match: line when one is locatable (falling back to the name key for use-only macros). Both are best-effort text-scan locators in proef-core::pack::locate, indexed at analyze time, following the existing sans-IO/text-scan idiom already used there — no parser change.

ADR-0018 — Named hurl fragments: a second macro body form

Status: Accepted · Date: 2026-08-11

Context

ADR-0004 made a macro step’s HTTP payload a raw hurl block embedded in the pack YAML. Field evidence says that was right: a real 844-line, 14-file hurl corpus was ported onto proef through the paste path with 100% coverage — all 844 lines used only [Asserts] (75) and [Captures] (15), with seven ordinary predicates, every one passing through untouched (OPEN-FINDINGS §“Positive evidence”). Nothing here weakens that.

Two things the embedded form cannot do, both structural rather than incidental:

  1. The block is not valid hurl. It carries ${url:…} / ${secret:…}, so parse_hurl_file only ever sees it after a probe substitution that tries {{probe}}, then 1, and accepts whichever parses (pack/validate.rs, probe_lower). Editors, hurl’s own tooling, and hurl itself see a file they cannot read.
  2. A block has no name, so it cannot be shared. use: composes macros, not payloads; two macros wanting the same request duplicate its text. For a corpus somebody else owns, the only route in is transcription — and a transcript drifts from its original the day after it is made.

The adoption question the worklist is now on (M1–M3) is not “can a suite be ported” — it can — but whether a team can keep both suites alive long enough to trust the new one. That needs one source of truth, not two copies.

Decision

A macro step’s body is hurl: | (as today) or ref: <fragment>, never both. A fragment is one hurl entry in a real .hurl file, named by a comment directly above it:

# tests/hurl/admin.hurl — runs under stock hurl, unmodified
# @proef admin.search
GET {{base}}/api/v1/admin/search/{{index}}
Authorization: Bearer {{apiToken}}
[Query]
q: {{q}}
HTTP 200
[Captures]
recordId: jsonpath "$[0].id"
bind:                                    # pack scope
  base:     ${url:base}
  apiToken: ${secret:apiToken}
macros:
  searchRecords:
    params: [q, index]
    defaults: { index: records }
    bind: { q: "${q}", index: "${index}" }   # macro scope (quote in flow style)
    steps:
      - ref: admin.search
  • The annotation carries a name and nothing else, permanently. No retry=, no key/value growth. A comment holding one identifier and zero behaviour cannot become a second configuration language, needs no parser or schema of its own, and cannot drift from the YAML. All orchestration stays in the pack.

  • Names are free-form dotted, globally unique (as macro names already are), with file.hurl#name as a disambiguator (as pack.yaml#name already is).

  • One annotation claims exactly one entry. Forced, not preferred: a fragment reused by several macros must be composable into any position of any of them, and a multi-entry region is a fixed sequence only its original neighbours can reuse.

  • bind: maps a fragment’s {{names}} to proef values, at pack, macro and step scope, most specific winning, nothing implicit. A foreign corpus names its own variables; convention-matching would only work on files we wrote. Values resolve once per scope instantiation — one binding is one value, two bindings are two values.

  • Non-secret bindings emit as per-entry [Options] variable:; ${secret:…} never does, and routes to insert_secret as it always has, so no secret value enters an artifact (ADR-0005 intact).

  • Every {{placeholder}} must be bound, produced as a [Captures] name by a preceding step, or supplied by the fragment’s own [Options] variable:. A fragment’s interface needs no declaration anywhere: what it reads is read off hurl’s own AST and crosses the seam as ScannedFragment::placeholders.

    The third source was missing from this list until it was found by audit, and its absence contradicted the decision above it: a file that answers its own question needs fewer variables passed in, which is precisely what makes it runnable on its own — so refusing those files refused the ones this ADR exists to accept. What a fragment supplies itself crosses the seam as ScannedFragment::supplied_variables.

    It is a supplier, so it also collides with a bind: of that name. Both reach the entry as variable: <name>=, hurl takes the last, and the fragment’s own line is last — so the bound value would never be sent, and hurl’s variable: assigning into the run-level set rather than scoping means the loss persists into every later entry. Refused as pack::option_declared_twice, the same rule a doubly-declared retry: gets, rather than resolved by an implicit precedence nobody wrote down.

    What earlier steps produce is derived differently, and deliberately: the core scans the emitted text of the steps it has already lowered (emit::capture_names). It has to, because a preceding step may be an inline hurl: block, which no scanner ever saw — there is no ScannedFragment for it. So the two halves of this check reach the core by different routes, and the produced half is the one that knows hurl’s [Captures] syntax inside proef-core. A second engine would need produced names on the seam instead; nothing is scheduled, and the note is here so the asymmetry is a recorded decision rather than a discovery.

Both body forms stay, because they are not two spellings

Inline does lower-time text splicing; ref: does run-time binding. Neither subsumes the other, and the boundary is demonstrable rather than stylistic: tests/features/packs/breadth.yaml’s postCustomNote splices ${docstring} — a multi-line Gherkin docstring — in as a request body. A hurl [Options] variable: value is a single-line scalar (VariableValue is Null/Bool/Number/String), so no binding can express it. Conversely no inline block can be named, shared, or run by stock hurl.

Splicing can substitute anything anywhere and is private to one macro. Binding is limited to what hurl can template, and buys a name, reuse, standalone runnability, and a static interface check. Choose by capability, not taste.

Consequences

The same file runs under proef test and under hurl file.hurl --variables-file … — one source of truth, no transcription, and ADR-0004’s “copy-paste flows both ways” becomes “no copy at all”. A fragment’s required and produced variables are machine-known, so an unbound placeholder is a load/lower-time error with a real file:line — a check the inline form structurally cannot perform. probe_lower’s two-candidate guessing does not apply to fragments: they parse as authored.

Costs, stated rather than discovered later:

  • A test spans three files (.feature → pack → .hurl) instead of two. explain and LSP go-to-definition have to earn that back — and both now do. Go-to-definition on a ref: line lands on the annotation; a ref: step records the fragment it ran as file.hurl#name in step_finished (an additive event field — ADR-0008 — absent for inline steps, so no pre-existing record changes a byte), which explain prints under a failure as via …. The qualified spelling is the one ref: itself accepts, so a post-mortem line pastes straight back into a pack.

    The record carries the name because it must stand alone: by the time anyone reads it, the pack that named the fragment may say something else. That is also why the path is shortened to a project-relative spelling before it is stored — [run] fragments resolves against the config file’s directory, and an absolute root would put a machine-specific path in a durable artifact and stop two checkouts’ records from comparing equal.

  • A bound value gets exactly one expansion pass, and no more. hurl’s eval_template is a single non-recursive pass, so a rendered variable’s value is never re-parsed as a template. But a [Options] variable: value is itself evaluated as a template before it is stored (hurl-8.0.1/src/runner/options.rs:508-511), which is one pass more than the entry body gets. So bind: { recordUrl: "${url:record}" }, where record is "${url:base}/api/v1/records/{{recordId}}", does work: {{recordId}} expands at option-eval time from a preceding capture, and GET {{recordUrl}} then renders the finished URL. [url]’s path table keeps working for fragments, and ADR-0012 is unchanged.

    The limit is the second level: a value that expands to text still containing {{…}} will emit those characters literally. In practice that is the same requirement the bound-or-captured rule already enforces — every placeholder must be in scope at the entry that reads it.

    (An earlier draft of this ADR concluded that fragment paths had to move into the .hurl file and that [url] would narrow to base. That was wrong: it read the non-recursive eval_template as applying to the binding path too. Recorded because the wrong version would have forced a needless proef.toml migration.)

  • Two ways to write a request body — accepted deliberately above, and the reason is recorded so the 844-line evidence is not forgotten by someone later tempted to deprecate the inline form.

  • proef reads files it does not own, so it must never write them: fmt refuses .hurl in directory discovery, and a foreign corpus stays byte-untouched.

Charter

PRD §3’s hurl non-goal is amended in the same change (PRD “Amendment (2026-08-11)”): it forbids generating Gherkin/macros/prose from hurl, not hurl text being an input source. Nothing here generates anything — features and macros stay hand-authored, and a fragment is inert until a macro names it. ADR-0016 stays declined on the untouched reasoning. The amendment records that OPEN-FINDINGS M3 asked for this re-examination to arrive with a measured port cost, and that it has not.

This is not M2 (mechanical equivalence between a hurl corpus and its proef port). The integration test that runs one fragment both ways proves the file is dual-runnable; it does not compare two suites’ results, and M2 stays open.

Alternatives considered

Region-claiming annotations (one comment claims every entry until the next) — fewer comments, but a claimed region is a fixed sequence, so the reuse requirement kills it; and pairing YAML modifiers positionally with claimed entries is exactly the coupling emit.rs’s MapEntry.step comment already warns against (“explicit, never positional”). One file per macro — trivial rule, but it dictates the layout of a corpus we do not own. Metadata in the annotation (# @proef.step retry=10x300ms) — maximum locality, at the price of a second configuration language with its own parser, schema and finite-retry lint, inside somebody else’s file. Replacing inline entirely — rejected: the 844-line corpus is evidence the paste path is sufficient for real work, and ${docstring} splicing has no binding equivalent. Implicit binding from config keys — shortest packs, but a name would bind from a file the pack never mentions, and a foreign corpus’s names rarely match anyway.

Prior art

The shape is well-established: -- name: magic comments naming queries inside valid .sql files (yesql, HugSQL, aiosql, sqlc); # @name naming requests in JetBrains’ .http client, which can import and run them by name across files; and the OpenAPI Initiative’s Arazzo specification, a separate declarative workflow document whose steps reference operations by operationId defined elsewhere — the same dependency direction chosen here, where the definition file knows nothing about its consumers. Reuse across hurl files is an acknowledged, unresolved gap upstream (Orange-OpenSource/hurl #317, #4574), so nothing here conflicts with a shipped hurl feature. The known failure mode of magic comments — invisible coupling — is answered by the name-only rule and by reading the annotation off hurl’s own AST, where comment-to-entry attachment is already modelled.

Amendment — 2026-08-12 (the scanner reports unannotated entries, by line)

FragmentScanner returns ScannedFile { fragments, unannotated } rather than Vec<ScannedFragment>. The named half is unchanged; the addition is the 1-based start line of every entry carrying no # @proef annotation.

The original contract dropped those entries at scan time, on the argument — still correct — that nothing downstream can use one, and that a corpus proef did not write is expected to be mostly unannotated, so building a ScannedFragment for each would be the bulk of a scan for nobody’s benefit. That reasoning covers building fragments. It does not cover counting, and the difference showed up in the field: a 97-entry corpus port found that missing an annotation on one entry produces a green dry-run and a silently absent test, with no signal at scan, bind, or run time — because the entry that would prove it was never built. Neither could any command state how many entries a corpus held, so there was no denominator against which the gap could be noticed.

A line number costs a push and is all a listing can point at, there being no name to print. The performance argument is therefore preserved intact: nothing extra is constructed, and proef fragments consumes what the scan already had to walk past.

Unannotated is not an error. proef fragments --check fails on annotated fragments no scenario runs; failing on unannotated entries requires --require-annotated. During a port “unannotated” means not done yet; in steady state it means deliberately not exposed, which is the premise that lets this ADR promise that pointing at a corpus you did not write costs nothing. Gating every adopter on the porting reading would have contradicted it, so the porting team asks for that check explicitly.

StepKindSpec also gained options, an engine-contributed recogniser mapping a raw option key to what the core’s ADR-0007 budget rules should make of it. The fragment half of that rule already crossed the seam (ScannedFragment::declared_options) while the inline half matched "retry-interval:" as a literal inside proef-core — one rule at two altitudes, and a second engine would have had its fragments linted and its inline blocks not. Option spellings now live only in the engine that owns them. Option baking (lower.rs) still writes hurl syntax directly; the emitter is hurl-shaped by ADR-0010 and is a separate question this amendment does not address.

Amendment — 2026-09-05 (a fragment’s file assets are its own, and staging is what delivers them)

“The same bytes run under stock hurl and under proef” was stated for the entry’s text. It was never true for the files that text reads. hurl resolves file,…; — a request body, a multipart part, a file-valued assert, and [Options] output: — against the directory of the file that wrote the reference (--file-root, defaulting to the .hurl file’s own parent). proef resolved every such reference against the feature, so a fragment’s asset, sitting where its own author put it, was unreachable: the same file passed under stock hurl and failed under proef as exit 2, blaming the author for a path that was correct.

Nothing worked around it. Moving the asset beside the feature breaks the standalone run this ADR exists to guarantee; a reaching ../ path is refused by hurl’s sandbox; and the advice that refusal prints — check –file-root option — names a flag proef does not expose and no proef.toml key supplies.

Per-source resolution, delivered by staging. Each asset is copied from beside the source that referenced it — the feature for an inline hurl: block, the fragment for a ref: — into that scenario’s own asset root, which is then the engine’s context dir.

Two roots at once was the obvious alternative and is not available: hurl exposes one context dir per run of entries and no per-entry override (its 42 OptionKind variants contain no file-root), while a single batch may mix both body forms — measured, not assumed: an inline step and a ref: step in one macro lower to one batch. Honouring two roots would therefore mean splitting batches on the authoring layout, which trades away the “batch maximally” rule TECH-SPEC §5 derives from run_entries building its client per call. Staging keeps batching intact, and copying fixtures into the build output is the standard answer to exactly this problem.

Three further consequences, each a correction rather than a cost:

  1. The root is per scenario. Flat staging was keyed by the asset’s bare name, so two features that each kept a data.json staged to one file — last writer wins, silently — and the loser’s artifact replayed against the other’s bytes. artifact_slug already refuses that trade for the .hurl text; the files it reads now match it. Two sources claiming one name within a scenario, which no per-scenario root can separate, is refused (proef::run::asset_unstageable).
  2. Staging is load-bearing, so its failure is fatal. A missing asset used to be skipped in silence because the copy only fed the record. It now feeds the run, and a file that did not arrive is a request reading nothing, not an incomplete record.
  3. ADR-0010’s promise widens. “Artifacts are the executed input” held for text while assets were read from the suite. The artifact directory is now the whole executed input, and an artifact that reads a file says so in its replay line (--file-root assets/<slug>). One that reads none is byte-identical to before, which is why exactly one snapshot in the corpus moved.

The sandbox also narrows: the context dir is a directory holding only what proef staged, rather than the suite tree it used to be (TECH-SPEC §13).

ADR-0019 — Reserved tags and the authored skip

Status: Accepted · Date: 2026-08-24

Emerged from the Robot Framework capability audit (OPEN-FINDINGS, “RF wave 2”); every design fact below was verified against the tree or reproduced empirically before acceptance.

Context

proef had no way to park a scenario. A test that must not run — mid-migration, a known-broken dependency, a seasonal flow — could only be deleted or dodged with --tags, and both are invisible: nothing in any report says “this exists and was deliberately not run”. Robot Framework’s SKIP model (its 4.0 headline design, replacing criticality) is the industry convergence point: skip is a first-class visible status with a reason that survives into every report.

Mechanically, proef already had a scenario-level Skipped status — but it arose only from cancellation, its reason existed nowhere, and two consumers had baked “Skipped means never-ran” into their logic (--rerun re-queues Skipped-on-cancelled; diff reads Failed→Skipped as fixed — and --fail-on-regression certified it). An authored skip that ignored those two would have shipped a laundering bug, not a feature.

Decision

  1. A reserved tag namespace, recognized at the CLI edge. @quarantine and @skip are the reserved tags; recognition lives in exactly one place (front::reserved), and core never reads tags — the front computes an instruction (ScenarioSpec.skip, like exclusive before it) per ADR-0014’s split. Reserved tags in [run] setup/teardown features have no effect: phases never pass through build_specs, and skipping your whole setup deliberately is spelled by deleting the config key.
  2. The spelling is @skip or @skip:<reason-token>. The gherkin grammar accepts skip:migration-pending as one tag (verified empirically through the real pipeline). The recorded reason is the pasteable tag spelling itself — "@skip" / "@skip:migration-pending" — the same philosophy as the fragment field’s file.hurl#name.
  3. Authored reasons start with @; mechanical reasons never do. That is the contract --rerun keys on: Skipped ∧ cancelled ∧ reason not authored re-queues as never-ran; an authored skip never re-queues. Pre-field records (no reason) read as mechanical, which they were.
  4. No tag-list normalization. An earlier draft injected a canonical skip atom beside @skip:x so --tags "not @skip" excluded both. Tag globs shipped first, and not @skip* says the same thing without proef ever rewriting an authored tag list. Authored tags stay exactly authored.
  5. A skipped scenario is selected, counted, and reasoned in every sink: console (∅ … — @skip:x), JUnit (<skipped message>), TAP (# SKIP @skip:x), the record (ScenarioFinished.reason, additive, schema stays 1), the HTML report, explain, flows --format json ("skip"), and the harness (libtest’s ignored flag). --tags remains the unselection mechanism — the two semantics stay distinct, as in RF.
  6. All-selected-scenarios-skipped exits 0. Exit 2 is for faulty input; the empty-selection refusal exists for the typo’d filter whose silent green run nobody sees. An all-skipped run is neither silent (every surface prints the totals and reasons) nor accidental (each skip is authored, versioned, and visible in review). RF and pytest agree; pytest reserves its special code for empty collection, which is exactly the case that stays exit 2 here.
  7. diff gives skip transitions their own bucket. Into-Skipped is neither fixed nor regressed (now skipped (was failing/passing)); out-of-Skipped has no meaningful baseline and takes the added shape.
  8. A quarantined test-failure reaches JUnit as skipped-with-message. The exit code already said “non-gating”; the XML said <failure>, so Jenkins marked UNSTABLE and every dashboard contradicted the verdict. RF converts the status for the same reason. User/System faults stay failures — quarantine is for flaky tests, not broken input.
  9. --dry-run still validates skipped scenarios. Skip is an execution-time decision, not a validation waiver — a broken-but-skipped scenario still fails --dry-run, deliberately.

Consequences

  • Library-breaking (clean break, no shims): ScenarioSpec.skip, ScenarioOutcome.reason, Event::ScenarioFinished.reason, ScenarioRun.reason, write_junit/write_ci_reports gain the non-gating list. Wire-additive; EVENT_SCHEMA_VERSION stays 1.

  • The sink wrappers that rebuild scenario events field-by-field (stamp_scenario_timing, phase_sink) must thread every new field — the exhaustive constructions turn forgetting into a compile error, and the e2e test pins the stamped stream.

  • A related stance this ADR writes down because the audit found it held but unwritten: control flow lives in packs (when: conditional skip at step level, optional: soft-fail, finite retry:) — prose stays declarative; there is no scenario-level IF/WHILE/TRY and none is planned.

  • 2026-09-06: a tag within a short edit distance of a reserved one (@quarantined, @skipped, @Skip) stays an ordinary, inert tag — but it now warns (tags::reserved_tag_typo) with the spelling it likely meant, since a scenario its author believed quarantined would otherwise gate the build in silence. Short reserved words get only a case-fold or a suffix match (ship/slip/step are one edit from skip); the long quarantine affords a distance-2 backstop.

ADR-0020 — Run metadata is explicit-injection-only

Status: Accepted · Date: 2026-08-24

The RF-audit wave-2 companion to ADR-0019; codifies the boundary R12-1 drew and the JUnit provenance decisions applied.

Context

A proef record could not say which commit, build, or environment produced it — run_started carried schema and run_id alone. Robot Framework’s --metadata name:value fills this in reports, and CI post-mortems genuinely need it: diff across records from different commits or environments has no context for what changed.

The hazard is on the other side. R12-1 removed harvested machine identity from every artifact because absolute paths broke the two-checkouts byte-equality ADR-0010 guarantees, and the JUnit sink deliberately omits timestamp/hostname for the same reason. Metadata must not reopen that door.

Decision

  1. The axis is harvested vs. handed-over, not automatic vs. manual. proef never reads git, the hostname, wall-clock provenance, or CI environment variables (GITHUB_SHA, CI_COMMIT_SHA, …). If the user wants the SHA recorded, their shell harvests it: --meta commit=$(git rev-parse HEAD). What the user explicitly hands over, proef records verbatim.
  2. One precedence chain, three scopes: [meta] < [env.<name>.meta] < --meta k=v — the same base < env < flags shape as jobs and [url]/[vars]. A duplicate key among the flags is exit 2 (loud over last-wins); a flag overriding a config key is the designed use. There is no PROEF_META_* — the values that motivate env vars are already in the shell where the flag is typed.
  3. The active --env profile name is recorded automatically, as its own field (run_started.env). It is user-chosen input to the invocation, not an observed machine fact — and without it the record is uninterpretable: the same suite deep-merges different [url]/[vars] per profile, so diff warns loudly on a cross-env comparison.
  4. shuffled rides the same head: with the permutation seeded by run_id, the bool plus the id reproduces an order exactly (deferred out of --shuffle’s own change so run_started moved once, not twice).
  5. Metadata reaches the record, explain, diff, the HTML report, the GitHub summary and the --format json body — and nothing else. Never artifacts (.hurl bytes stay identical across checkouts and commits — ADR-0010, R12-2); not TAP (no slot a consumer reads); not JUnit <properties> (GitLab ignores them, Jenkins reads them only behind a non-default opt-in — same named-consumer method as R3-6, additive later if a consumer asks); not the console (the record and explain own it).
  6. Everything passes the sink-boundary mask — keys and values both: a secret-bearing URL pasted into either position must not survive into the record or the body. The known limit stands recorded: a token proef was never told is a secret matches no needle, the same standing as any CLI argument.

Amendment (2026-09-07) — a computed input fingerprint is not harvested metadata

proef flaky’s equivalence-class fingerprint (the inputs.json sidecar) prompted the obvious question: does §1 forbid it? It does not, and the boundary is worth stating so the next reader does not re-litigate it.

§1 forbids proef from harvesting an environment fact — reading git state, the hostname, or CI variables and putting them in the record. The input fingerprint reads none of those. It is a hash of proef’s own inputs — the feature sources, the loaded macros and fragments, the resolved config scope — the same category as the artifact slug (emit::artifact_slug) or the shard hash: a derived identifier over data proef already holds, not a fact lifted from the surrounding machine. Derived identifiers have never been in scope here; ADR-0020 governs [meta]/--meta metadata, which this is not.

The git-commit case remains exactly as §1 requires: a user who wants commit-based grouping hands the commit over (--meta commit=$(git rev-parse HEAD), §1’s own worked example) and proef flaky --by commit groups on it. proef never runs git itself. So the survey’s “git tree SHA equivalence class” splits cleanly along this ADR’s own axis — a computed fingerprint proef may derive, plus a commit the user may hand over — and needs no new decision.

The sidecar is a derived aid like timings.json, not a second record (ADR-0008): the JSONL event stream remains the only record format, and the event schema is untouched (no run_started field was added).

Consequences

  • run_started gains env, metadata, shuffled — additive, skip-serialized when unset, EVENT_SCHEMA_VERSION stays 1 (ADR-0008 erratum extended). The empty case is byte-identical to every existing record.
  • Library-breaking: ProjectConfig/EnvProfile gain meta, exec::execute takes the merged map, RunRecord::open takes the head trio; clean break per policy.
  • proef stays sans-IO in core: the CLI merges and injects; core never reads an environment.

ADR-0021 — Run discovery and rotation are separate questions

Status: Accepted · Date: 2026-09-02

Context

--run-id accepts any single path component, and TROUBLESHOOTING.md demonstrates --run-id pr. A run named that way writes a complete, valid record. It is then invisible to proef explain, diff, flaky, report and --rerun whenever they resolve the latest run, because every one of them enumerates through record::all_runs, which admits only 36-character uuid names.

That predicate is not an oversight. rotate_runs consumes the same function and says so:

record::all_runs is the one answer to “what is a run record here” — sorted, uuid-named directories only. Rotation adds a single further exclusion (the in-flight run), not a second enumeration rule.

and is_run_id states the reason:

rotation deletes the oldest run-shaped directories, so breadth here is a deletion hazard when the runs dir points somewhere shared.

Both are right. runs-dir may be ., so a broad predicate would let rotation delete directories proef never created. That risk is real and the narrow answer is the correct one — for rotation.

The defect is that one predicate serves two questions whose risks point in opposite directions:

QuestionUnsafe whenConsequence
May I delete this directory?too broaddestroys user data
Is this a run I can show you?too narrowhides a real record

Sharing the predicate meant the deletion-safety choice silently became a visibility choice. CONFIG.md documented the rotation consequence of a custom id (“this will never delete it”) and not the discovery one, so the surprising half was the undocumented half.

Decision

Split the predicate along the risk, not along the file.

  1. Rotation keeps the uuid rule (is_rotatable, formerly is_run_id) — uuid-named directories only. Unchanged in behaviour, for the reason it was written; renamed for the question it answers. A custom-id record is still never deleted by [run] keep-runs, and that stays documented.

  2. Discovery admits any directory containing an events.jsonl. A directory holding proef’s own record file is a proef record; the test cannot mistake target/ or node_modules/ for one, and — decisively — it authorises no deletion. Reading a directory that turns out not to be a record fails as a parse error naming the file, which is already how a corrupt record behaves.

  3. Ordering stops relying on the name — but keeps relying on the uuid. all_runs documented that “uuid-v7 names sort chronologically, so lexical order is time order”, which a pr directory breaks. The fix takes the timestamp rather than the spelling: a uuid-v7 name carries 48 bits of unix milliseconds — the moment proef minted it — which is precisely why the lexical sort worked. Runs order by that where it exists, and by directory mtime where it does not (a custom --run-id carries no time). Both are wall-clock unix time, so the two sources compare directly.

    Not read from the record, though the first draft of this ADR said it would be. The head event carries event/run_id/schema and no timestamp at all — per-event times are injected observability on scenario_started (ADR-0015) — so “the record’s own run_started timestamp” does not exist. Reading it returned None for every real record and silently ordered everything by mtime; the test that covered it passed only because its fixtures fabricated a field no record has. An ordering that depended on the suite having run at least one scenario would not be an ordering anyway.

Consequences

  • proef explain, diff, flaky, report and --rerun find custom-id runs. “Latest” means latest in time rather than latest in the alphabet, which is what every one of those commands already claimed to mean.
  • Rotation’s blast radius is unchanged — the one property that could have made this dangerous.
  • Two predicates now exist where the code deliberately had one. That is the cost, and it is why this is an ADR rather than a patch: the earlier single-predicate statement was a considered position, and superseding it needs to be on the record. The two are named for their questions (is_rotatable / holds_a_record) so a future reader cannot reach for the wrong one by picking the shorter name.
  • Ordering costs a stat per custom-id run directory and nothing at all for a uuid-named one, whose time is read straight out of its name. Bounded by [run] keep-runs (200 by default) and paid only by commands that resolve “latest”.

Alternatives considered

  • Leave it, document it. The status quo before this ADR. Rejected: the invisibility surprises exactly the user who chose a memorable id so they could find the run again, and --rerun silently operating on a different run than the one just produced is the worst shape of it.
  • Make --run-id reject non-uuid names. Honest, and it would remove the trap — but it also removes the feature’s point. CI archiving .proef-runs/pr-1234/ by a known path is the use case --run-id exists for.
  • Rotate custom-id directories too, and keep one predicate. Rejected outright: it makes runs-dir = "." a data-loss configuration, which is the hazard is_run_id was written to prevent.

Erratum — 2026-09-06

The JUnit report’s own uuid still had the trap this ADR removed elsewhere: a non-uuid --run-id parsed to the nil uuid, so every --run-id ci run reported 00000000-… and collided in any consumer keyed on it. A custom id now derives a stable UUIDv5 from its bytes; a uuid id passes through verbatim.

proef — Improvement Plan

Status: complete. Round 1 (§1–§11) and Round 2 (§12) are both shipped — every batch (N2·N6a·N9·N8·N7·N4·N6b·N3), with N1 deferred per ADR-0013, N5 rejected, and the OpenAPI generator settled in ADR-0016. This file is the historical feature roadmap and its file:line citations are from 2026-07/08; the live worklist is OPEN-FINDINGS and the running ledger is CLAUDE.md’s Status block. · Date: 2026-07-31, appended 2026-08-02, closed 2026-08-31 · Owner: Emre Companion docs: PRD (scope + the binding non-goals, §3), adr/ (the invariants every item must respect), TECH-SPEC (types/pipeline), IMPLEMENTATION-PLAN (milestones + definition of done).

0. What this is (and isn’t)

A competitive analysis of proef against the BDD and API-testing field, converted into a roadmap of candidate improvements — each validated against the current architecture at file:line, and filtered strictly through proef’s permanent non-goals (PRD §3).

Nothing here is committed work; it is the durable output of the post-M5 review round. Effort grades (S/M/L) and sequencing are advisory. Feature numbers (#1…#15) are stable identifiers used across the tables. File:line citations are as of 2026-07-31 and will drift — treat them as “start reading here”, not addresses.

Every item is in scope by construction: none proposes a second engine, mocking, contract testing, load testing, a dashboard/server, OpenTelemetry, or importing hand-written hurl — all permanent non-goals (§3 below). The work is almost entirely surfacing data the sans-IO core already computes, not new engine capability.

1. Headline finding

proef’s engine is already ahead of its cohort; the gaps are in reporting surfaces and authoring/maintenance DX, not in HTTP power. Because proef pins hurl 8.0.1, it already ships hurl 8.0’s full assertion arsenal (RFC 9535 JSONPath with filter functions, type predicates isUuid/isIsoDate/isString/isObject/isList, response-time duration < ms, the filter chain split/count/toDate/base64Decode/daysAfterNow). Pack authors can write Karate-grade assertions today, inside raw hurl: blocks — they are just undocumented and under-surfaced. The roadmap is therefore mostly exposure and tooling, which is cheap, rather than engine work, which is done.

2. Where proef already wins (positioning to defend)

proef strengthCompetitor weakness it beats
One canonical way, raw-hurl-only, no escape-hatch languageKarate’s most-cited flaw: a three-language model (Gherkin + DSL + embedded JS) the moment anything gets non-trivial
Deterministic sans-IO coreKarate/Tavern/Newman are non-deterministic; reproducibility is now a headline selling point
Artifacts = executed bytes (hash-locked, git-diffable, replayable hurl --test)Postman’s opaque JSON collections (unreviewable merges) — the reason Bruno is displacing it
Finite retries + budgets + watchdog (ADR-0007)hurl itself has no cancellation and unbounded retries — proef fixes its own engine’s biggest gap
Property-tested secret masking + typed exit codes (ADR-0009)Most tools treat masking loosely and lack a stable exit-code contract
Dev-maintained macro packs = an enforced “step dictionary”The AI-authoring trend is groping toward exactly this; proef has it structurally
No cloud, no account, plain textPostman’s 2026 pricing exodus is driving the whole git-native wave

3. Scope guardrails — what this plan will NOT propose

Permanent non-goals (PRD §3) — never revisit under “competitor parity”: further engines (browser/gRPC/etc.), API mocking, contract testing (OpenAPI drift / Pact / Schemathesis), load testing, a desktop dashboard or server mode, OpenTelemetry export, dynamic plugin loading, importing hand-written hurl into Gherkin (artifacts flow outward only), static musl/Windows binaries.

Named anti-patterns (from the 2025–2026 trend research) to avoid: silent retries / green-on-attempt-2; retry ceilings > 3 or unbounded (proef’s finite cap is already correct); permanent quarantine without owner+expiry; treating masking as a security boundary; LLM self-healing / non-deterministic test mutation; config sprawl.

A Karate-style marker DSL (#uuid, ##optional) is explicitly rejected — it would be a second assertion mechanism competing with raw hurl predicates. Achieve the same readability by surfacing hurl’s native predicates (item #2), not by inventing a layer.

The overriding design gate is one-canonical-way (see §6): four items must replace or augment an existing mechanism, never add a parallel knob.

4. What code validation changed about the roadmap

Three deep code-validation passes (reporting, CLI/tags, authoring) reshaped the outside-in list. Six meta-findings:

  1. “Surface, don’t build” is confirmed at file:line. The sans-IO core already computes and records the data behind most items: attempts/duration_ms/detail on StepFinished (proef-core/src/event.rs), the winning macro_name per bound step (bind.rs:20), deterministic seeded fakes, and hurl’s own curl_cmd (already returned by EntryResult, currently discarded in engine-hurl session.rs).

  2. Two items are already ~90% built — validation downgraded them.

    • #14 seeded fakes: fakes are already deterministic (hand-rolled SplitMix64, no rand crate, fake.rs) and already seeded by the injected run_id, which is already recorded in RunStarted. The feature collapses to “let test pin run_id the way artifacts --run-id already can.”
    • #9 stub-gen: the “did you mean” fuzzy suggestion already exists (bind.rs:214, levenshtein); only the paste-ready stub template is missing.
  3. One item’s premise is broken — #13 impacted-only re-run. There is no stored content hash anywhere; ADR-0010 is enforced as a byte-identity assert_eq! on outputs (crates/proef-cli/tests/execute.rs:188), not a reusable digest. The only honest impact fingerprint — the emitted .hurl — is deliberately not stable run-to-run (run_id lives in runtime globals). #13 is gated behind a determinism prerequisite.

  4. Two plumbing gaps gate a cluster. Scenario tags stop at the CLI edge — they never reach the runner (ScenarioSpec/ScenarioOutcome in runner.rs carry no tags) or the event stream — so #15 (quarantine) and part of #8 (rerun) need new plumbing. And #4 (boolean tags) is hard-blocked by value_delimiter=',' on --tags (main.rs:57); it is a contract-changing replace, not an add.

  5. The recurring architectural rule is one-canonical-way. Four items (#4, #9, #14, #15) each risk a second mechanism; each must replace or augment an existing one.

  6. The exit-code contract (ADR-0009) is a live wire for #15. Quarantine changes which scenarios feed RunSummary::exit_code (fine — the fold stays pure in core) but must extend the pinned assert_cmd tests and never let a quarantined system fault mask exit 3.

5. The validated roadmap (master table)

Status is what exists in the tree, re-verified against main on 2026-09-08 — --help for flags, source for the rest. Verdict is the 2026-07-31 judgement of whether the idea fits the architecture; it never meant “done”, and reading it that way is why this table looked like a backlog when, by 2026-08-10, 13 of its 16 items had shipped (14 today; one partial, one gated).

Status: shipped · partial · open · gated (premise rejected). Verdict legend: ✅ FITS · ⚠️ NEEDS-ADAPTATION · 🚫 premise broken. Effort: S ≤ ~1 day · M ~days · L ~weeks.

#ItemStatusVerdictLives inEffortArchitectural truth (as of 2026-07-31)
2Assertion cookbook (surface hurl-8.0 predicates/filters)shipped✅ docs-onlydocs/SLowering copies all but ${…} verbatim (resolve.rs:163); load runs the real hurl_core parser (engine-hurl/src/lib.rs:95). Predicates already work. Caveat: grammar validation, not JSONPath semantics.
1GitHub ::error file=,line=,title= annotationsshipped✅proef-cli ci_reports.rsSfile+line+detail already flow to write_github_summary (ci_reports.rs:98). Gate stdout vs --output json; percent-encode multiline detail. Line-only (no byte-span at runtime).
5--curl export per requestpartial✅engine-hurl session.rsShurl’s EntryResult.curl_cmd is already returned, just discarded (session.rs:346). Must redact (holds resolved secrets, ADR-0005). Fold into the existing reproduce: block (exec.rs:230). Partial as of 2026-08-10: the curl line is surfaced on a failing step (exec.rs), which covers the debugging case; a per-request export flag for passing steps is still open.
3a“passed on attempt N” badge (JUnit/summary)shipped✅proef-cli ci_reports.rsSattempts:u32 already on StepFinished (event.rs:72) + StepOutcome; JUnit ignores it today (ci_reports.rs:43).
9Stub-gen for unbound stepsshipped⚠️proef-core bind.rsSAugment the existing did-you-mean help (bind.rs:89), zero-match arm only — not a new command. Derive {param} from quoted tokens (matcher already sheds quotes, matcher.rs:85).
10SARIF export of --dry-run diagnosticsshipped✅proef-cli new sarif.rsS–MDiag (diag.rs:56) → SARIF result ~1:1: code→ruleId, byte span→region.byteOffset. Pre-populate rules[] from the closed diagnostic-code set. A parallel serializer to render.rs.
14--seed (reproducible fakes)shipped (as run-id)⚠️proef-cli main.rs/exec.rsSThread into front::run’s existing run_id param (artifacts already exposes --run-id, main.rs:101). Caveats: arbitrary seed breaks JUnit’s UUID parse (ci_reports.rs:22); occurrence is per-scenario, not per-run — identical ${fake:X} at the same position in two different scenarios still draws the same value (a known limitation; see OPEN-FINDINGS). Note: the per-step reset this row originally cited (Refs::default() on every lower() call) was fixed in 0.6.0 — the counter now threads through lower with a high-water mark. Only the cross-scenario half remains. Shipped as the one knob §7 demanded (no separate --seed): test --run-id pins the fakes, and --shuffle seeds its permutation from the same id.
7Dead-macro / usage reportshipped✅proef-cli new macros --usageS–MBoundStep.macro_name (bind.rs:20) vs packs.macros. Count use:-only macros (pattern:None, pack/mod.rs:165) as reachable via the use: graph. Report the whole corpus, not a --tags subset.
6Self-contained HTML reportshipped✅core render_html(&[Event]) + cli writeMPost-hoc proef report <run-id> replaying events.jsonl like explain (explain.rs:12) — the command is the HTML report, so it took no --html flag; -o picks the file. Bodies live in artifacts/ — deep-link, don’t inline. Derived view, never a second record (ADR-0008).
12proef diff between two runsshipped⚠️proef-cli diff.rsMIdentity (file,scenario) (why ADR-0008 added file, event.rs:86); key step diffs on text not line (lines shift on edit). attempts+duration_ms → free flakiness/perf-regression detector. Pre-file records replay file="".
4Boolean tag expressions (@a and not @b)shipped⚠️ (replace)grammar in core, apply in cli front.rsMvalue_delimiter=',' (main.rs:57) actively breaks and/or/not; must replace the CSV/OR contract (front.rs:388), keep empty-match=exit-2 (front.rs:382). Grammar/evaluator is deterministic → proptest/fuzz-shaped, belongs in core.
8Rerun-only-failures (--rerun)shipped⚠️proef-cli, reuse explain replayMexplain already reads the latest record + failed (file,name) (explain.rs:100, event.rs:83). Needs a multi-identity predicate (today’s scenario/scenario-file filters are single-valued, exec.rs:301,321). Factor a shared record::failed_scenarios.
15@quarantine non-gating tagshipped⚠️ (contract)thread gating:bool core+cliMTags must first reach ScenarioSpec/ScenarioOutcome (P1). exit_code() stays pure in core (runner.rs:89) and skips non-gating outcomes. Events still emit the scenario → not hidden. Extend the pinned assert_cmd tests; never mask a Fault::System (exit 3).
3bTrue <flakyFailure> with earlier-attempt detailshipped⚠️schema + engine-hurlMNeeds an additive attempt_details field (ADR-0008 additive-only) + engine-hurl collecting per-retry bodies before the final one. Bigger than 3a.
11proef lsp (feature/pack language server)shipped✅new proef-lsp crateLAll diagnostic substrate is headless/sans-IO already (bind, pack::load, resolve Probe mode, matcher). New: a sync lsp-server (tokio ban forbids async), a byte-offset→token API (not exposed), and a partial-results wrapper (bind/load are all-or-nothing today). Karate notably lacks good IDE support → differentiator.
13Impacted-only re-run (content-hash)gated🚫— (gated)LNo input hash exists; raw-input hashing is unsound (shared packs, use: nesting, config vars fan out). Honest fingerprint = per-scenario emitted .hurl, but it is not run-to-run stable (run_id in globals). Needs a determinism prerequisite first; silent-green risk. 2026-09-07: a suite-level input hash now exists (inputs.json, for flaky windows) — not per-scenario; still gated.

6. Prerequisites that unlock clusters

  • P1 — carry scenario tags + a gating flag past the CLI edge into ScenarioSpec / ScenarioOutcome (runner.rs:30,121) and, additively, the event stream. Done (RF waves): tags ride ScenarioSpec/ScenarioOutcome and scenario_finished.tags; gating became the reserved-tag instruction plus the non-gating list (ADR-0019) rather than a bool. #15 and #8 both shipped on top of it.
  • P2 — a shared record::failed_scenarios(run_id) + a multi-identity scenario predicate. Reused by explain, --rerun (#8), and proef diff (#12).
  • P3 — a deterministic emitted-.hurl fingerprint (stable run-to-run despite run_id). Prerequisite for #13; do not attempt #13 without it. 2026-09-07: proef_core::fingerprint now hashes a run’s whole input set (feature sources, loaded macros and fragments, the resolved [url]/[vars] scope) into inputs.json — proef flaky’s equivalence class. Suite-level and deliberately coarse (any edit ends a window), so it is not the per-scenario, run_id-independent fingerprint #13 needs; P3 stands.

7. The one-canonical-way watch-list

Each of these must fold into an existing mechanism, never ship beside it:

ItemMust replace / augment (not duplicate)
#4 boolean tagsReplace the CSV/OR --tags semantics — no second tag syntax
#9 stub-genAugment the existing did-you-mean Diag.help — no separate proef stub command
#14 --seedFold into the existing run_id determinism knob — no parallel seed unless it replaces run_id-keyed fakes
#15 quarantineExactly one non-gating tag name; must not spawn a second “skip” concept
#5 --curlAttach to the single reproduce: mechanism, not a parallel debug path
  1. Batch A — free / small, all FITS, all reuse existing data. #2 cookbook (docs) → #1 annotations → #5 --curl → #3a attempt badge → #9 stub.
  2. Batch B — small, high-leverage. #10 SARIF · #7 dead-macro · #14 --seed.
  3. Prereqs → Batch C — medium, now unblocked. Build P1+P2, then #6 HTML · #12 diff · #8 rerun · #15 quarantine · #4 boolean tags · #3b flaky-detail.
  4. Batch D — strategic. #11 LSP. And #13 only after committing to P3.

9. Per-item detail & competitor provenance

Each entry: what it borrows from whom → the validated architectural note. Numbers cross- reference §5.

#2 Assertion cookbook — from Karate’s fuzzy markers + hurl’s own docs. Confirmed a docs task: resolve() leaves everything but ${…} byte-for-byte (resolve.rs:163-208, test runtime_tier_passes_through), and pack load validates the full grammar via hurl_core::parser::parse_hurl_file (engine-hurl/src/lib.rs:95), re-checked on the emitted artifact (front.rs:159). Extension: ship a tests/features/ reference feature exercising each predicate so the cookbook is snapshot-locked against hurl upgrades (the canary catches drift).

#1 GitHub annotations — from the 2025–2026 CI-reporting shift (annotations displace log-diving). The failures loop already prints `{file}:{line}` — {detail} (ci_reports.rs:98). Emit ::error workflow commands as a sibling; title = scenario + step text. Risk: stdout is owned by --output json (exec.rs:114) — gate it.

#5 --curl export — from hurl’s loved --curl; Bruno/Postman “copy as curl”. hurl hands us EntryResult.curl_cmd already (session iterates result.entries at session.rs:346 but reads only captures/errors/duration). Must pass Redactions (session.rs:396) before any sink — the curl line contains resolved secrets. Cannot be derived pre-execution (needs runtime {{…}}).

#3a “passed on attempt N” — from the flaky-test-honesty consensus (never hide a retry). attempts is first-class (event.rs:72, step.rs:149) and already printed on the console (report.rs:220); JUnit simply drops it. Count-based badge is S.

#9 Stub-gen — from Cucumber/Behave snippet generation. The matcher already computes the nearest macro via closest_pattern/levenshtein (bind.rs:214, matcher.rs:248); add a match:+hurl: | skeleton to the help text for the zero-match arm only (an ambiguous step, bind.rs:99, must not get a stub).

#10 SARIF — from SARIF’s rise for static/validation findings inline in PRs. Dry-run diags are a structured Vec<Diag> before miette (diag.rs:123, front.rs:64). Diag maps ~1:1 to a SARIF result; the closed code set (one per tests/errors/ dir) pre-fills rules[]. cli-edge serializer, no core change.

#14 --seed — from seeded-faker reproducibility. Fakes already deterministic (fake.rs:12 SplitMix64/FNV, seeded fnv1a(run_id) ^ …), seed already recorded in RunStarted{run_id} (event.rs:26). Design fork: alias run_id (zero core change, but must stay uuid-parseable for JUnit) vs a dedicated recorded seed field (cleaner, but a second knob — resolve per one-canonical-way). Known limit: the occurrence counter is scoped per scenario (lower.rs:69-86, Refs::fakes, threaded through resolve()’s fakes: &mut usize parameter, not reset per call) → cross-step uniqueness within a scenario now holds, but cross-scenario uniqueness still does not: two different scenarios each resolving ${fake:X} at the same position in their own step order get the same value. Document before advertising “unique fakes” — it means per-scenario, not per-run or per-entity.

#7 Dead-macro report — from Cucumber’s usage formatter marking UNUSED. Binding records BoundStep.macro_name (bind.rs:192); iterate front.features[].scenarios[].bound.steps[].macro_name vs packs.macros.keys() (pack/mod.rs:116). use:-only macros (pattern:None) need reachability via the use: graph (pack/validate.rs:628) to avoid false “unused”. Report the whole corpus.

#6 HTML report — from Cucumber/Karate/hurl HTML reports; the industry’s convergence on the Cucumber-Messages/JSONL stream. Core render_html(&[Event]) -> String, cli writes; best as post-hoc over events.jsonl (explain.rs already replays it) so historical runs render. Events are pre-redacted at the sink (report.rs:124).

#12 proef diff — from Allure history / test-observability-without-OTel. Identity is (file, scenario) (report.rs:147 ScenarioKey; ADR-0008 added file for exactly this). Key step diffs on text, not the volatile line. attempts+duration_ms make it a flakiness/perf-regression detector. run_id is uuid-v7 → chronology recoverable.

#4 Boolean tags — from Cucumber tag expressions (and/or/not/()). Single filter fn tag_selected (front.rs:388), three callers. value_delimiter=',' (main.rs:57) blocks the operator syntax → drop it, take one expression string, replace the CSV contract. Grammar/evaluator → core (deterministic, fuzz-shaped). Preserve empty-match=exit-2.

#8 Rerun-only-failures — from Cucumber’s rerun formatter (@rerun.txt). explain already discovers + replays the latest record and extracts failed identities (explain.rs:63). Add a multi-identity predicate reusing build_specs (exec.rs:289). Empty failure set → reuse no_scenarios_matched (exit 2), never silent-pass.

#15 Quarantine — from the flaky-quarantine-with-owner+expiry consensus. Thread gating:bool from CLI (which sees scenario.lowered.tags, exec.rs:318) into ScenarioSpec/ScenarioOutcome; exit_code() (runner.rs:89) skips non-gating outcomes and stays pure in core. Events unchanged → scenario still reported. Extend the cli.rs/execute.rs exit-code assertions; a quarantined Fault::System still exits 3.

#3b Flaky-failure detail — from JUnit <flakyFailure> / Allure retries. Needs an additive attempt_details on StepFinished (ADR-0008 additive-only) and engine-hurl collecting per-retry messages (hurl retry is per-entry internal — verify the adapter isn’t already discarding earlier bodies).

#11 proef lsp — from Cucumber’s language server (unbound-step diagnostics, go-to-def, completion); a gap Karate never closed. Reuses feature::parse, bind, pack::load, resolve Probe mode, matcher — all headless, all with stable codes + byte-offset spans that already map to editor ranges. New work: sync lsp-server (tokio banned), byte→token API, and a “collect diags, don’t early-return” wrapper (bind/load are all-or-nothing today).

#13 Impacted-only re-run — from selective/affected-test re-run. Premise broken: no reusable input hash (ADR-0010 is a byte-identity assert_eq! on outputs, execute.rs:188), and the honest fingerprint (emitted .hurl) is not run-to-run stable because runtime globals include run_id (execute.rs:304). Gated on P3; a hash miss must never skip a scenario that would fail (needs --force/first-run fallback).

10. Karate feature ledger — considered / adopted / rejected

Karate (github.com/karatelabs/karate) was the closest competitor and the deepest research stream. This ledger makes the “considered → decision” trail explicit, so each Karate idea is an intentional call rather than an omission.

Adopted — drove a plan item or the positioning:

Karate featureproef outcome
Inline fuzzy markers (match response == { id: '#uuid', age: '#number' })#2 assertion cookbook — the same readability via hurl 8.0’s native predicates (isUuid, isIsoDate, …) surfaced in docs, not a new marker DSL.
Weak/immature IDE support (Karate’s own gap)#11 LSP — reframed as a differentiator proef can win, since Karate never closed it.
HTML report with a timeline view#6 self-contained HTML report (with Cucumber’s and hurl’s).
Three-language cognitive load (Gherkin + DSL + embedded JS)proef’s headline positioning (§2): one-canonical-way, raw-hurl-only, sans-IO — the inverse of Karate’s most-cited flaw.
call / callonce cross-feature reuseAlready covered by proef’s use: / with: macro composition (ADR-0004) — no new work.

Rejected — with the reason (so it stays rejected):

Karate featureWhy not
A marker DSL (#uuid, ##optional, #? _ > 0)A second assertion mechanism competing with raw hurl predicates — violates one-canonical-way (§3). Readability comes from #2 instead.
Embedded JavaScript escape hatchConflicts with the sans-IO deterministic core and one-canonical-way — it is the thing proef exists to avoid.
Soft assertions (configure continueOnStepFailure)Conflicts with proef’s deliberate stop-at-first-failed-step model (the ∅ cascade); a failed step’s downstream is intentionally not run.
Service mocking · karate-gatling perf · UI automationPermanent non-goals (PRD §3 and §3 above).

Deferred — genuine candidates, not yet planned:

Karate featureNote
Dynamic data-driven Examples (rows from a read('data.json') array)Table-driven coverage from an external data file. Plausible, but in scope-tension with product-neutrality and the sans-IO/determinism line (an external read at lower time). Revisit if a real need appears.
match each / schema-as-a-value reuseAchievable today via the #2 cookbook’s hurl predicates and reusable expect: macros — no new engine feature needed.

11. Sources (competitive research, 2026-07-31)

12. Round 2 — post-execution competitive re-review (2026-08-02)

Round 1 (§1–§11) is largely shipped (across the releases that followed). Round 2 re-ran the Karate + Cucumber + adjacent-landscape survey against the post-execution codebase, then put every surviving candidate through a four-stream deep code-validation pass (matcher/binding · reporting/events · lifecycle/i18n/snapshot · scope-boundary), mirroring Round 1’s method. Identifiers N1–N9 are stable and doc-local — this registry is their only sanctioned home; they never appear in code comments (per the no-task-ids-in-source rule). File:line citations validated 2026-08-02 and will drift.

12.1 What the re-survey confirmed is already shipped (positioning to defend)

The external agents, blind to the just-landed work, flagged many “gaps” that Round 1 already closed — recording them so the omission-vs-decision trail stays explicit:

Re-flagged “gap”Already shipped as
Cucumber snippet/stub suggestion on unbound step#9 stub-gen (unbound-step diagnostic prints a paste-ready macro)
Rerun-only-failures#8 --rerun (keyed on (file, name))
Soft-fail / allow-failure tag#15 @quarantine (non-gating, still reported)
Boolean tag expressions#4 proef_core::tags (fuzzed grammar)
“retry until assert passes” (Karate retry until)macro retry: → hurl [Options] retry (retries until asserts pass or budget ends)
Data-table → step argumentsbind.rs merges | key | value | rows into macro args
One trial per Examples rowoutline expansion → ScenarioDef per row → one harness Trial
Explicit skipped/pending statusStatus::Skipped (post-failure steps) + Warned (optional)
Whole-run JSONL event record; git-native plain text; JUnit/SARIF/HTMLADR-0008 event spine; .feature+YAML+proef.toml; the reporter family

Convergent-evolution note (validates the architecture, nothing to adopt): Karate v2’s karate-events.jsonl (2025) is proef’s ADR-0008 event stream re-invented; Bruno’s plain-text git-native rise is proef’s text model; both confirm the design is industry-aligned.

12.2 The validated Round-2 roadmap (master table)

Verdict legend as §5 (✅ FITS · ⚠️ NEEDS-ADAPTATION · 🚫 rejected/premise-broken). Effort: S ≤ ~1 day · M ~days · L ~weeks.

#ItemVerdictLives inEffortArchitectural truth (validated 2026-08-02)
N2Run-level SLA thresholds (p95/max(duration) gate)✅proef.toml [sla] + cli exec.rsS–MStrongest — zero schema change. duration_ms already on StepFinished (event.rs:74) and in the record; a pure CLI fold. Config as [sla] (env-overridable, ADR-0012), not a flag. Breach = TestFailure (exit 1) folded before exec.rs:361; malformed table = exit 2; no new exit code. Must be opt-in by presence of [sla] — absent, behaviour is byte-identical, so pinned exit-0 tests + reference snapshot are untouched. Distinct from hurl per-request duration < (aggregate vs per-entry) → keep SLA aggregate-only, one home.
N6aHTML per-scenario timing waterfall✅core html.rsSZero schema change. Step start-offset = cumulative sum of prior duration_ms in the scenario, width = own duration_ms; new render in html.rs:176. Cannot show cross-worker occupancy (no clock/worker id) — intra-scenario only. Ship this first.
N8i18n # language: — verify, fix, test, keep the claim⚠️core feature.rs + testsSClaim asserted twice (PRD.md:104, TECH-SPEC.md:110) but unverified. gherkin-0.16 honours the header transparently; proef strips no keywords. One English-only bug: feature.rs:167 detects outlines via keyword.contains("Outline"/"Template") — false under any dialect (fr Plan du scénario, de Szenariogrundriss). Blast radius small (only a no-Examples malformed outline degrades). Fix: detect via !examples.is_empty(); add a localized fixture + byte-span test. Keep the docs claim — fix, don’t retract.
N9Curated expect: shape-macro library (expectUuid, expectIsoDate, expectNonEmptyList…)✅helpers/*.yaml + docsSNew — surfaced by validation. Augments the existing expect:/MergedAsserts mechanism (step.rs:60, emitted emit.rs:192); zero engine/core change; product-neutral (generic shapes only). Same lever as the §12.5-B ergonomic uplift. Narrows the deep-equality ergonomic gap — not the semantic one (§12.4).
N7Near-duplicate macro lint (extend proef macros)⚠️core sim-fn + cli commands.rsS–MAbsent; reuse literal_skeleton (matcher.rs:236) + levenshtein (matcher.rs:260). Extend the shipped dead-macro report (commands.rs:270), not a load pass (those are hard errors). Tight heuristic — skeleton-equal-modulo-captures — or it false-positives on the shipped corpus (boardShows* family; activateChannel “…and ready”). Advisory JSON field beside unused (commands.rs:306), exit 0, never a gate. Drop the conjunction + “organize-by-domain” sub-lints (false-positive on shipped prose; no machine model of “domain”).
N4TAP reporter⚠️cli new tap.rs via --output tapMValid, but the “surface hurl’s native TAP” rationale is wrong — proef calls run_entries in-process, never shells out; TAP must derive from the event spine (scenario = test point), like every reporter. Live Reporter (report.rs:120) → inherits sink redaction. Plan count from exec.rs:194. @quarantine → # TODO needs the non_gating set injected (it’s computed at exec.rs:313, not in the stream). One surface: --output tap (reuses the stdout-ownership machinery), never also a proef tap replay.
N1Typed parameter types in the matcher ({int}/{uuid}/custom, bind-time)⚠️matcher/bind in coreMGenuinely absent (captures are untyped strings, matcher.rs:221; params: Vec<String>). Must be declaration-site, not inline {name:type}: 3 of 4 arg sources aren’t captures (data-table, defaults, with:, and use:-only macros have no pattern), and inline typing is invisible to proef schema. One-canonical forces a single params spelling (params: {q: uuid}, bare = any) via custom Deserialize → breaking pack migration → needs an ADR. Model on the fake::GENERATORS typed registry. Two-tier caveat: skip validation when the raw arg contains ${/{{ (resolves later) → a best-effort literal-args lint, not a type system. Diagnostic proef::bind::param_type_mismatch at bind.rs:193 (+ defaults at validate.rs:57, with: at validate.rs:353).
N3Suite-level setup/teardown (once-before / once-after)⚠️proef.toml [run] + cli exec.rsMReal gap (only per-scenario Background; teardown is engine-internal session.finish()). Premise correction: tags never reach the core runner — ScenarioSpec/ScenarioOutcome carry no tags/gating (runner.rs:30,131); quarantine is a CLI-edge non_gating set (exec.rs:313). So use a proef.toml [run] setup/teardown construct, not a tag (a tag would entangle with --tags/--rerun/flows/dedup + undefined ordering). Orchestrate in execute() around runner::run (exec.rs:220); state crosses only via saveAs: global, which must merge before the parallel pool snapshots the store (runner.rs:439). Explicit failure short-circuit required — an assert-failed setup is fault:None→exit 1 and would not abort the pool (cascading failures on un-seeded state); refuse to launch and surface the fault. Excluded from build_specs so it never double-runs; --dry-run unaffected (never calls runner::run).
N5Golden response snapshots🚫—LRejected. Response bodies exist but are discarded (HurlResult…calls[].response.body; session reads only captures/errors). It is a second assertion mechanism competing with hurl body-asserts + expect:, and whole-body regression is already proef diff’s job. No normalization machinery exists (sink redaction masks only known injected values, report.rs:29). Secret-leak risk: backend-minted tokens/PII would be committed unredacted — against ADR-0005 (session.rs:346 already refuses to persist a capture equal to a secret). Do not build. If ever needed: diff-time over run records, never committed goldens.

12.3 Verdict-change ledger (validation overturned the first sketch)

The deep pass is on the record because it changed conclusions — the point of validating:

ItemFirst sketchAfter validationWhy
N5 golden snapshots⚠️ candidate🚫 rejectedDuplicates hurl asserts + diff; no normalization; leaks backend-minted secrets
N4 TAP rationale“surface hurl’s native TAP”corrected: derive from the event spineproef never shells out; hurl --report-tap is unreachable + per-file, not per-scenario
N3 selector@setup/@teardown tagproef.toml [run] constructtags never reach the core runner; a tag overloads “filter” with “phase” and races the pool
N6 timelineone “timeline” itemsplit N6a (zero-schema, now) / N6b (injected timestamp_ms, later)true cross-worker occupancy needs an injected clock/worker id
N1 typed params“small matcher tweak”M + ADR + breaking migrationdeclaration-site forced by schema coherence; single spelling forced by one-canonical; literal-args-only forced by two-tier vars

12.4 The named architectural ceiling (accepted, not a defect)

Karate’s match response == { id:'#uuid', items:'#[]' } — order-insensitive whole-body deep-equality with type-holes, exhaustive-key checking, and one readable structural diff — cannot be assembled under hurl-only (asserts are path-at-a-time). Reusable expect: macros over hurl jsonpath cover per-path type/value/shape, collection membership (contains), cardinality (count), optional keys (exists/not exists), and RFC-9535 filtered queries (AUTHORING.md:100) — the ergonomic gap, narrowed further by N9. The semantic gap (single order-insensitive whole-body diff + exhaustiveness) stays open by design. The Round-1 marker-DSL rejection stands (§3, ledger §10 — a second assertion mechanism). This is a deliberate ceiling of the hurl-only bet, stated honestly, not engineered away.

12.5 Non-goal-adjacent — explicit scope decisions (keep excluded absent an ADR)

The two biggest capabilities a market reviewer would name are on/over the PRD §3 line. The governing boundary the validation extracted: CLI-edge IO that injects values into the sans-IO core is sanctioned (ADR-0012); IO that re-shapes the corpus or acts as a recurring oracle is contract testing (out).

  • A — OpenAPI → scenario generator (proef generate). Verdict: needs an ADR; default = deferred/out-of-scope. Strict generate-then-freeze clears sans-IO/determinism (ADR-0012 precedent) and echoes #9 stub-gen at suite granularity — but the bright line is “the spec may be a one-shot seed; it may never become a recurring oracle.” It sits one --check flag from OpenAPI-drift (§3 non-goal), introduces a new inward generation direction, and pressures one-canonical-way on regeneration (a second maintenance path). Only an ADR that bans the oracle/drift mode and accepts the OpenAPI dependency can green-light even the narrow scaffolder. → Now settled in ADR-0016 (Proposed): the oracle/drift mode is permanently rejected; the narrow one-shot scaffolder is deferred (output-quality + dependency cost) but buildable later under the bright line.
  • B — JSON-Schema conformance assert. Verdict: shape/type conformance is already-achievable today via expect: + hurl type predicates (the #2 cookbook — zero new features); the ergonomic uplift is N9 (curated shape macros). Full external .schema.json whole-body validation is out — hurl has no jsonschema predicate, and using the API’s canonical schema as oracle is drift-detection.

12.6 Prerequisites & one-canonical-way watch-list (round 2)

  • P4 — injected per-event timestamp_ms/worker (Option, skip_serializing_if), stamped by a CLI sink-wrapper on the worker thread (the run_id injection pattern), left None by the sans-IO core; kept off RunStarted (exact-bytes pin event.rs:157). Note: old records still parse (additive holds), but the new-run reference snapshot changes and needs a new insta filter + deliberate review. Unlocks N6b.
  • ADR needed: N1 (params-shape migration + literal-args-only semantics); Tier-3-A OpenAPI (bans the oracle mode).

One-canonical watch-list — each must fold into an existing mechanism, never ship beside it:

ItemMust replace / augment (not duplicate)
N1 typed paramsOne params spelling (name→type map); no second inline {name:type} form
N3 setup/teardownExactly one [run] setup + one teardown; not a tag, not a second Background
N4 TAPOne surface (--output tap); no parallel proef tap replay
N9 shape macrosAugment the existing expect: mechanism; never a schema/marker DSL
N2 SLAAggregate run/scenario budget only; per-request latency stays hurl duration <
  1. Batch E — free / small / zero-schema — SHIPPED 2026-08-02: N2 SLA (opt-in) · N6a waterfall · N9 expect: library · N8 i18n verify+harden.
  2. Batch F — small–medium — SHIPPED 2026-08-02: N7 near-duplicate lint · N4 TAP (--output tap).
  3. Batch G — medium, design/ADR call first — mostly SHIPPED 2026-08-02: N6b full timeline (ADR-0015, P4 delivered) · N3 setup/teardown (ADR-0014, [run] setup/teardown). N1 typed params — deferred (ADR-0013): the research-grounded call given proef’s deferred-heavy, string-ish corpus.
  4. Blocked pending an ADR: Tier-3-A OpenAPI generator. Rejected: N5 golden snapshots.

12.8 Sources (round 2)

Competitor sources unchanged from §11 (Karate/Cucumber/Hurl/Bruno/Schemathesis/Pact/k6/ Playwright). Round-2 findings are code-internal — every verdict is anchored to a file:line validated 2026-08-02, not to an external claim.

proef — open findings

This is the worklist. Every open defect and gap lives here, whichever review found it. Each entry is self-contained: the evidence, the reasoning, and — where something was declined — why.

Companion: IMPROVEMENT-PLAN is the feature roadmap (own numbering, 14 of 16 shipped) and stays separate because five ADRs cite it by section number. CHANGELOG records what shipped, per release.

Provenance. Three reviews fed this list, each validated claim-by-claim against the tree and then retired into it:

ReviewScopeContributed
v0.5.3 external (2026-08-06)40 claims → 38 confirmed, 1 partial, 1 already fixedthe A/B/P/Q items below
first-run UX (0.5.3, engineer’s first 30 min)F1–F4R1–R2
non-technical UX (0.8.0, PRD §4 P1 calibration)N1–N5R3
round-7 pre-merge review of PR #134 defects + residue§2.1–§2.4 below; three shipped in #31
corpus-port report (0.8.0, a real 844-line hurl suite ported)12 items → 3 shipped (#41, #43), 2 premises correctedM/E/D items below

The review documents themselves were removed once their open items landed here; their full text, transcripts and citations are in git history (git log --diff-filter=D -- docs/FIRST-RUN-UX-REVIEW.md docs/NON-TECHNICAL-UX-REVIEW.md). The shipped/open split was re-checked against main on 2026-09-11 (after 0.19.0, which closed H3, H4 and H5).

Read the citations as “start reading here”, not as addresses. They were accurate on 2026-08-06 and files have moved since; locate symbols with rg, not line numbers.


Noted after the 0.18.0 release (2026-09-10)

The exit-130 interrupt test asserts a race nothing holds open (open)

a_second_interrupt_hard_exits_with_130 (crates/proef-cli/tests/execute.rs) failed once on gates (ubuntu-latest) and then passed on a re-run of the same commit with no change: run 34341778587, attempt 1 red, attempt 2 green, both at a802bfe. The diff under test was documentation only, so it cannot have been a regression. Filed because a flake that is only ever re-run is a flake nobody is counting — and because the test is young, added by #168 as the first assertion anywhere on exit 130.

Verified from the failed attempt, not inferred. The panic carries the child’s stderr, and the interrupt notice is in it:

stderr:

interrupt — cancelling after current batches (a second interrupt hard-exits)

  left: Some(1)
 right: Some(130)

So the sequencing the test is built around worked — the first signal landed and the handler announced itself — and the process still exited 1, the graceful cancelled code, rather than 130. Two further facts bound what can have happened. nextest timed the whole test at 57 ms; the scenario’s only request is GET /slow, which the fixture answers after a deliberate sleep(5s) (proef-fixture/src/lib.rs — the one documented exception to that server’s own “never sleep-raced” rule). A 57 ms test never waited on that sleep. And the banner the test synchronizes on, running N scenario(s), is written at exec.rs:661 — before runner::run is called at exec.rs:680.

The most-supported reading, to be confirmed by a Linux reproduction rather than assumed: the banner proves the run started, not that a batch is in flight. Cancellation is cooperative at batch boundaries (ADR-0007), so a first signal that wins the race against dispatch has nothing to wait for — the pool starts already-cancelled, the scenarios record as skipped, the record closes and the process exits, all inside the time it takes the test to spawn an external kill(1) for the second signal. The window the test needs is the 5-second sleep; on that run the window never opened.

Why this is not just one red run. TESTING-STRATEGY §5 already names the rule — assert normalized event order, never raw interleaving, and treat wall time only as a generous upper bound — but it says parallel tests, so a signal-delivery test sits outside its letter while squarely inside its intent.

The fix shape, not applied here. The assertion is worth keeping (nothing else pins exit 130), so the answer is to make the window deterministic rather than to weaken it: the test needs a synchronization point proving the request reached the fixture, not that the run began, so the 5-second sleep is genuinely in flight when the first signal arrives. That is a fixture and test change on the one path that exercises the second-signal escape hatch; it wants its own change, and a Linux reproduction first — macOS has not reproduced it, and this is exactly the trap P5’s atomic-save half is held away from (“do not chase it on a Mac — that is how it gets fixed by coincidence”).


Ingested — the 0.18 survey (2026-09-06), validated then implemented

A check-the-world round over the CI-consumer surfaces — delivery failures, signals, staging, the ADR-0007 budget family, the redaction boundary, the machine sinks, and the flaky predicates — executed as waves A–F (#168–#175) and then swept by a /simplify pass (#176, #178–#179). Recorded like every external round so its verdicts are not re-derived.

Premises that did not survive validation — do not re-raise as filed:

  • “The flaky predicates are missing.” The hardest one, broken≠flaky, already existed in flaky.rs, with transition-counting, latent, and the quarantine lifecycle. Four gaps remained (a sample floor, hysteresis, an outage guard, an input equivalence class) — all pure folds over the retained history, no new state, advisory by design. Designed first (#174, closed unmerged once approved; its decisions: the fingerprint is the default key, new → insufficient-data is a MINOR break, and ADR-0020 takes a clarification rather than a new ADR), then shipped as #175.
  • “Group flakiness by git tree SHA.” Collides with ADR-0020 (proef never harvests git state). Split along the ADR’s own axis: a proef-computed input fingerprint (inputs.json — feature sources + loaded macros/fragments
    • the resolved [url]/[vars] scope; a sidecar, so ADR-0008’s schema stays frozen where the design had proposed a run_started field) plus a handed-over commit via --meta commit=… and --by commit.
  • “Five sinks leak secrets.” They bypassed the masker for identity fields only; secrets lower to {{name}} and the engine pre-redacts details, so no live leak was found. An unenforced boundary, not a leak — closed per sink (#171), then made structural (Redactions::apply_outcome, #178).
  • “Adopt cargo-auditable, attestations, machete.” All already in place (release.yml, just gates); the genuine gap was on-demand coverage (#173).

Shipped: #168 (the record’s own write failure and the GitHub summary’s reach exit 3; SIGTERM/SIGHUP graceful; a second signal exits 130 without printing; a UUIDv5 JUnit identity for a custom --run-id) · #169 (staging beside the file the parser read — the feature-side twin of H5, updated in place below; --sarif lines from the carried source; the symlink and case-insensitive edges; a 120-byte slug cap) · #170 (ADR-0007 amendment: max-time:/retry-interval: capped, a four-hour batch ceiling, timeout-ms = 0 refused) · #171 (every sink masks identities; lsp --env; the CTRF key set pinned; three tests de-flaked) · #172 (warned/cancelled in --format json, JUnit and CTRF; tags::reserved_tag_typo) · #173 (just cover, the ratchet policy) · #175 (the flaky guards).

Deferred, with dispositions:

  • Restoring the Given/When/Then keyword and the Rule name into the reporters — schema-additive, but it moves the pinned event snapshot: its own review, unscheduled.
  • Removing the two fake-generator aliases — a breaking change; bundle it with the next MINOR that already breaks.
  • secret list --format json and macros --check — surfaces the survey wanted and nothing yet needs; build on a request.
  • A cargo-mutants CI job, an llvm-cov + coverage-service job, and the immutable-releases repository setting — cannot be validated without triggering CI, and their cadence and cost are a maintainer’s decision. The coverage job, when it lands, must be a ratchet (TESTING-STRATEGY §3).

Noted while simplifying, not filed: the quarantine-match closure is spelled three times (tap, ctrf, ci_reports) — one helper would do; and apply_outcome clones an outcome even when the needle set is empty, a Cow/is_empty short-circuit away from free on a secret-free run.


Ingested — the hurl-coverage audit (2026-09-05), validated claim-by-claim

The question: can every hurl test case now be wrapped in Gherkin? Answered by enumerating hurl 8.0.1’s own surface from its AST — 8 section kinds, 7 body byte kinds (5 multiline variants incl. GraphQL), 42 options — and checking each against both body forms. Coverage is near-total by construction: the fragment scanner implements hurl’s own Visitor, so it has no per-construct enumeration to fall out of date, and only 4 of the 42 options are constrained at all (the ADR-0007 budget rules on retry/repeat/delay/retry-interval).

Shipped: the two defects the audit found — a file,…; body in a ref: fragment was unresolvable, and staged assets collided across scenarios. See the ADR-0018 amendment of the same date.

A premise of the audit’s own first pass that did not survive validation. Path-valued options were reported as sharing the file-body defect. They do not: runner/options.rs never consults context_dir, so cacert, client-cert, client-key and netrc-file reach curl as raw CWD-relative strings under both runners, identically. [Options] output: is context-dir mediated (runner/output.rs), via a later path than the option table — which is what made the first reading look right.

Open — the two remaining gaps are by design, and stay that way

H1. One ref: names one entry. A .hurl file is usually one test case spanning several chained entries; wrapping it means annotating each entry and writing one ref: step per entry. Deliberate (ADR-0018 fixes the annotation at one entry, permanently) and not silent: proef fragments prints UNANNOTATED — not referenceable per entry with its line, and --require-annotated exits 1. Declined rather than open: a multi-entry ref: would have to decide where the run ends, which is the orchestration ADR-0018 keeps in YAML.

H2. A cross-entry [Options] variable: does not carry into a fragment. hurl’s variable: assigns into one shared set that persists forward, so a corpus file whose first entry declares variable: term=ok and whose second reads {{term}} runs standalone but is refused by proef at --dry-run (proef::lower::unbound_placeholder). Correct as it stands: a fragment is independently runnable by definition, so its inputs must be satisfiable without a neighbour having run first. The refusal is early, names the variable, and offers the fix that preserves standalone runnability — give the fragment its own [Options] variable:. Worth a porting note in AUTHORING if adopters hit it; not worth weakening the check.

H3. --dry-run does not notice a missing file,…; asset. Verified: a suite whose asset has been deleted still reports dry-run OK. A missing file is statically knowable and --dry-run is the gate CI runs before standing an environment up, so catching it there is the right end state. Deliberately not done in the same change as the staging fix, for two reasons worth writing down. dry_run has its own path and never calls build_specs, so the check would be second code walking artifacts for assets — which “one way to do one thing” says should instead be one shared checker both paths call. And the reference corpus is run from temp working directories with settings passed by environment (TESTING-STRATEGY), so a new filesystem requirement at validation time needs its own regression pass over those tests before it can be trusted. Until then the run-time failure is early (before the request is sent), names the file and the directory it was sought in, and cannot be reached silently.

Closed 2026-09-11, on both of the conditions this entry set. The check is one function with two callers, not a second walker: stage_assets split into assets::resolve_assets — every refusal that is statically knowable, and no destination touched — plus the copy, and --dry-run calls the first half. Staging and validation cannot disagree about whether a suite’s assets resolve, because they are the same code.

And the regression pass this entry asked for came back clean without needing anything: all 712 tests pass, including the reference-corpus suites that run from temp working directories with settings passed by environment. The reason is the other H-item — since the feature-side twin of H5 landed, staging resolves against LoadedFeature::read_from, the path the parser actually read, so a new filesystem requirement at validation time does not inherit a cwd-dependency. The concern was correct when it was written and had been retired by a change filed under a different number.

H4. The file,…; scan in proef-core is hurl grammar the grammar guard cannot see. emit::file_refs_in finds asset references by scanning for the literal "file," and a closing ;. That is engine syntax living in core, and source_guards.rs::hurl_grammar_in_core_is_the_closed_set_the_adr_names does not catch it: engine_grammar_kind classifies fences, HTTP, [Section] headers, method lines and key: value options, and a body reference matches none of those — so the literal is neither on the sanctioned list nor detected as missing from it. Pre-existing, not introduced by the staging change (git show confirms the scan body is byte-identical to the former file_references), which is why it was not fixed alongside it. Two ways out, both real work: widen engine_grammar_kind so the set is closed over the shapes ADR-0002 names rather than the shapes the guard happens to classify — the same correction the method-line arm already records — or move the scan behind the seam, where proef-engine-hurl’s Visitor already reads filenames from hurl’s own AST (fragment.rs, visit_filename). The second is the ADR-0002 answer; it needs a StepKindSpec entry beside validate, fragments and options, and hurl_core supplies the hooks for it already (visit_file for Bytes::File, visit_filename_param/visit_filename_value for multipart parts — hurl_core-8.0.1/src/ast/visit.rs).

Closed 2026-09-10 — by the second way, and the first way as well. The scan is now StepKindSpec::assets, a fourth engine hook beside validate, fragments and options; emit() takes &[StepKindSpec] to reach it and FrontEnd carries kinds beside the kind_to_engine table registry already documents as a pair that must not be re-derived apart. The guard was widened too, rather than left blind because nothing currently trips it: a bare lowercase keyword followed by a comma (file,, hex,, base64,) is now classified, and planting the literal back in emit.rs fails with body "file," in emit.rs. Two things this entry predicted came true on contact. The engine hooks are exactly the two named above — and the tempting third, visit_filename, is the wrong one: hurl routes the [Options] file paths through it, and output: names a file the run writes, so staging it would demand a source that cannot exist. And the AST reading fixed the defect this entry recorded as a consequence: file, inside a JSON body is no longer an asset. What this entry did not anticipate is that ADR-0002’s amendment had miscounted — it says thirteen literals, and this was the fourteenth, missing for precisely the reason that amendment had already written down about the method line.

The second consequence recorded below stands unchanged: collect_assets still inspects only StepPayload::HurlEntries, never Structured. Recognition is now the engine’s, but which payload variants carry assets at all is still core’s assumption.

Two consequences of the text scan worth recording with it. It cannot tell a real file,…; body from the same six characters inside a JSON or text assertion body. And collect_assets only inspects StepPayload::HurlEntries, never StepPayload::Structured — the variant reserved for a future non-hurl engine — so the root (assets/<slug>/, per scenario) generalizes while the recognition of what belongs in it does not. ADR-0002’s acceptance test (“adding an engine leaves proef-core diff-empty”) is what would catch that, and the seam above is what would satisfy it.

H5 — updated 2026-09-07. The 0.18 survey found and reproduced the feature-side twin of this finding, worse than the fragment side it records: the feature’s staging root was parent_dir(portable name) resolved against the cwd, so a typed-absolute or config-written suite path run from any subdirectory failed staging with exit 2 (a name’s anchor — project root, or as-typed — is not recoverable from the string). Closed by exactly the fix this entry prescribes, applied to the feature side: the resolved discovery path travels beside the name (LoadedFeature::read_from) and staging is a lookup, not a re-parse. The fragment side below still resolves by name-join (correct while both are seeded from config.root(), per the original analysis) and this entry stays open for it.

H5. A fragment’s directory is re-derived from its display name, inverting SourceNaming without its canonicalize fallback. assets.rs::AssetRoots:: source_dir turns a recorded file.hurl#name back into a directory by splitting the qualifier and joining against the project root. But that name is produced once, at what the codebase calls the naming boundary (front::read_corpus → naming.name(&path)), and SourceNaming::relative is more than a strip: it falls back to comparing canonical forms precisely because a lexical-only version already shipped a bug (a suite reached through a symlink — macOS /tmp → /private/tmp — silently failed to match, R11-9). The inverse here has no such fallback. The two agree today because both are seeded from config.root() and discovery walks from that same root, so only the lexical case is exercised; nothing enforces that they stay inverses, and AssetRoots’ unit tests hand-build the struct rather than going through a real SourceNaming. The deeper fix is to carry the resolved source directory through the data model — Fragment/ScannedFragment holding the real PathBuf beside file: String, threaded onto AssetRef — so staging is a lookup rather than a re-parse. Not done here because it is a data-model change across three crates, and because the record must keep carrying the portable name: the resolved path would have to travel beside it, never replace it.

Closed 2026-09-11 — and it was one crate, not three. The prescription above aimed the change at Fragment/ScannedFragment/AssetRef, which would have put host paths into proef-core. The feature side had already answered this differently and better: FeatureFile.path (core) carries the portable name and LoadedFeature::read_from (CLI) carries the IO path beside it. Doing the fragment side the same way keeps core untouched and makes the twins symmetric — front::CorpusDirs records the directory each fragment file was read from, at the naming boundary where both the name and the path are in hand, and FrontEnd carries it beside kinds. AssetRoots::source_dir is a lookup; there is no inverse left to drift.

Two things fell out. AssetRoots loses its project field and build_specs its project_root argument — with nothing recomputed, the project root was staging’s business only as the join’s left-hand side. And a fragment the corpus never read is now a named error rather than a directory guessed from its name; it is unreachable from a loaded suite, which is exactly why the old code’s silent guess would never have been noticed.

Both new tests were checked against the old resolution and fail under it. The regression test is deliberately a case the join gets wrong rather than a symlink reproduction, because this entry is right that the two resolutions agree on every path a suite takes today: the defect was that nothing held them together, not that they had already come apart.


Ingested — validation round 19 (2026-09-02), validated claim-by-claim

An external round against v0.15.0+v0.16.0 (66 commits). Every finding was reproduced against the tree before being acted on, and the round’s own correction of two earlier rounds (the Rust pin) is accepted — see below.

Shipped: the P1 and all eight P2s. --rerun on a truncated record (a silent green over a suite that never ran); the artifact slug collision (ADR-0010, silent overwrite); diff‘s phantom “now skipped (was passing)”; the tab exempted from the control-character guard; the unreachable Warned scenario status and its four dead consumers; rerun composition (headline vs page, and a non-transitive overlay); --shard-weights’ zero pileup; the two ADR-0020 §5 metadata consumers that never received any; and the ADR-0002 grammar guard’s blind shapes — which, once taught method lines, surfaced exactly the token the report predicted.

Two P2 sub-claims declined, with reasons:

  • Header lines in the grammar guard. The report names method and header lines as undetectable. Method lines were taught and found a real token. Header lines were not: no instance exists in core today, and the only workable heuristic (a Capitalized key with a colon) fires on ordinary diagnostic prose. Trigger to revisit: the first header literal that appears in core — at which point it should be pinned by hand rather than by pattern.

  • production_text truncation was latent, not active. The report calls it “already the shape of html.rs and pack/validate.rs”. Checked: both do carry a second #[cfg(test)] mod, but neither has production code after one, so nothing was actually unscanned. Fixed anyway (the scan now excises every test module) because it was one edit away from real.

P3s — shipped

.cargo/audit.toml’s stale quick-xml ignores (it claimed to mirror deny.toml, which had deliberately removed them — the lockfile is on the patched 0.41.0 line, so the nightly job was suppressing for no reason, and would have silenced any new advisory against that line); explain dropping a step’s authored name: while step_label’s own doc enumerates explain among its six readers; the HTML “Slowest” section counting [run] phases into “% of run time” while the tag table on the same page excludes them (ADR-0014); the toolchain policy stated correctly in RELEASING.md/CLAUDE.md but not in the normative spec that rust-toolchain.toml cites as its authority; and five stale --output json spellings in documents describing current behaviour, now guarded — narrowly, by an allowlist of present-tense docs, because CHANGELOG/RELEASING/this file quote the flag as it really was.

P3 — closed by ADR-0021

  • --run-id records are invisible to latest, flaky, diff and --rerun (closed 2026-09-02 — ADR-0021, the decision this entry asked for). Split along the risk rather than the file: rotation keeps the uuid predicate (is_rotatable), discovery asks whether a directory holds an events.jsonl (holds_a_record), and ordering follows the uuid’s own embedded timestamp, falling back to directory mtime for a custom id. Not the record’s run_started, which is what this entry and the ADR’s first draft both proposed: the head event carries event/run_id/schema and no time at all, so there was nothing there to read. The analysis below stands as the reasoning; it is kept because the tradeoff it names is what the ADR decides, not because the item is open.

  • --run-id records are invisible to latest, flaky, diff and --rerun. record::all_runs filters on fsutil::is_run_id, which requires a 36-character uuid, so a --run-id pr directory (which TROUBLESHOOTING demonstrates) is never enumerated.

    Do not “just widen it”. rotate_runs consumes the same predicate and says so in its own words — “all_runs is the one answer to what is a run record here” — and its narrowness is what keeps rotation from deleting user content under runs-dir = ".". Broadening the shared predicate broadens deletion. The two uses have opposite risk profiles: discovery is unsafe when narrow, rotation is unsafe when broad.

    So the fix is to split them, which contradicts an explicit design statement and therefore wants a decision on the record. A safe discovery predicate exists (a directory containing events.jsonl cannot be mistaken for target/ and deletes nothing), but ordering does not come free: all_runs documents that uuid-v7 names sort chronologically, so lexical order is time order — a pr directory breaks that, and latest would need mtime or the record’s own run_started. CONFIG.md documents the rotation consequence of custom ids; it does not document the invisibility. That gap is real either way.

Noted while reviewing ADR-0021 — recorded, not scheduled

  • One doc check is still in the binary-half’s file. docs.rs states its own charter — it holds the checks that need a built binary, because they ask clap rather than parsing help text — and xtask docs-check states the mirror rule for the checks that only read files. Three of the four tests left in docs.rs genuinely need the binary; no_current_behaviour_doc_spells_a_format_as_an_output_path reads files and nothing else, so it belongs in docs_check() beside check_examples and check_links. Consequence, the same one that moved the changelog check: it never runs in the fast doc-only CI step. Not moved with that one because it depends on collect_markdown and the DESCRIBES_TODAY allowlist, both local to docs.rs — porting them is a real change, not a relocation, and it earns its own. Closed 2026-09-10: ported to xtask docs-check as check_output_path_spelling, and cheaper than this entry expected — living_docs() already collects the ADRs, so collect_markdown was deleted rather than ported and “which files are documentation” stays one answer. Only DESCRIBES_TODAY moved. The shrink guard was tightened in the move: it had counted ADRs into the same total, so checked >= DESCRIBES_TODAY.len() could be satisfied by docs/adr alone, masking the one failure it exists to catch.

  • --rerun reads the base record’s events.jsonl twice. exec.rs calls record::read_events(&dir) for the JUnit overlay, then record::rerun_candidates(&dir), which calls read_record → read_events on the same directory. Two full reads and two full deserializations of one file, bounded only by the 256 MiB record ceiling. Pre-dates ADR-0021 and is untouched by it. The fix is small and shaped like the rest of the module — rerun_candidates takes &[Event] rather than a &Path, and the one caller passes the events it already has — but it is a signature change on a path --rerun alone exercises, so it wants its own change, not a ride on this one. Closed 2026-09-10, in exactly that shape. rerun_candidates also became infallible, which surfaced a second defect this entry had not seen: the caller’s first read swallowed its error with .ok() and the second rediscovered it a line later, so which call reported a read failure was an accident of ordering. One read now, one error path.

  • Discovery now costs a second stat per custom-id run, and that population is the one nothing bounds. all_runs stats each directory once for holds_a_record; began_at then reads a uuid-v7 name’s time out of the name itself (no syscall) but falls to std::fs::metadata for any other name. Since rotation deliberately never deletes custom-id directories, [run] keep-runs does not cap that set — so a CI job minting --run-id per build pays one extra stat per historical build on every command that resolves “latest”. Accepted, not a defect: the stat is what buys correct interleaving of custom-id and uuid runs in one time order, which is the point of the ADR. Recorded because it is the one cost here that grows unbounded, and a future reader measuring a slow flaky on a long-lived runs dir should find it named.

Corrections this round made to earlier ones (accepted)

Rounds 17 and 18 reported the 1.97.1 pin as “overdue”. It was not: R18-2 changed the policy to latest stable adopted at its x.y.1 point release, and 1.98.1 does not exist yet. The round is right that the remaining defect is documentary, and right about where — the correction had reached RELEASING.md and CLAUDE.md but not TECH-SPEC §15, which rust-toolchain.toml names as its authority. Fixed in all four places.


Ingested — the 2026-09-02 survey (internal), validated then implemented

A deliberate check-the-world round: repo state against upstream releases, standards movement, and the open list itself. Recorded like every external round so its verdicts are not re-derived.

Validated as needing nothing — do not re-raise without new evidence:

  • The hurl pin is current. 8.0.1 is the latest upstream stable (2026-04-28); the canary covers the next one.
  • The Rust pin is correct per the written policy. 1.98.0 landed 2026-08-20; no 1.98.1 exists yet, and policy adopts at x.y.1 — a calendar item (~mid-September 2026), not a drift.
  • notify 9.0 is still a release candidate (rc.4, 2026-05); =8.2.0 stands.
  • Release engineering already ships the modern supply-chain story — Sigstore attestations (attest-build-provenance@v4), .sha256 sidecars, Homebrew tap, binstall metadata, SHA-pinned actions gated by pinned zizmor. The survey’s own candidate (“add attestations”) died against the tree.
  • Competitor movement is OpenAPI-generative testing (Schemathesis et al.) — a different product shape (generated negative tests vs. declared business scenarios); no charter-fit gap. The hurl-fidelity niche is uncontested.

Shipped from the survey (this series): [http] cookie-store = false (hurl 8.0’s env-shaped option; the one [http] key with no per-entry spelling at all); --ctrf (CTRF report off the JUnit fold, quarantine parity per ADR-0019, real retryAttempts); the mid-run console write failure latch (the deferred v0.6–v0.8 item, to its own written design); emit::feature_stem/emit::artifact_slug closing Q6 structurally; Q2 re-verdicted closed (the #146 cache had already closed it).

Still open from the survey, dispositions unchanged: [source-links] (build verdict of 2026-09-01, unscheduled); P13 (the CI llvm-cov job — its local half, just cover, shipped 2026-09-07 in #173); the text-scan honesty bundle (capture-name charset / ≤2-char methods / key_line_spans flag — see the deferred list); P12 (measure first, alone, per the complexity-guard lesson). Decision items untouched: E2’s split-invocation remainder (trigger not fired), E3 (wants an ADR), R1 (wants its own spec). The shipped-changelog duplicate headers (maintainer’s call) — closed 2026-09-10: no release carries a repeated kind heading any more, the regrouping is recorded in CHANGELOG.md’s own preamble, and xtask docs-check’s check_changelog_kinds fails if one returns, so the call does not need making twice.


Shipped since validation

Kept here so the list reads as live rather than stale, and so a finding is not re-reported after it is fixed.

IDFindingShipped in
P10Abandoned-scenario events appended after RunFinished — worse than reported: past the run’s terminal event, not just the scenario’s#15
B11${fake:*} collided across a scenario’s steps#15
P9.map.json gained phantom capture rows (fence-unaware scan, unrecognised custom methods)#15
B1Whitespace-only expect: produced an inverted sidecar span [9,8]#15
P2Non-UTF-8 PROEF_KEY/PROEF_ENV/PROEF_SECRET_<NAME> read as absent#18
P6Full disk: --output json exited 0 with truncated JSON#18
P7No stdout-side pipe-close test (both existing ones closed stderr)#18
P1Tee re-wrote the full slice on every write_all retry, duplicating run.log tail bytes#18
P8proef fmt rewrote CRLF → LF wholesale#18
Q3report -o outside the run dir shipped dead relative artifact hrefs#18
B8diff flagged a brand-new retried step as flaky#18
A3CLAUDE.md status stopped at post-M5#21
N1First run reported system error with no explanation (NON-TECHNICAL)#24
N2proef macros printed identifiers, never the match: sentence#24
N3macros refused to list when any step failed to bind#24
N4unbound_step’s help led with the pack maintainer’s action#24
N5No document described the scenario author’s workflow#24
§8init announced four files and reported five#24
Q5Ctrl-C skipped teardown silently — cleanup never ran, nothing said so#26
Q4--dry-run validated neither [run] setup nor [run] teardown#26
R2doctor did not report a missing pack schema (FIRST-RUN F4b’s second half)#29
B7secret set --value put the secret in argv, and the error text steered to it#29
§2.2init destroyed an authored proef-pack.schema.json (round 7)#30
§2.3a mixed suite+phase failure lost the phase label exactly when it disambiguated#31
§2.4--rerun after a phase-only failure blamed filters never passed#31
§2.1pre-0.6.0 records reported the wrong verdict with confidence#31
—diff counted a failing teardown as a test regression#31
P4fmt homogenized mixed-endings files beyond its hurl-blocks-only promise#33
P4proef --help described macros with pre-prose wording#33
P4WRITING-SCENARIOS’ two sample outputs drifted from the binary#33
—init destroyed an authored proef-pack.schema.json (round-7 §2.2)#30
Q7fuzz_tag_expr compiled but was in neither fuzz loop#30
B3windows.yml built and tested without --locked#30
B13justfile gate list omitted public-api (and the fuzz gate)#30
B5explain/diff/report each inlined ProjectConfig::load()#30
A6TROUBLESHOOTING’s exit table omitted 130#30
A4README’s ADR range and flag rows, and TECH-SPEC §10’s command surface, were stale#30, #34
A5TECH-SPEC’s publish claim and its run-dir inventory were stale#30
A1EDITORS.md claimed go-to-definition cannot land on a match: line#34
A7GETTING-STARTED’s copy of the scaffold comment had a word the scaffold does not#34
P11ADR-0015 described a worker on ScenarioFinished that is always None#34
B2a templated retry:/delay: under-counted the batch budget, abandoning healthy scenarios#35
B4--output json’s exit_code disagreed with the real exit after a JUnit failure#35
B6LSP completion snippets did not escape $/}/\#36
B9GitHub annotation file= and job-summary table cells were unescaped#36
—the --dry-run nudge echoed a command that was not the run validated (round-7)#37
P3--sarif emitted no startLine, so it annotated nothing#37
P5--watch did not retrigger on proef.toml#37
—a run against untouched scaffold routes got no coaching (round-8 §5)#38
—truncated-record fallback totals dropped Warned scenarios (round-7)#39
—fmt rewrote any file handed to it, not just a pack#40
—fmt trimmed the YAML skeleton, turning --check red outside its scope#40
C1negative-case authoring had no signposted catalogue form#43
C3expect: composition documented as a mechanism, never shown as the pattern#43
R9-1no proef fragments listing — neither way a fragment dies had a denominator0.11.0
§2.1a bind: key nothing reads passed silently — the one authoring mistake with no signal0.11.0
§2.2duplicate_fragment said “in both x and x” and offered a remedy that cannot work0.11.0
§2.3unbound_placeholder named two of ADR-0018’s three supply routes0.11.0
§3.1doctor did not know fragments exist — a path error surfaced as a name error0.11.0
§3.2config discovery searches only up, undocumented; no way to name the file0.11.0
§3.3init scaffolded only `hurl:, so ref:` was invisible to the persona built for it
—ADR-0007 value caps never crossed to fragments: retry: -1 validated clean0.11.0

Q7 is now closed (#30): fuzz_tag_expr is in both fuzz loops as well as the compile gate.


Ingested — round 19 (2026-08-31), validated claim-by-claim

Five confirmed defects, all shipped; five checks that cleared; four external triggers re-tested. The round’s shape: the heavily-audited paths (scheduler, record gate, outline expansion, the shard×shuffle×rerun composition) were probed and found correctly defended, so the yield came from what the output surfaces contain rather than from what the core computes.

R19-1 — a step’s name: reached the artifact and nothing else (shipped)

A macro with more than one step turns one feature sentence into several engine steps sharing a StepRef exactly. The emitter always wrote the authored name: into the artifact’s entry comment; StepRef never carried it, so the console, HTML report, JUnit, TAP, the job summary and explain printed the same sentence once per step with only the status glyph between a warning and the failure beside it. In a fresh reference run, 17 of 44 step identities were duplicates, and the pinned event snapshot was encoding the defect — three byte-identical step_finished for the cookie session is exercised.

Fixed by mirroring fragment (StepOutcome + step_finished), not by extending StepRef: several engine steps share one StepRef, so the label belongs to the engine step. Additive on the wire; schema stays 1. Retires two untrue claims — AUTHORING.md’s “they anchor artifacts, events, and failure output” and LoweredStep::label’s own “(events/console)”. Same class as reproduce_hint in the R18 wave: computed all along, printed all along, dropped by the record.

R19-2 — report -o wrote the machine into the shared file (shipped)

Absolute artifact hrefs, 12 per report, naming the author’s home directory — in the one output built to be uploaded. 0.13.0 scrubbed machine identity from the record (R12-1) and the record is clean; the HTML put it back. The absolute path was deliberate and pinned by a test, but it resolves only on the machine that produced it, which is exactly where -o output is not read. A relative href strictly dominates. Windows CI then caught a second half the local gate could not: the href was built with Path::display, and \ is not a separator in a URL, so a Windows-generated report’s links were dead either way — it is now built from components joined with /. The known macOS-only-gate hazard, paid again.

R19-4 — one palette token failed WCAG AA, and every dark pill did (shipped)

--skip was the single token the dark block does not redefine: a grey chosen against #0d1117 left carrying white text on white at 3.45:1. Writing the guard rather than the fix found the larger one — .pill painted color:#fff on status colours the dark palette tunes as text on a dark ground, so all four dark pills sat between 2.52:1 and 3.45:1. The pill foreground is a token now. Tests assert the ratio, not the hex, and that both palettes define the same token set (the absence that caused it).

R19-5 — the report had one heading and no outline (shipped, narrower than filed)

Filed as “no headings at all”; the timeline already had an <h2> — the first inventory ran against a record with no timing, so the timeline never rendered. Corrected before implementing: only the tag table and the scenario list lacked one. Both gained one, sharing the class the timeline already used.

R19-3 — the three post-run commands had no machine output (shipped)

explain, diff and doctor. A run directory carries no structured summary, so anything analysing a run it did not launch had to fold events.jsonl itself — the fold proef’s own two copies disagreed on three ways. Each object mirrors its prose field for field; doctor had to start collecting its checks before rendering them, so JSON is a second rendering rather than a second walk.

Cleared — checked, not defects (do not re-raise as omissions)

  • RecordGate’s (file, name) identity is safe: feature.rs’s dedup_names guarantees uniqueness feature-wide and its doc names this consumer.
  • Duplicate step rows are not a counting bug — the record keys steps by (text, occurrence ordinal). Only the surfaces were blind (R19-1).
  • Report keyboard focus is intact: no :focus rules, but no outline:none either, so native rings survive on button/anchor/summary.
  • The tag table is a real <table>; it renders no rows only when a run carries no tags.
  • crates.io showing no homepage for 0.14.0 is publish lag — the field landed after that release was cut (verified by ancestry), and appears on the next publish. Confirmed 2026-09-09: publishing 0.18.0 carried homepage = https://emrecdr.github.io/proef/ through, closing that half. documentation is still unset: for a binary crate that falls back to a docs.rs library page rather than the book, worth setting deliberately.

External triggers re-tested 2026-08-31 — three hold, one has since fired

  • OpenTelemetry export stays a non-goal. OTel graduated CNCF (2026-05), so the umbrella argument weakened, but the attributes that would carry a test run — test.case.name, test.case.result.status, test.suite.name, test.suite.run.status — are all still Development stability. PRD §3’s stated reason is current as written; only the re-check date moves.
  • CTRF was declined here and has since shipped. As re-tested on 2026-08-31 this read “still community-adoption phase; Microsoft’s test platform has a discussion issue, not an implementation. Trigger unfired” — accurate for its own date. The 2026-09-02 survey shipped it anyway as --ctrf (#160), rendered off the same fold as JUnit. Corrected 2026-09-10; the two sibling statements of the same deferral, in the RF audit below and at R3-5, were stale with it.
  • Both sacred pins are correct. hurl 8.0.1 is the latest release (2026-04-29) — no 8.1, no 9.0. Rust 1.97.1 is right under the written policy: stable is 1.98.0 (2026-08-18) and channel-rust-1.98.1.toml 404s, so the point release the policy waits for does not exist yet. R18-2’s refutation survives contact with the calendar.
  • An MCP server is declined, with a named trigger. The largest ecosystem shift since the last research round — Playwright, Cypress, BrowserStack, Maestro and ReportPortal all ship one, and Claude Code / Cursor / Windsurf consume them natively. They shipped MCP because their primary surface is a GUI or a cloud API and an agent had no other way in. proef is CLI-first with --format json and a pinned four-code exit contract: an agent already has a complete interface, and a second one is a second way to do one thing. The only real gap an agent hit was R19-3, now closed. Trigger: a concrete agent workflow that --format json plus exit codes cannot express. Recorded so proef lsp’s precedent is not read as an open door.

Noted, not filed

source_guards’ malformed-plural scan matches the literal (y) anywhere in a non-comment source line, so any code with a single-character y parameter false-positives (x.max(y) did). The guard’s intent is user-facing strings; scanning all code is broader than that. Left alone — it is working as a guard and tightening it to string literals is more risk than the trap is worth — but the next author to trip it should know why.

Ingested — the deep improvement report (2026-08-25), validated claim-by-claim

A twelve-stream self-audit plus competitive/ecosystem research (five code audits, three UX audits, four research streams; ~125 findings), every load-bearing claim re-verified against the tree before acceptance and three proved empirically (measured stack-overflow abort and exponential backtracking in the tag glob; observed nondeterministic gherkin error ordering). The full report is the session artifact “proef — deep improvement report”; this section records the verdicts and what remains open.

Wave 1 — shipped (#112–#116)

  • #112 — a comment on a section header no longer blinds any scan ([Options] # tuning + retry: -1 dry-ran clean — ADR-0007’s named hole; [Captures] # ids dropped sidecar rows; [Asserts] # note doubled a section); the delay cap learned hurl’s h unit (delay: 5h validated clean at 5× the cap); pack-scope bind: resolves arg-free instead of in whichever macro ran first; the tag glob is the two-pointer match (oracle-property-tested — the metachar branch previously had zero generated coverage); multiline_bind refuses \r/controls; the expect: merge shares the emitter’s hardened response-line check.
  • #113 — a Ctrl-C in --watch’s debounce window no longer launches one more full run; a delivered watcher error or rescan burst (queue overflow) retriggers instead of leaving the watch permanently deaf. Punctured and re-closed 0.12.0’s “staleness class closed for good” claim.
  • #114 — eleven silent-failure sites gained voices (store-poison save, suite walker, doctor/fmt over unreadable trees, non-UTF-8 env values, .map.json, LSP config, flaky degrade, docs-check vacuous pass, Sinks severity filter); fmt recognizes every literal-block spelling.
  • #115 — a travelling record can no longer lie (scenario_finished.file = "" key mismatch silently emptied every step map — flaky’s Latent verdict was unreachable and diff --fail-on-regression certified green), crash (256 MiB read ceiling; saturating sums; saturating Span::len), or steer (rerun_of/--run-id single-component validation; [tag-links] URL percent-encoding + http(s)-only in both sinks).
  • #116 — saveAs: global refuses a secret it can find (needle set, in core’s World, every engine covered) rather than one it can equal (engine-side, raw values only); the invariant is now property-tested as CLAUDE.md had claimed. The SLA gate applies the same @quarantine non-gating list as the exit code.

Waves 2–5 — shipped (#118–#128), audited 2026-08-31

This section said these were open for a week after they landed. The list exists so a finding is not re-reported once it is fixed, and it failed at exactly that: a re-read sent one round toward rebuilding wave 2, and repeated two of its claims to a reader as open work. Corrected by checking the tree for each item rather than trusting the entry.

WaveShipped inSpot-checked by
2 — CI-sink conformance#118failure_detail_reaches_attribute_and_text_node_alike, illegal_bytes_and_ansi_never_reach_the_xml, an_oversized_summary_truncates_and_says_so, annotations_cap_at_ten_with_an_honest_notice, composed_identities_form_a_set, times_are_three_decimal_seconds
3 — UX#119–#122console is_terminal colour, clap_complete/clap_mangen in the archives, doctor’s project block, the report’s jump nav + data-f filter, the --format / -o split
4 — diagnostics#123–#125suggest_or_enumerate, code_description, proef::config::* codes, match_span in use
5 — docs & distribution#126–#128docs/INSTALL.md, .sha256 sidecars, the README comparison

Two wave items did not ship, and one of them should not:

  • [[ATTACHMENT|path]] in a testcase’s system-out — declined, with a trigger. It is a Jenkins-plugin convention: GitLab and GitHub ignore it, so it buys a link for one vendor’s users who also installed the JUnit Attachments plugin. It would put a filesystem path inside an artifact built to travel — the class of defect R19-2 had just finished removing from the HTML report — and the reader’s need is already met twice over, by the reproduce-hint curl in the failure content and by the HTML report’s own artifact deep-links. Trigger: a user on Jenkins reporting that neither reaches the artifact for them.
  • llms-full.txt — still unshipped, and the entry that proposed it already records that the SEO case for it is empirically dead. Left as-is.

Corrections to this list’s own claims (all four were stale)

  • “the exclusive-tags scheduler and RecordGate have no direct tests” — false. proef-core/tests/runner.rs carries an_exclusive_scenario_never_shares_the_pool, back_to_back_exclusive_scenarios_each_get_the_pool_alone, cancelling_during_an_exclusive_drain_still_completes_the_run, and — for the gate — abandoned_scenario_emits_nothing_after_run_finished.
  • “bake_entry_options deserves a proptest” — it has one, in lower.rs.
  • “the --format/-o split is open” — shipped in #122.
  • “match_span is computed and unused” — it is used; the diagnostics wave wired it.

Verified against the tree (each entry says whether it is open or closed)

  • --shard balances by hash while the timing data to balance by duration is already retained. Closed (2026-09-02) by --shard-weights, after validation found the obvious design silently wrong.

    shard_bucket(file, name, count) took identity only, so a 4-way split was balanced by count and not by time — and a CI matrix finishes when its slowest shard finishes. The weight now shipped is the sum of a scenario’s step durations (record::StepRun::duration_ms), which measures work rather than queue wait; the wall-clock span would have been the wrong number and the record reader does not retain it anyway.

    The hazard this entry existed to record. The natural implementation — “weight by the latest record in runs-dir” — is silently incorrect for the only case sharding exists to serve. Each shard of a CI matrix runs on a separate machine with its own (usually empty) runs-dir, so every job would compute a different weight table and therefore a different assignment. Scenarios would run twice or not at all, and the suite would still report green. Nothing about that failure announces itself.

    Shipped shape: every run that reaches its suite writes a small timings.json into its run directory (an aborted setup has no suite to weigh, and writing its own scenarios would skew the next split with identities that never run), CI archives that one file, and each matrix job points --shard-weights at the same copy — so the split is a pure function of (selected scenarios, that file). Weighted scenarios are placed longest-first; unweighted ones fall back to the frozen hash, and the two rules partition rather than compete, so a test added after the timings were captured still runs exactly once. A three-way matrix test asserts set equality both ways; mutating placement by one bucket drops two scenarios and the test names them.

    What it gives up is what hash mode was chosen for — a balanced split is not stable under insertion — which is why the flag is opt-in. See CONFIG.md.

  • lower.rs threads the same mutable trio through twelve functions. Closed (2026-09-02), and the premise it was filed under was wrong.

    The original filing said “mechanical, no behaviour change: introduce a context struct and make them methods”. That would have broken the code. The closures (resolve_in, resolve_pack_scope) take refs and sinks as explicit parameters rather than capturing them, precisely so they remain callable while other state is mutably borrowed — and a method on &mut self cannot be called while self is borrowed elsewhere. Threading was not an oversight; it was load-bearing, and validating that is what turned a rename into a design.

    Shipped: three bundles, each a type the code already implied — Emit { out, refs, sinks } (the mutable outputs, always passed together), StepScope { step_ref, ctx, at } (what stays fixed for one authored step however deep expansion recurses), and Finished for the four values describing a completed step. The threading discipline is unchanged; only the arity is. Arity suppressions workspace-wide: 13 → 6, lower.rs at zero.

  • Hurl grammar in proef-core vs ADR-0002’s diff-empty claim. Closed (2026-09-01) by an ADR-0002 amendment plus a guard — and this entry was wrong three times over. “~290 lines” counted #[cfg(test)] fixtures; the correction to “19 lines, all in lower.rs, four concerns” fixed the count and kept two errors. It is 19 lines across three files — lower.rs, emit.rs and pack/validate.rs — and the four “concerns” mostly are not concerns: is_method_line, is_section_header, is_response_line and is_header_line are already one canonical pub(crate) set that all three files share.

    What the entry missed entirely is the finding: the seam already solved this once, on the reading side. StepKindSpec::options exists, in its own words, as “the seam that keeps option spellings out of proef-core” — added because matching "retry-interval:" as a literal meant “one rule lived at two altitudes.” It covers recognising options. The core still writes retry:, retry-interval:, delay: and variable: as literals, so the same rule still lives at two altitudes, in the other direction. That, not the line count, is the actual asymmetry.

    Resolved as: the thirteen-token vocabulary is sanctioned and closed on the ADR record, pinned by source_guards::hurl_grammar_in_core_is_the_closed_set_the_adr_names (which fails on growth and on an existing token spreading to another core module, and names both remedies in the failure). Moving the written half behind the seam is deferred with a named trigger — a second engine being scheduled — because until then it relocates seven literals that exactly one implementation will ever supply, at the cost of a public-API break.

    The meta-lesson, and the fifth instance of it this programme: a claim that lives only in prose decays, and decays in whichever direction makes the writer’s point. Every wrong version of this entry overstated the problem.

  • Reading events.jsonl is spread across seven files. Closed (#150), and it was not the duplication it was filed as. Chasing it found that the 256 MiB record ceiling reached two of its four readers: explain and report each opened the file with a bare read_to_string, so neither had it — report even used the guarded reader for the base record two dozen lines below the raw read of the primary one. Both go through record::read_events now, and a source scan in source_guards makes the next reader use the same door. The folds that could disagree were already unified (record::parse_record, report::suite_totals), and explain --format json hands consumers the canonical answer rather than inviting an eighth reader.

  • captures_before is O(steps²) — deliberately left. It runs only when a ref:/bind: consumes it and is bounded by scenario size, so threading a running set through lowering is churn against a bound that is not tight.

  • Redaction runs inside the reporter mutex. Fixed (#149). The allocation cost had already been addressed (the miss path no longer allocates, and a clean field keeps its Arc); the structural point stood until the masking simply moved above the lock(). It reads the event and the needle set and writes neither, so it never needed the lock at all.

Analysed 2026-09-01 — none of the three was an ADR question

This section carried three items as charter questions needing a new or amended ADR. Checked against the tree, none of them needs one, and two had the wrong governing principle attached. Each entry below states the verdict and the evidence; the decision to act is a one-word answer, not a design exercise.

  • GitHub-summary permalinks to failing lines — build it; ADR-0020 is untouched. The entry assumed the commit must come from GITHUB_SHA, which §1 forbids by name. It need not, on two counts. First, a link to the failing line already ships: github_annotations (ci_reports.rs:363, live at exec.rs:1444) emits ::error file=…,line=…:: per failing scenario, and --sarif (sarif.rs) carries startLine. GitHub resolves both against the commit the job checked out — proef never reads a SHA to make that work. The residual gap is narrow: the job-summary tables are inert text where the annotations are linked. Second, closing that gap needs only two mechanisms that are already accepted and already shipped — [tag-links] (config.rs, glob → URL template with {tag} substituted, applied by the HTML tag table and the GitHub summary, documented as “base config only — a link is a project fact, not an environment one”), and ADR-0020 §1’s own worked example, --meta commit=$(git rev-parse HEAD). A [source-links] table over {file}/{line}/{commit}, with the commit handed over as metadata, sits inside both rules unchanged. §1 forbids proef harvesting the variable; it does not forbid a user handing it over — that is precisely the distinction the ADR was written to draw. It also serves GitLab, Bitbucket and self-hosted forges, which a GITHUB_SHA read never would.

  • Chrome-trace export of the scheduling timeline — decline; the filed reason is false and the conclusion survives on a different one. “A second rendering of what the HTML timeline already shows” is wrong: render_timeline (html.rs) draws one bar per scenario per worker lane, while a trace’s whole value is step-level nesting and zoom, which the report does not have. (Cited without a line number on purpose — the first version of this entry named one, and adding the neighbouring render_slowest moved it.) The adjacent gap that entry implied — that the page could not say which scenarios cost the most — is closed separately by the report’s ranked Slowest section; the trace question is unaffected, because that section ranks scenarios and a trace nests steps. The correct reason to decline is that the JSONL record already carries every step’s start and end (ADR-0015 injected timestamps), so a trace is a short transform of data proef publishes in full — and a second export format for already-published data is what one canonical mechanism forbids. No consumer has asked, which is the same CTRF/TAP-14 discipline applied above. Action: document the conversion recipe instead of building an exporter. Trigger: someone who has run the transform and hit something it cannot express.

  • Templated report output paths — decline; the one-path rule was the wrong lens. That rule governs the resolution base (a path in proef.toml resolves against the config’s directory, a flag against the cwd); templating touches neither, so the two never conflicted. The governing principle is ADR-0020 §1’s axis again: -o "reports/$RUN_ID/index.html" is shell interpolation the caller already controls, and asking proef to interpolate it is asking proef to own a value that is already handed over. Per-run records also already have their mechanism — [run] runs-dir plus keep-runs rotation (config.rs:101-111), deliberately project-wide so there is one record store and one policy — so a second per-run path scheme would be the duplicate, not the gap.

The pattern is now consistent enough to be worth stating: an item filed as a charter question is usually an item whose governing principle was guessed. The only-failed console sat here until #147 shipped it as --console failed, a fourth mode on the existing flag — exactly what the entry asked for and no ADR at all. Two more had already been decided. Check the tree and name the actual principle before filing the next one.

Declined — do not re-raise

  • In-run scenario @retry (cucumber-rs’s headline feature, ranked first by one research stream): the CI-standards stream independently established retry-until-green as the anti-pattern (a 25%-failure bug passes 99.6% of the time under three retries) and proef’s detect-then-quarantine shape as the consensus architecture; per-step retry: already covers polling. Document the stance in TESTING-STRATEGY instead — it reads as a gap until stated. Done 2026-09-10: stated in TESTING-STRATEGY §5, with the arithmetic, beside the determinism rules it belongs with.
  • CTRF — shipped 2026-09-02 as --ctrf (#160); the “deferred trigger was checked and has not fired” verdict recorded here is superseded. It was right about the ecosystem — no CI platform ingests CTRF natively (GitLab/CircleCI are JUnit-only, GitHub has no format at all, Buildkite has its own JSON) — and that is why the entry is kept rather than deleted: the trigger genuinely never fired, and the format shipped for a different reason, that rendering it off the existing JUnit fold cost a renderer rather than a mechanism. Buildkite JSON remains the higher-yield target if a consumer ever does materialize (its span model maps 1:1 onto step outcomes). Corrected 2026-09-10.
  • TAP 14 (unratified branch, zero declared consumers) · Bruno-style granular exit codes (ADR-0009 is a contract) · an OS-keychain secret backend (second storage mechanism) · a user-level personal config file (proef.toml is the one channel; NO_COLOR covers terminal taste) · Karate match within sugar (hurl predicates cover it).

Environment note (machine-side, not repo-side)

Homebrew’s Rust (1.98.0) shadows rustup on this machine’s PATH (/opt/homebrew/bin/cargo first), which breaks cargo +nightly and the public-api gate and silently un-pins builds; Wave 1 gates were re-run under the pinned 1.97.1 explicitly. Owner action: brew uninstall rust or reorder PATH. Resolved 2026-09-10: the PATH is reordered — ~/.cargo/bin precedes /opt/homebrew/bin, and a login shell now resolves both cargo and rustc to the pinned 1.97.1. The Homebrew formula is still installed and harmless where it now sits; nothing needs uninstalling.


Ingested — Robot Framework capability audit (2026-08-24)

A deliberate mining of Robot Framework 7.x for transferable ideas, run as five extended-context investigations (one per adoption candidate, plus a counter-audit attacking the first-pass verdicts), every load-bearing claim reproduced against the tree before anything shipped.

RF wave 1 (shipped, #96–#101)

Detail cap at the engine boundary (RF’s 40-line rule) · tag-atom globs (*/?, anchored; the silent-no-match became the intended selection) · flows feature descriptions (parsed, was dropped) · --shuffle seeded by the run id (R3-9, one determinism knob) · reproduce_hint into the record (the console knew more than explain did). The shard parity fix (#96) was round 18’s, not RF’s, but shipped in the same wave.

RF wave 2 — the schema wave (all three shipped, #103–#105)

  • RF-W2-skip (shipped — ADR-0019, #103) — @skip / @skip:reason-token (prefix verified to parse as one tag); reason on ScenarioFinished+ScenarioOutcome; one reserved-tag module (quarantine moves in); the two mapped collisions are the point of the work: --rerun re-queues Skipped-on-cancelled (authored reasons start with @, mechanical never do), and diff buckets Failed→Skipped as fixed — three-way bucketing required. All-skipped → exit 0 (ADR-0009 argument recorded); harness → libtest-mimic ignored.
  • RF-W2-tags (shipped — #104) — tags on ScenarioFinished only (the cancel-skip path emits no Started), exclusive on ScenarioStarted (closes R11-6), schema stays 1; HTML + GH-summary per-tag tables, suite-only per ADR-0014; the quarantine non_gating list re-derivation collapses into the new one owner; D1 becomes its predicted recipe. NOT building RF’s tagstat combine/link/doc knobs.
  • RF-W2-meta (shipped — ADR-0020, #105) — --meta k=v + [meta]/[env.<name>.meta] on the existing precedence chain; RunStarted.env auto-recorded (handed-over, not harvested — R12-1’s real axis); values through the one sink-boundary mask; JUnit <properties> deferred by the named-consumer method; never in artifacts; ADR codifying explicit-injection-only ships with it. A shuffled: bool marker rides the same RunStarted change (deferred out of #100 for one wire change instead of two).

Hazard both schema items must clear: stamp_scenario_timing and phase_sink (exec.rs) rebuild scenario events field-by-field — a new field compiles clean and is silently stripped from every stamped stream unless threaded there, with an integration test per field.

RF wave 3 (shipped, #106–#108)

  • Rerun merge-at-report (shipped — #106) — the record-composition half of E2: a rerun’s JUnit carries the base’s not-re-run suite scenarios as ordinary testcases, and report overlays the base for one whole-suite page. Composition over records — the record files themselves never merge (ADR-0008); totals and the exit code stay the rerun’s own (ADR-0014).
  • Console modes (shipped — #107, extended #147) — --console full|failed|dotted|quiet. The OSC-8 hyperlink half did not ship: printed paths stay plain until a terminal consumer asks (the same named-consumer method as JUnit <properties>).
  • Quarantine in JUnit (decision taken; shipped with ADR-0019, #103) — a quarantined test-failure maps to <skipped message="quarantined failure (non-gating): …">, so Jenkins, the dashboards and the exit code agree.

Deferred with named triggers

Report-size mechanism (first >1k-scenario record; failures-only render, one mechanism not three knobs) · --runemptysuite (first CI consumer; settle the early-error record first) · JUnit <properties> (a Jenkins-keepProperties user). The --tagstatlink analogue left this list: it shipped as [tag-links] (#108).

Rejected, and where the basis actually lives

--nostatusrc (ADR-0009 is a contract) · argfiles/ROBOT_OPTIONS (proef.toml is the one channel) · pre-run modifiers, custom parsers, listener API (PRD §3 non-goals + product identity) · GROUP · Set Test Message · --exitonerror (--max-fail is the one early stop) · robot:private (the macro listing exists to show the vocabulary). Two stances the counter-audit showed are held but unwritten — control flow lives in packs (when:/optional:/retry:), prose stays declarative; and proef has no runtime extension surface, the record is the observation API — both belong in AUTHORING or a short ADR when wave 2’s ADR is written anyway.

Counter-audit corrections (for the record)

“--rerun ahead of RF” was half wrong — ahead on selection (cancelled-tail union), missing the merge half entirely; see E2. “HTML report ahead of log.html” — ahead on visualization, was behind on forensics (the reproduce_hint gap, now closed by #101; request/response excerpts remain a deliberate non-goal until asked). “@quarantine ≈ --skiponfailure” holds only for exit-code CI (see wave 3). RF 7.4 added a Secret type — proef’s redaction invariant predates it; banked as an ahead.

Ingested — round 18 (2026-08-24), validated claim-by-claim

R18-1 — the shard hash collapse, round two (confirmed — shipped)

Round 18 tested this registry’s R17-2.1 refutation instead of restating the round-17 claim, and won the half that matters. The mechanism is arithmetic, not statistics: FNV-1a’s multiplier is odd, so the accumulator’s low bit is exactly the XOR-parity of the input bytes’ low bits; a scenario named after its feature file — the commonest Gherkin convention — duplicates content across the (file, name) identity, whose parity contributions cancel, leaving a corpus-constant bit: N=2 → [20,0], odd buckets empty at N=4 (reproduced against the real shard_bucket, then pinned red in the balance test before the fix). The R17 balance test could not see it by construction — all three corpora held the file constant, the one condition under which raw FNV behaves. Shipped: Murmur3 fmix64 finalizer on shard_bucket (Breaking: every matrix re-deals), the mirrored corpus in natural_corpora_spread_across_shards, and bounds recalibrated to what a well-mixed hash yields (no empty shard at any N; the 3× skew bound at N=2 only — a fair deal of 20 over 4 buckets legitimately produces [2,7,5,6]). The reviewer’s own concession stands for the record: fmix64 is mildly worse on constant-file corpora ([7,13] vs [10,10]), which is randomness, not structure — no empties.

R18-2 — Rust pin “four days overdue” (refuted — and the policy is now written)

The pin follows the practiced policy — adopt a new stable at its x.y.1 point release, ~3–4 weeks after x.y.0 (1.98.1 expected mid-September) — but the reviewer read CLAUDE.md’s “always latest stable Rust”, which said otherwise. An unwritten policy that contradicts the written one is a docs defect on our side: the policy now lives in CLAUDE.md and RELEASING.md, and the pin bump lands on 1.98.1, as it always would have.

R18 closures

Eleven round-17 closures re-verified by the reviewer against cc75129 with original repros; nothing reopened. The review singles out the machine-body funnel and the flags-direction docs gate as the durable forms of their fixes.

Ingested — round 17 (2026-08-23), validated claim-by-claim

Two P1s filed; one confirmed both ways it can be read, one refuted by measurement. Every confirmed item reproduced against b5b320a before any fix.

R17-2.1 — --shard hash collapse at power-of-two counts (refuted for constant-file corpora; corrected by R18-1)

The filed claim: FNV-1a’s unmixed low bits collapse the distribution at N=2/4/8 (“100% of scenarios land in one shard”), fix with an fmix64 finalizer. Measured with a model calibrated against the frozen-literal test (exact match on every pinned value), the claim inverts. Natural corpus shapes — numbered scenarios, prose names, outline #N instances, camelCase, verb templates, multi-file — are near-uniform under the current fnv % count: [10,10], [11,9], [4,5,6,5] at their widths. The proposed fmix64 is worse on the same corpora ([7,13] where FNV gives [10,10], empty shards at N=8 that FNV does not produce): FNV’s parity-structured low bit behaves like round-robin on templated names, which real suites are full of. Collapse requires a degenerate corpus — every name an even-length run of one character — which no suite exhibits. The filed measurement tables do not reproduce from the calibrated function. What survives: no test asserted balance — shipped as a distribution test over natural name shapes, so a future hash change that does skew fails loudly.

Round-18 correction: the refutation above held only where its evidence did — every corpus it measured kept the file path constant. Round 18 showed the varying-file half was real (see R18-1): the low bit of raw FNV is byte parity, and a scenario named after its feature file cancels to a corpus-constant parity — [20,0] at N=2. “Degenerate corpus only” was this registry’s error, not the reviewer’s.

R17-2.2 — bind: validation refused input the engine accepts (shipped)

Confirmed, both halves, plus a third the round missed:

  • {{newUuid}} in a bind value was refused as an unbound variable; it is a hurl function (ExprKind::Function), and stock hurl 8.0.1 runs the equivalent line (reproduced both directions). “What does this text read” is now the engine’s answer — FragmentSupport::template_reads, the same AST walk the fragment scanner uses — so the tree holds one answer, not two disagreeing ones.
  • A sibling literal bind sorting before the bound key is a real supplier (injected lines are written and evaluated in name order) and is now accepted; a later-sorting sibling stays refused, with the ordering named in the help.
  • The round’s fix list missed the ordering half: injection landed at the head of an author [Options] section, so the fragment-supplies-it route the check accepts was assigned too late to be read at run time. Injection now lands at the section’s end; pinned by a_fragments_own_variable_evaluates_first.

R17-2.3 / 2.4 / 2.5 — machine output and phase reporting (shipped)

An empty shard wrote prose (plus a stray-space run) where a --output json/TAP body belongs; a setup abort wrote JUnit but zero machine-stdout bytes; a failed teardown reached no report at all. Shipped as one mechanism each way: emit_machine_body is called by every terminating path (pool, empty shard, both setup aborts) with ADR-0014 suite-only totals and the path’s own exit code — the note moved to stderr under machine output — and a failed teardown’s outcomes ride into write_junit as their own suite (#78’s rule made symmetric; a green phase stays out). Deliberate scope as recorded then: the GitHub summary keeps pool-only totals. The second audit pass showed the code does not hold to it — a setup abort passes the setup summary as the primary, so setup failures render in the GitHub summary while teardown failures do not. Queued: unify all three CI sinks on “a phase appears when it fails” at the write_ci_reports boundary, with totals staying suite-only everywhere (ADR-0014).

R17-2.6 — batch (shipped)

README omitted --shard/--max-fail (and the #73 gate was blind to the flags direction) — closed with a reverse-flags gate whose measured burden was exactly three flags; explain’s truncated-record fallback now filters through is_suite() (the fourth consumer #72’s helper was built for); the canary refuses a backport older than the pin by semver ordering, not equality; identical warnings collapse to one with a repeat count (every class, at the front-end aggregation — bind_shadows_capture was the motivating fifty-warning wall); quick-xml rides at quick-junit 0.7’s in-tree copy again, one generation in the lock. P4s (all shipped in #87): the pages workflow comment now states that upstream/’s .patch files are served; docs/runbooks/ entered both living_docs scanners; outline identity’s positional #N is documented in AUTHORING with the column-placeholder remedy; the #79 comment stopped claiming file:line survives in the failure detail.

Standards note

Rust 1.98.0 released 2026-08-20 (verified against the channel manifest). The round calls the pin overdue; house policy waits 3–4 weeks after x.y.0 and targets x.y.1 — the window opens ~2026-09-10.

Open — round-9 residue (ingested 2026-08-12)

The review’s P1/P2 and three P3s shipped in #48 and #50. What follows is what was verified and deliberately not built, so none of it depends on remembering.

R9-1 — proef fragments has no listing command (shipped)

flows lists scenarios and macros lists the vocabulary; nothing lists the corpus. There is no way to ask which fragments exist, which are referenced, or which .hurl entries carry no annotation — and an unannotated entry is dropped at scan time by design, so the tool structurally cannot report what it never built.

Raised by a consumer migration whose coverage gate (“every @proef name is referenced, every entry is annotated”) had to become a script that repo owns. Not built for 0.10.0 on purpose: new public surface, and the migration was unblocked by correcting its own gate instead.

Shipped. A second migration report (ADOPTION-REQUEST.md, 97 entries) supplied the field evidence this entry was waiting for and ranked it first of seven. proef fragments now names both death modes apart, lists unannotated entries by line, and gates CI with --check; --require-annotated is opt-in because an unannotated entry is inert by design (ADR-0018), so “not done yet” is a porting team’s reading of that signal and not every adopter’s.

R9-2 — fuzz coverage does not reach the fragment surfaces (shipped)

fuzz_pack_load runs with an empty corpus, so ref:/bind: clash logic never executes under fuzzing; the annotation scanner’s entry-boundary arithmetic — proef’s own code, not hurl’s — and bake_entry_options’ textual injection are unfuzzed entirely. Split the fuzz input into pack and corpus halves, and consider a fuzz_fragment_scan target (nightly, accepting the native-libs cost).

Shipped, and the prescription was half wrong — measurably. Splitting the input into pack and corpus halves was tried first and did not work: a byte-oriented target never resolved a single ref: in 1.45 million runs, because reaching the rules means discovering valid YAML and a matching corpus name simultaneously. Verified by probe (panic on a resolving ref:, run the fuzzer, see whether it fires) rather than assumed from coverage numbers — which is the same mistake this finding is about, one level up.

What shipped instead is fuzz_fragment_binding, structure-aware: it builds a well-formed pack and corpus from the input and spends the budget on the name space, so every run reaches the rules. The probe fires in seconds. fuzz_pack_load stays byte-oriented and unchanged — parser totality is a real job and the split would only have diluted it.

The fuzz_fragment_scan half was declined for a concrete reason, not on cost alone: cargo dependencies are package-level, so adding proef-engine-hurl to the fuzz crate compiles hurl for all five targets and drags native libraries into a job that has none. Hurl’s scanner is instead property-tested in proef-engine-hurl, where those libraries already are — pinning that every reported line lies inside the file, that entries are accounted for exactly once in order, and that no fragment’s text runs into the entry after it. The last assertion was added after mutation testing: the first draft passed with the boundary deliberately broken.

Still open from this entry: bake_entry_options’ textual injection is unfuzzed. It is lower-time, not load-time, so it sits behind lowering rather than pack::load and needs its own target.

R9-3 — no resource bounds on the corpus read (shipped)

No per-file or file-count cap: a multi-GB .hurl is read whole on every command that loads packs. Pairs with the read-resilience work in #48, which made the read survivable but not bounded.

Shipped, and worse than filed by one word: not “a multi-GB file” — a 279 MB file cost 601 MB of resident memory on proef flows, a command that never looks at a fragment, over a file carrying no # @proef annotation at all. The doubling is read_to_string into a String and then Arc::from(&str), which copies.

Bounded now at 8 MiB per file and 64 MiB per corpus, measured from the directory entry so an oversized file is never allocated (601 MB → 15 MB on the same input). Reported through the per-file diagnostic channel unreadable_file already established — skipped, never fatal — and applied in proef lsp too, where the corpus is held between requests rather than for the length of one command. The laziness promise is intact: a corpus nothing ref:s still reports nothing and exits 0, pinned by a test.

The Arc<str> copy itself was left alone. Removing it means changing PackSource’s type across every reader, which is a wider change than a bound and buys a constant factor on an input that is now capped anyway.

R9-4 — a bind that shadows a capture is silent (shipped)

hurl’s variable: assigns into one shared set, so a pack- or macro-scope bind: re-assigning a name an earlier entry captured overrides it for every later entry, with no diagnostic. A warning shaped like option_declared_twice fits — the difference is that this one is only decidable where the capture set is known, at lower time.

Shipped as proef::lower::bind_shadows_capture, a warning per the verdict above — a fixed value over a live session is sometimes deliberate. Only a literal bind warns: a secret bind skips the [Options] path entirely, so the earlier capture’s assignment stands and there is nothing to warn about (pinned by a unit test). En route it was validated that unread_bind_key already narrows the surface to binds a fragment in scope reads — the live gap was exactly the capture-shadow shape.

R9-5 — {{x}} inside a bind value is unvalidated at lower time (shipped)

It fails at run time instead of at --dry-run: loud, but late, and the late half is what --dry-run exists to prevent.

Shipped as the same proef::lower::unbound_placeholder the fragment check uses — one code for one defect class — naming both the placeholder and the bind key, anchored on the feature step (pack-line anchoring from lower time is R1’s recorded deferral). The accepted suppliers, each pinned: an earlier step’s capture, the fragment’s own [Options] variable: (authored lines precede the injected ones), and a secret in scope — a run-time {{secret}} reference never puts the value in an artifact, unlike the ${secret:…} splice that secret_in_composite_bind refuses.

R9-6 — provenance is cwd-dependent (shipped)

Run from a subdirectory and step_finished.fragment, explain’s via, JUnit and the diagnostics carry an absolute machine path; the record-portability claim holds only from the project root. Relativize against the config root rather than cwd — the same boundary [run] fragments already resolves against.

Shipped as part of R12-1, which found the same defect reaching further than this entry describes — the safe case it names, running from the project root, had stopped being safe. The prescription here was the right one and is what landed: one anchor, the config directory, for every input kind.

R9-7 — smaller edges, verified and recorded

Artifacts written inside a fragments root poison the corpus with proef’s own output (loud, but the remedies misdirect — skip files carrying the artifact header, or document it); a step-scope bind: key the fragment never reads is silently baked as a run-level variable: and can shadow a later capture, and an unused ${secret:} bind silently widens the required-secret set (warnable at step scope, where it is decidable); a # @proef annotation placed mid-entry is silently ignored and the resulting unknown_ref does not hint at misplacement; proef macros prints a corpus error twice on the degraded path; same-file duplicate annotations read as “declared in both f.hurl and f.hurl”.

The double print is broader than filed (verified 2026-08-14 while adding the corpus bound, which inherits it). It is not specific to macros: proef fragments does it too, and to any corpus diagnostic — unreadable_fragment_file and the new oversized_fragment_file alike. The mechanism is that commands::fragments renders corpus.diagnostics() itself and then loads the suite, whose failure path renders the same diagnostics again. Both land on stderr, so the count line reads 1 error(s) under two rendered copies. Left here rather than folded into the bound: it is a rendering decision about which of the two sites owns corpus diagnostics, not a property of any one diagnostic.

Open — round-10 residue (ingested 2026-08-12)

Found by a cleanup review over the fragments branch, after its own gates were green. All three are consequences of what that branch added; none is a defect in what shipped before it. Recorded rather than fixed in place because each is a behaviour change, and the branch was already carrying two correctness fixes.

R10-1 — --config is honoured by the runner and ignored by the editor (shipped)

--config <path> bypasses the upward search so a proef.toml beside the suite becomes usable. proef lsp never sees it (it re-discovers via ProjectConfig::load_from), and --watch watches the config found by its own fresh upward search, not the one the run was given.

So in exactly the layout the flag exists for, proef test --config … runs green while the editor gets no [run] fragments and reports every ref: as unknown — diagnostics disagreeing with the runner, which is the drift that makes an editor untrustworthy.

Shipped. ProjectConfig now keeps the file it was read from and derives root from it, rather than storing the directory and leaving every consumer that needed the file to search again. --watch watches the config the run resolved through; proef lsp takes the flag and lets it outrank even the client-announced workspace root, since a named file is not a guess to be improved on. The free config::config_path() — the fresh upward search both bugs went through — is gone, which is what stops the class recurring. proef lsp still starts when a named config is missing (an editor offering less beats one that will not boot), where the runner exits 2; the asymmetry is deliberate and documented.

R10-2 — proef fragments judges reachability over a smaller universe than the runner (shipped)

[run] setup / [run] teardown are not loaded, so a fragment used only by a phase feature counts as never run and fails --check — a false CI failure in the workflow --check was asked for, unless the phase feature happens to sit inside the suite directory. exec::execute already threads one corpus through both phase validations and both phase runs; the listing needs the same universe.

R10-3 — three predicates answer “is this a fragment file?”, and they disagree (shipped)

front::fragment_extensions (exact match, and its doc claims to be “the one place that answers this”), pack::scan_fragments (exact), and the LSP’s own is_fragment (case-insensitive). api.HURL therefore invalidates the editor’s corpus but is never scanned by core or discovered by the CLI.

The shared home is proef_core::engine, beside StepKindSpec — it is pure logic over the registry, so it is sans-IO-legal, and proef-lsp cannot reach proef-cli’s copy. Worth pairing with the deeper question the LSP predicate raises: membership in discover_fragments() is the real test, and an extension match also claims emitted artifacts that happen to end in .hurl.

R11-1 — proef.toml resolved its paths against two different roots (shipped)

[run] fragments resolved against the config file’s directory and suite, setup, teardown and runs-dir resolved against the working directory, so the same relative spelling meant two directories depending on which key it sat under. .proef-state.json and .proef-secrets.json were cwd-anchored too and appeared in no inventory, making two shells in one project two Worlds and two secret stores. One rule now: written paths resolve against the config, typed paths against the working directory.

R11-2 — --watch retriggered on a config it then ignored (shipped)

The loop watched proef.toml and reran on an edit while the rerun used the startup snapshot, so changing [url] base produced a rerun that called the old host. Fixed by re-reading per rerun — and by moving the startup config out of scope, which makes the stale value unreachable from the rerun closure and the invariant a compile error rather than a habit. Which directories are watched is still fixed at startup, so [run] fragments and [run] suite need a restart to be watched. runs-dir was in that list until R11-8 showed it did not belong there: it is not a watched root but an excluded one, and freezing it was the bug rather than the limitation.

R11-3 — --config was honoured, swallowed, or ignored depending on the command (shipped)

doctor printed the error for a missing named file and then reported on defaults, exit 0; fmt, init, schema and secret accepted a nonexistent path silently. Three documents called the flag global to every subcommand. A named-but-missing file is exit 2 everywhere now; doctor stays lenient about discovery, which is a different claim.

R11-4 / R11-5 — [run] exclusive-tags did not validate itself (shipped)

--dry-run never parsed the expression, and a well-formed expression matching nothing was silent — both defeat the reason the setting is a config expression rather than a reserved tag name.

R11-6 — exclusivity is invisible in the run record (shipped — RF wave 2)

Event::ScenarioStarted carries no field saying a scenario ran exclusively, so a post-mortem cannot tell a deliberate drain from a stall: the record shows parallelism dropping to one and nothing explaining why. An additive field is permitted by ADR-0008, and the reporters would need to decide whether to surface it. Filed rather than built — it is a design question about what the record should say, not a defect, and the run behaves correctly either way.

Closed (2026-08-24): scenario_started carries additive exclusive — the very bool the scheduler read, never re-evaluated. Surfaced in the HTML timeline title only; every other reporter deliberately ignores it.

R11-7 — the corpus-read rule is shared, its discovery is not (closed 2026-08-23 — discovery unified: one walker, one claims predicate; the surviving asymmetry is size measurement — fs::metadata vs text length — deliberate and documented, an unsaved buffer has no file to stat)

FragmentCorpus::unreadable_file now gives both readers one diagnostic, but the CLI walks the fragment root with std::fs while the LSP reads through its overlay provider. That difference is real — the editor must see unsaved buffers — so the readers stay separate. What is worth watching is that “which files are in the corpus” is still answered twice, and only the meaning of a failed read was unified here.

R11-8 — a runs-dir edited mid---watch fed the loop its own output (shipped)

R11-2 made each rerun re-read the config, so records went to the new runs dir while the watcher’s exclusion still named the one frozen at startup. Every rerun’s artifacts/*.hurl, now under an unexcluded directory, requeued the next run: 39 runs in 12 seconds, firing real traffic, from one edit. The third outing for this class, so the fix removes the second answer rather than resynchronising it — each rerun registers where it is about to write, before it writes, and the exclusion is derived from the same config the run is. Deliberately not a uuid-shaped exclusion: --run-id names a run directory that is not uuid-shaped.

R11-9 — a relative --config was never the file --watch matched (shipped)

The watcher compared the config by exact path while notify reports events under the spelling the OS resolved them to, so --config proef.toml matched nothing and config edits produced no rerun — silently, because feature edits kept firing and the loop looked alive. Two questions had been conflated: where a path points (answered once, lexically, when the flag is stored) and whether two paths are the same file (answered by comparing canonical forms, since absolute is not enough — macOS’s /var → /private/var aliasing and symlinks both survive it). The same relative path had been costing proef lsp --config go-to-definition across the whole corpus, because documents::name_to_url refuses a relative name.

R11-10 — doctor reported on defaults over a proef.toml that would not parse (shipped)

R11-3’s discovery arm became a silent unwrap_or_default, dropping the parse error the previous code printed: a malformed config left doctor reporting on invented defaults and printing “all checks passed”, exit 0. A project: row now, so it reaches worst and the exit code CI reads. Leniency still means absent — doctor must run outside a project — not broken.

Ingested — competitive research v2 (2026-08-16), validated claim-by-claim

An external research pass (prototyped against the built 0.12.0 binary) plus its round-14 companion review. Each actionable claim was re-reproduced here before anything was written down. Disposition:

S1 — an encoded reflection of a secret defeated redaction (shipped)

The one defect in the set, confirmed by live reproduction: a server reflecting the bearer token base64-encoded put dG9r… (trivially decodable) into an assert-failure detail; the raw needle never fired; the encoded credential reached the console and events.jsonl. The raw-form invariant was intact — this violated its intent. Shipped as derived needles inside Redactions::new (see the changelog and the ADR-0005 amendment); property- and mutation-tested, pinned end-to-end against a fixture introspection route.

Not covered, on purpose: hashed/split/re-encrypted reflections (not needle- matchable), double encodings (an unbounded tower; echo endpoints produce one level). The research doc’s companion ideas — a redaction-verifying scan over a finished run record, GitHub ::add-mask:: for captured secret-typed values, RF-style secret-typed macro arguments — are enhancements, not part of the defect, and await triage.

Corrections to the research set, so they are not re-litigated

  • S4 (Trusted Publishing plan) rests on a false premise: it plans a first publish with a classic token, but all four crates have been live on crates.io since 0.5.1 (0.12.0 current). Trusted Publishing can be configured directly against the existing crates; the token sequence is unnecessary.
  • S2’s exposure check is right and already satisfied: Cargo.lock carries curl-sys 0.4.90+curl-8.21.0, past the June-2026 CVE batch. The detection blind spot (RUSTSEC carries no advisories for *-sys-bundled C libraries) is real; the proposed libcurl-version print in release artifacts awaits triage with the rest.
  • R3-16/R3-17 (browser and Android engines) are foreclosed, not deferred: proef is API-testing-with-hurl only — a standing decision, not a gap the research reopens. The M6 line in CLAUDE.md is architectural readiness, with nothing scheduled. The seam-hygiene half of R3-15 stands on its own merits and awaits triage like the rest of the registry.
  • The round-14 review audited 214a39d (a pre-amend commit never pushed; what merged is c3ac752, differing by one deliberately-removed proptest seed), counted 464 tests where 462 exist, and credited #63 with the LSP corpus-holding change that shipped earlier — recorded here because review counts have now drifted by +2 for three consecutive rounds.

The R3 registry — triaged 2026-08-17

Triaged as a set against the PRD, the ADRs, and current industry practice, with each seam re-validated against the tree first. The v1 research document was confirmed absent (only v2 exists on disk), so items defined only there are one-line summaries with no spec — that fact drives several verdicts below.

Built:

  • R3-1 --max-fail N (shipped with this triage). The convention is universal — Playwright --max-failures, pytest --maxfail, nextest --max-fail — with one shared semantics: stop after N failures, un-run tests report as not-run rather than passed. proef’s seams made it a CLI-only change: a sink wrapper counts suite-scenario failures (the phase field keeps setup/teardown out of the count) and cancels the run token, which is the tested Ctrl-C drain path — in-flight batches finish, the rest record as skipped, teardown still runs on its own token, and the record is a complete cancelled run. That last part is free correctness: diff --fail-on-regression already refuses to certify a cancelled run, which is exactly right for a deliberately-partial one.
  • R3-4 diff takes a record path — shipped earlier (#65), with the research’s --baseline flag spelling declined as a second name for the same positional.

Build next (validated, in order):

  • R3-2 a flakiness verdict — (shipped as proef flaky). The 2026 pipeline is detect → quarantine → resolve, and proef already owned the middle step (@quarantine runs-but-does-not-gate); flaky is the missing detect, a fold over the records runs-dir already retains, so the history window is [run] keep-runs and no new state exists. Transition-counting separates flaky from broken (a mutation test proved the test suite could not initially tell that apart from a naive fail-rate — the F,F,P,P case now pins it), per-step attempt counts surface the pass-only-on-retry latent class, and a cancellation-skipped row is not evidence. No --check gate, deliberately — its sibling fragments has one, but a flakiness verdict is advisory by nature and @quarantine owns the gating decision; the asymmetry is a choice, not an omission, and the thresholds become contract (and move to proef.toml) only if a gating mode ever exists.
  • R3-3 sharding, hash-mode only — (shipped as --shard I/N). The measured stability argument held end to end: the mutation test swapped index-slicing back in and the insertion case (prepend, which shifts every position) caught it — the append case did not, which is itself the finding’s point. The assignment is frozen by literal-pinned tests; changing the hash is a breaking change to every sharded matrix. Filter→shard order pinned; an empty shard of a non-empty selection exits 0 with a note.
  • R3-6 JUnit attributes — (shipped, from the fresh spec the triage required). The spec was written from what the two consumers actually parse, at source level: GitLab’s docs enumerate testcase classname/name/file/ time plus suite and root time — and explicitly ignore the count attributes and timestamp; Jenkins’ SuiteResult.java reads suite name/package/id/time/timestamp and case classname, and never reads hostname. What shipped, and why:
    • Identity became classname + name — Jenkins keys test history on the pair, GitLab’s MR widget diffs head against base by it, and the old single name embedded file:line, so an edit above a scenario re-identified every test below it (a fleet of “new” tests on both tools). classname carries the feature file, name the scenario alone — unique per file by construction (outline instances are #N-disambiguated). Breaking for anything keyed on the old names.
    • file on the testcase (GitLab source linking), time on suite and root (both consumers), and the suite skipped count spelled skipped (quick-junit 0.5 → 0.7; 0.5 wrote disabled, which neither consumer reads).
    • timestamp and hostname deliberately absent — GitLab ignores both, Jenkins substitutes its own build clock and never reads hostname, and naming the machine would undo R12-1. Additive later if a consumer asks.

Deferred, with the trigger named:

  • R3-5 CTRF output — shipped as --ctrf (#160, 2026-09-02). The deferral read “a seventh format needs a consumer, not a trend”, against the six proef already emits (JUnit, TAP, JSONL, a GH summary, SARIF, HTML). What it had not weighed is that the seventh shares the JUnit fold, so it cost a renderer rather than a mechanism, and ADR-0019 quarantine parity plus real retryAttempts came with the fold. Corrected 2026-09-10 — it sat under a “deferred, trigger named” heading for the eight days after it shipped. R3-9, four bullets below in this same list, was annotated the moment it shipped — that is the convention this entry missed.
  • R3-18 generated pack documentation — when pack-vocabulary discovery becomes a reported adoption pain; the LSP currently serves that need interactively.
  • R3-15 pre-M6 seam refactors — when a second engine is actually scheduled (M6 has nothing scheduled; the snapshot corpus already provides the golden artifact-diff prerequisite).
  • R3-7 --affected-by, R3-10 fake variants — defined only in the absent v1 document; need the source or a fresh spec before any verdict.
  • R3-9 seeded shuffle — shipped as --shuffle (RF-audit wave 1). The old pointer here was dangling: IMPROVEMENT-PLAN #14 is the fakes seed and never mentioned order. The shipped form honors #14’s actual rule anyway — the permutation is seeded by the run id, no parallel seed.

Declined — do not re-raise (moved to the standing section’s rules):

  • OTel trace export (R3-11) and Cucumber Messages (R3-12). ADR-0008: the JSONL event stream is the record, no second record format. Both are re-encodings of the record for ecosystems that can convert from JSONL outside proef; building them in creates permanent format-tracking obligations against moving upstream schemas.
  • Browser/Android engines (R3-16/R3-17) — foreclosed by the standing hurl-only decision, not deferred.
  • S4’s first-publish token sequence — false premise; the crates have been live since 0.5.1. The worthwhile residue (crates.io Trusted Publishing for the existing crates, then the token-delete) is an owner-side dashboard action, recommended to the user rather than something the repo can do.

Open — adoption report on 0.12.0 (ingested 2026-08-14)

From a suite that ported to ref: at scale — 15 hurl files, 112 fragments, 21 scenarios — and ran 0.12.0 as an installed release. Three items, each reproduced here against the tree before being written down. Two shipped in the same change; the third is recorded because the report’s diagnosis was wrong even though its observation was right, and that distinction is the finding.

R12-1 — provenance named the machine that produced the record (shipped)

[run] suite resolves against the config directory (R11-1), so a path-less proef test handed the front end an absolute path and every emitter printed it: the .hurl # source: header, .map.json’s feature.file, every step_finished event, the console, and pack diagnostics. Two checkouts of one suite stopped producing equal artifacts, which is exactly the property ADR-0010 exists to guarantee.

Worse than R9-6 filed it. R9-6 says the portability claim “holds only from the project root”; this reproduces from the project root with the config in it. R11-1 was the right fix — one resolution rule — but resolution produces absolute paths, and nothing was named at the other end.

Shipped, and R9-6 with it. front::SourceNaming is the one naming boundary: resolve against the project, then name against the project again. A relative path is left exactly as it arrived (machine-independent already, and the caller’s own spelling, which their terminal can open); an absolute one is spelled relative to the config directory when it lies inside it. This also replaced the fragment corpus’s cwd-relative strip, which was a second anchor for the same question — the drift R9-6 predicted. The four ways to name one suite (derived, typed, typed absolute, from a subdirectory) now emit one artifact byte-for-byte, pinned by crates/proef-cli/tests/provenance.rs.

Two limits, deliberate: a corpus genuinely outside the project keeps its absolute name, because no project-relative one exists; and DiskSourceProvider (proef lsp) still yields absolute names, because it keys document identity on them.

R12-2 — the run-record ceiling was a constant no project could reach (shipped)

Retention was const RUN_RETENTION = 200 with only runs-dir configurable, and artifacts are byte-identical across runs of an unchanged suite — so a suite re-run on every save accumulated identical bytes for a day before anything signalled a ceiling existed. [run] keep-runs makes the policy expressible; 0 keeps none but the run in flight.

The report’s inference that artifacts should therefore not be stored is wrong, and it said so itself: an old record’s artifacts are what that run executed, and once the corpus changes proef artifacts no longer reproduces them. Bound the cost, do not drop the evidence.

Not closed by this, and not reported: rotation only ever deletes directories named by a generated run id, so --run-id <name> records sit outside the budget entirely. A CI minting a fresh id per build accumulates without bound. Guessing at user-named directories is the worse failure — runs-dir may be . — so this stays, documented in CONFIG.md rather than fixed.

R12-3 — a [run] setup test failure is invisible to JUnit (shipped)

Reproduced: a setup feature whose assertion fails exits 2 with summary: 0 passed · 0 failed · 0 skipped, and --output junit writes an empty report, because the abort precedes the reporter. A CI reading JUnit sees nothing at all.

The exit code is not the defect. ADR-0014 decided it explicitly — a setup failure maps to a user (2) or system (3) fault, never a test failure, “the same distinction Playwright draws between a clear setup error and a cryptic test failure”. Changing it needs a superseding ADR, not a bug fix.

Three of the report’s supporting claims do not survive checking, recorded so they are not re-litigated:

  • “teardown already has a distinct code; setup collapses both into one” — false. A teardown assertion failure exits 3, not 1: both phases map a test failure onto a non-test code (phase_failed(…, UserError) / …, SystemError). Neither distinguishes, by design.
  • “appears in nothing explain/diff consume” — false for explain, which prints failed (setup — excluded from the totals above) with the assertion detail and the artifact reference; the events are in the record with phase: setup.
  • “previously raised, still open” — no entry in this file matches it.

So the open item is narrow: the phase reporters run only for the pool. Worth fixing at the reporter, not the exit code.

Shipped at exactly that boundary: the CI-report block (JUnit, GitHub job summary, PR annotations) is one function both enders call, so a setup abort now writes the reports from the setup phase’s own summary — one testcase, failed, suite named by the setup feature file. On main the gap was worse than filed: no JUnit file was written at all (the finding said “empty”). Exit codes are untouched, per ADR-0014. Nothing is fabricated for the pool that never ran — the test pins that too.


Open — round-7 residue (ingested 2026-08-10)

The round-7 pre-merge review of PR #13 never entered any worklist; a round-8 revalidation re-reproduced its findings against v0.8.0. §2.2, §2.3, §2.4 and the diff item shipped in #30/#31. What remains, carried on that report’s evidence rather than re-reproduced here:

  • The early-error record — reproduced 2026-08-11, needs a decision. proef test --tags <nothing-matches> prints the error and then a summary: 0 passed · 0 failed · 0 skipped line, and the record it leaves is run_started + run_finished 0/0/0 — byte-indistinguishable from a clean run of an empty suite. A post-mortem reader cannot tell “errored before dispatch” from “ran nothing successfully”.

    The fix is a design call, not a patch. Suppressing the tail on this path would leave the record incomplete, which the tooling already banners correctly — but RunRecord emits its tail structurally, on Drop, precisely so no return path has to remember it, and adding an exception reintroduces the fragility that design removed. Opening the record later is blocked by setup, whose scenario events need it. The third option is an additive event carrying the early error (ADR-0008 permits it) — the most honest and the most work.

Open — residue of the two UX reviews

Verified against main on 2026-08-10. Everything else those reviews raised has shipped (first-run: F1, F3, F4a and F2’s did-you-mean in 0.6.0 · non-technical: N1–N5 and the init count in #24).

R1 — missing_config_var’s span points at the sentence, not the pack line

The diagnostic reports at the feature step that used the variable, e.g. suite/case.feature:3:5, rather than the pack line where ${url:bse} actually appears — so the reader goes hunting. The did-you-mean half shipped in 0.6.0; this half did not, deliberately.

Why it was deferred, in full — this is the whole reasoning, do not re-derive it: ResolveError carries no position, and resolve() is documented “pure and total”. The comparable diagnostic that does land on a pack line (pack::invalid_hurl) gets its position from hurl’s own parser reporting a line/column, which feeds locate::payload_line_span(…, rel_line); nothing computes a rel_line for a resolve failure. Supplying one means threading an offset out of a deliberately position-free pure function and carrying pack identity to the diagnostic site. That is a design change, not a fix — it wants its own spec.

Two sibling extensions were declined at the same time: resolve::missing_env must not suggest from the injected environment snapshot (it would surface unrelated environment variable names in diagnostics, against the secret-masking posture), and resolve::unknown_namespace already enumerates all seven valid namespaces. Sibling codes share a shape, not a candidate set.

R3 — the scaffold default is the dev fixture’s port (declined 2026-08-11)

init.rs writes base = "${env:PROEF_BASE_URL:-http://127.0.0.1:8787}", which is proef’s own dev fixture port — so to someone who installed a binary and has no fixture, the value looks configured and is not. The proposal was an obvious placeholder (https://api.example.com) to cover prevention, since a failing run already covers recovery.

Declined, with the reasoning recorded rather than a silent skip. Recovery is now covered on both halves: an unreachable target and untouched routes each get their own note (#28, #38). The remaining benefit is that the config file would read as obviously unfilled. Against that, init.rs’s module doc states the scaffold deliberately mirrors what GETTING-STARTED teaches — so changing the literal changes the tutorial too, and the tutorial’s “run it against xtask fixture with no PROEF_BASE_URL” flow stops working. That flow is a real onboarding asset for contributors. Trading a working tutorial for a more obviously-fake string is not worth it once the failure itself explains both halves.

Revisit if first-run drop-off is ever measured rather than reasoned about.

Decided against — do not re-raise

Recorded as decisions, so they are not rediscovered as fresh ideas.

  • Re-classify the unconfigured-scaffold failure from exit 3 to exit 2. Not a CLI-edge change: the verdict is set in proef-engine-hurl (classify_error’s _ => Infra arm), Fault::System(String) carries no kind to match on, and the exit derives in proef-core (RunSummary::exit_code_excluding). Both routes — string- matching the engine’s opaque message, or adding a structured kind to core’s public surface — cost more than the value, which is vocabulary. The note delivers that, and fires on the exit-1 placeholder-route path a re-classification would have missed.
  • Degrade proef flows the way macros degrades. flows promises every scenario; a list silently omitting the feature that failed to parse is a wrong answer, not a degraded one. macros degrades safely only because pack loading precedes binding and does not depend on it.
  • Ship proef-fixture in the binary so the scaffold’s first run passes. Needs a new ADR (it is dev-only today), enlarges the binary and the security posture of a test runner with a listening server — and R3 plus #24’s note remove the need.
  • A GUI, web UI, or “no-terminal” mode. PRD §3 forecloses dashboard/server mode. The P1 gap was always about vocabulary and error text, never a second interface.
  • Importing or round-tripping hand-written hurl, and anything OpenAPI-shaped as a recurring oracle. PRD §3 and ADR-0016 permanent non-goals.

Open — correctness

Q2 was the remaining Tier 1 branch (Q5 and Q4 shipped in #26); it closed with the #146 analysis cache — see below.

Q2 — the walk still happens twice per request (closed 2026-09-02)

Shipped in #27: the walk skips target/, node_modules/, vendor/ and dot-directories, is depth-bounded, and no longer aborts the whole discovery on one unreadable subdirectory (which analyze.rs swallowed into a silently empty analysis). Shipped in #32: the server adopts the workspace root the client announces — workspaceFolders, else rootUri, else the previous config-then-cwd resolution — so an editor launched outside the project no longer analyses the wrong tree.

Closed 2026-09-02 — by the #146 analysis cache, which this entry predated. The premise (“on every completion/definition/references request”) is no longer true: every request handler reads one cached Analysis through the single read path (server.rs — “the debounced diagnostics publisher and every on-demand feature go through here; they share one recompute per edit rather than one each”), edits mark the suite dirty behind a debounce, and the fragment corpus is held across recomputes (“called when a fragment file changes, never per request” — analysis.rs). The invalidation hook this entry said SourceProvider lacked turned out not to be needed: the whole analysis is invalidated on any edit, which at this suite scale (tens of small files, milliseconds per recompute) beats maintaining an incremental index — the module doc says so in as many words. The two walks inside one recompute remain, and are now a per-edit cost too small to file.

P5 — watch: the atomic-save half (remainder)

Shipped in #37: --watch now also watches proef.toml, matched by exact path.

Closed by inspection — the inspection was invalidated by a later change, and the bug shipped. The original argument was: the retrigger filter is an allowlist of .feature/.yaml/.yml, and no run-record file (.jsonl, .log, .hurl, .vars, .json, .xml, .html) matches it. ADR-0018 then added the engines’ fragment extensions to that allowlist — .hurl, named in this very paragraph as the thing that could not match — while every run writes .proef-runs/<id>/artifacts/*.hurl. A watched tree containing its own runs dir fed itself: 49 runs in 15 seconds, firing real traffic in a tight loop.

Now closed by construction, not inspection. The retrigger filter excludes generated trees by directory name, reusing discovery’s own skipped_dir, so there is one rule with two consumers rather than a second list to drift; the configured [run] runs-dir is passed in for the case where it is not a dot-directory. watch::tests pins both halves — that an emitted artifact never requeues, and that a fragment edit still does.

The lesson is the general one: a “closed by inspection” note records a conclusion whose premise nothing watches. This one even enumerated the fact that later became false. Prefer a test that would fail when the premise changes.

Still open: “a single watched file dies after an atomic save”. It did not reproduce on macOS/FSEvents; notify’s own docs say it is real but platform-dependent and worst on inotify. Do not chase it on a Mac — that is how it gets “fixed” by coincidence. It needs a Linux reproduction first.


Open — adoption and execution model (ingested 2026-08-11)

Source: a report written while porting a real 844-line hurl corpus onto proef — field evidence rather than inspection, which is why it found a different class from the review rounds. Every claim below was re-checked against main before filing; where the report was wrong, the correction is recorded with the item.

Already closed from it: the docstring-placeholder documentation gap (#41). Two of its claims did not survive checking, and both are noted in place (M1, D2).

The through-line. These are adoption, not correctness. The first-run path is finished and the correctness series closed its bug class; the next constraint is whether a team with an existing hurl suite can move onto proef and demonstrate they lost nothing. M1 and M2 are that story. E1 is the first wall a real suite hits afterwards.

F1 — proef.toml now has two path-resolution rules (closed — duplicate of shipped R11-1)

[run] fragments resolves relative to the config file’s directory (ADR-0018); suite, setup, teardown and runs-dir stay relative to the working directory. The reasoning that produced the new rule — the config is found by walking up, so a path in a config three levels above must mean “relative to the project” — applies verbatim to all five keys, and setup/teardown/runs-dir are consulted on every run rather than only when a path was omitted.

Cost: one file with two semantics and no marker distinguishing them. A user with setup and fragments in the same proef.toml gets one working from a subdirectory and one not, and every future path key re-litigates the choice against four precedents for the older rule.

Not fixed here on purpose. Changing the four existing keys is a behaviour change for every project that already relies on cwd-relative resolution, which is out of scope for the change that introduced the fifth. The fix is a single ProjectConfig::resolve_path used by every path accessor, shipped deliberately with a changelog note — recorded so it is a decision rather than an oversight.

ADR-0018 (named hurl fragments) lands into this section — read it against these items before assuming what it closes. It lets a pack ref: a named entry in a real .hurl file, so a corpus file is annotated once instead of transcribed, and stays runnable under stock hurl. Item by item:

  • M1 is not closed and must not be built concurrently — both touch fmt discovery. ADR-0018 requires the opposite of M1 at one entry point (directory discovery must never sweep .hurl into the pack formatter) while leaving M1’s actual ask untouched (an explicitly named .hurl may be canonicalized). Sequence them, either order, never at once.
  • M2 is not closed. ADR-0018’s integration test runs one fragment both ways, which proves a file is dual-runnable; it does not compare two suites’ result sets.
  • M3 is unanswered and now overtaken: the charter re-examination M3 asked for has happened (PRD §3 amendment) without the measurement it asked it to rest on. The amendment argues from the non-goal’s own rationale instead, and says so. Measuring the port cost is still worth doing — it now informs priority rather than permission.

Closed 2026-08-23 (premise false — the “not fixed here” above HAS since been fixed, as shipped R11-1). The exact fix this entry prescribed exists as ProjectConfig::resolve (config.rs:328-338): every path-valued key routes through it, its doc comment narrates this entry’s story, and CONFIG.md documents the one rule. This entry and R11-1 were the same finding filed twice.

M1 — fmt cannot canonicalize a standalone .hurl (closed — foreclosed by ADR-0018)

The report had this backwards and it is worth recording why. It claimed fmt refuses a file outside a pack, and proposed teaching it to accept .hurl as a small plumbing change. fmt in fact accepted any file and rewrote it — two defects fixed in #40, which now makes it refuse .hurl correctly, since applying YAML block-location logic to hurl syntax would be nonsense.

So the item survives but changes shape: making it real means teaching fmt to recognize a hurl file and run the block canonicaliser over the whole thing, with no hurl: key to locate. That is a feature, not a flag.

Why it still ranks first. It is what converts M2 from clerical to mechanical, and it is the cheapest unlock for the most valuable capability.

Closed 2026-08-23, without building it. This entry predates ADR-0018, which was accepted with the opposite principle: proef reads files it does not own, so it must never write them — fmt refuses fragment files (ADR-0018, “proef never writes”; carried as a hard constraint in CLAUDE.md). Building M1 would diverge from an accepted ADR without a superseding one. And the goal M1 served no longer needs it: it existed to make M2 mechanical — canonicalize both corpora, diff the text — but fragments removed the transcription M2 was guarding, so there is no ported copy whose equivalence needs proving. The file the backend team owns is what proef runs, pinned per-file by the both-runners test. Reopening this requires a superseding ADR, not a feature request.

M2 — no mechanical equivalence check between a hurl corpus and its proef port (deferred — trigger named below)

Verified when filed (diff now also accepts record dirs and .jsonl paths — R3-4/#65 — but still reads no hurl report); no path reads a hurl --report-json, which the pinned hurl 8.0.1 does emit.

Why it matters. The safe way to adopt proef is to run both suites until the new one is trusted. During that window nothing proves the two assert the same things, so the equivalence gate degrades to a hand-maintained mapping table reviewed once by a human — and that table is what a team’s decision to delete their old suite rests on.

Scope. Not the hurl-import non-goal in disguise (PRD.md:42). Import means reading .hurl and generating Gherkin. This compares two result sets, which is diff’s existing job with one more input format. The non-goal forecloses a direction of data flow, not the ability to check your own work.

Deferred 2026-08-23. The urgency rested on transcription drift — a port that could silently assert less than its original. ADR-0018 removed the transcription: a migrating team annotates the corpus it already has, and the same bytes run under stock hurl and under proef (fragments.rs pins it per-file against the fixture). What remains defensible is a results diff for the trust-building window when both runners run in CI side by side — diff’s job with hurl --report-json as one more input. Trigger: the first concrete migration that runs both runners and asks to compare outcomes mechanically. Building a seventh input format ahead of a consumer is the same mistake the CTRF deferral records.

M3 — the port cost has never been measured (closed — overtaken by ADR-0018)

PRD.md:42 makes hurl import a permanent non-goal, and that rests on persona P3’s “pastes between corpus and packs” (PRD.md:57) being cheap — which nobody has measured. A 14-file, 844-line port is the first real datum available. Recording the hours settles a recurring argument in one direction or the other: cheap vindicates the non-goal with evidence instead of assertion, expensive earns the charter a re-examination with numbers rather than opinion.

Closed 2026-08-23. The re-examination this measurement was meant to trigger happened: ADR-0018 narrowed the non-goal to generation and rewrote P3’s job from “pastes between corpus and packs” to “annotates once” — the exact charter change M3 said the numbers should decide. The two field data points stand recorded (an 844-line/14-file corpus ported by raw paste at 100% coverage; a 97-entry corpus that chose annotation and stopped the paste port deliberately), and no third answer would change a decision that has already been made and shipped.

E1 — no intra-run serialization primitive (report B1)

Verified. TECH-SPEC.md:313 — scenario ordering is “preserved for artifact naming, not execution order.” No serial tag or config key exists anywhere in core, cli, CONFIG.md or AUTHORING.md.

Why it matters. Real suites contain scenarios that mutate global state — the reporting corpus has two, one needing an empty database for absolute items[N] assertions and one installing a workflow definition governing everything created afterwards. Neither can run in a parallel pool, and proef offers no way to say so; the workaround is several CLI invocations driven by tag discipline in a Makefile.

Charter fit. Scheduling, not a new engine or execution mode — the orchestrator already decides what runs when, and [run] setup/teardown prove the surrounding concept is in charter. Those cover before and after the pool and nothing inside it.

Options. A reserved @serial tag, or [run] serial-tags = [...]. The config form is more explicit and keeps runner semantics out of the feature files — and E4 is an argument for it.

Shipped as [run] exclusive-tags, a tag expression rather than a list — the same language --tags takes, so group membership is answered exactly as selection is. Two corrections to this entry, both from checking before building:

  1. The filing describes one axis; the mature shape has two. cargo-nextest separates a group concurrency limit (max-threads, which bounds members against each other and leaves the rest of the pool running) from per-test weight (threads-required, which is what buys global exclusivity — they redefined it in 2024 precisely so limits “are never exceeded”, enabling mutual exclusion against all tests). Only the second is what was missing here, so only that shipped; a group table can be added later without breaking this key.
  2. Of the two motivating scenarios, only the first is a serialization problem. “Installs a workflow definition governing everything created afterwards” is ordering, which [run] setup already provides — a feature run once before the pool exists. Recorded so an ordering primitive is not built on the assumption that it was needed.

E2 — N invocations produce N run records, with no merge (report B2; consequence of E1)

Verified. Each run writes its own .proef-runs/<run-id>/ (TECH-SPEC.md:299).

E1’s workaround therefore yields N records, N JUnit files, N HTML reports, and pass/fail aggregation pushed onto the caller’s shell, while explain/diff operate per-run so a post-mortem reader must know which to open. Recorded as a consequence, not an independent item — solve E1 and this largely evaporates; solving it alone (a proef merge) treats the symptom.

Largely closed by E1 shipping: a suite whose isolation needs are expressed as exclusive-tags runs in one invocation, so it produces one record, one JUnit file, one report and one exit code. Kept open rather than closed outright because a suite may still split invocations for reasons E1 does not address (different environments, different --tags in separate CI jobs), and nothing merges those.

Shipped (2026-08-25) — the rerun half: run_started.rerun_of names the base; the rerun’s JUnit carries the base’s not-re-run scenarios (reconstructed from its record, exit code and totals untouched), and report overlays the base into a whole-suite page with a merged-view banner, degrading loudly when rotation ate the base. What remains of E2 is the original split-invocation case (different --tags in separate CI jobs), still open on its trigger.

Widened by the RF audit (2026-08-24): the class includes --rerun’s own CI story, which this entry never named — a rerun writes a new record whose JUnit/report contain only the re-run subset, so “the one JUnit at the end” of the standard retry workflow describes 3 scenarios of a 300-scenario suite. RF’s answer is rebot --merge. The proef shape, when built: overlay a rerun record onto its base at report/JUnit emission — composition over records, never a merged record file (ADR-0008); an additive RunStarted.rerun_of field would make records self-describing for it.

E3 — no per-scenario state reset hook (report B3)

Verified. [run] setup/teardown are whole-suite only, run once around the pool (CONFIG.md:63-64, 120-141).

Any suite against a real database wants before-each; today isolation is convention (title prefixes so scenarios do not see each other’s rows) and convention has no guardrail. proef knows nothing about databases, so “reset the DB” cannot be a proef feature — but framed as a feature file run before each scenario it is the same primitive as setup at a different scope, which is engine-agnostic by construction. The cost is real: it multiplies run time by scenario count and interacts with parallelism. This needs an ADR against ADR-0014, not a patch, and it may well be declined — deliberately rather than never asked.

E4 — nothing enforces tag-group discipline (report B4; record, do not build)

If E1 ships as a tag convention, a scenario added six months later lands untagged in the parallel pool and breaks isolation intermittently — the worst failure mode, because it reads as flakiness. A lint would have to guess which endpoints are global, which proef cannot know. Its value is as a marker: this is the follow-on cost of the tag form of E1, and therefore an argument for the config form.

D1 — no first-class requirement traceability

Verified. flows --format json prints one object per scenario (main.rs:128-137), which with tags like @FRD-3.1-create gets most of the way. Almost certainly a documented recipe rather than a feature — proef should not learn what a requirement is — but the recipe does not exist, so every team reinvents it and the capability is not advertised for this use.

D2 — report generation across N runs (premise partly corrected)

The report overstated this. It claimed a Makefile must capture the run id because proef needs proef report <run-id>; in fact run_id is optional and defaults to the latest run (main.rs:199-201), so the ordinary single-run case needs nothing captured.

What survives is the compounding with E2: with N invocations, “the latest” is one of N. Minor on its own, and listed because report-generation friction is felt by every CI integration rather than by one team.

Positive evidence — recorded so it is not undone

  • The raw-hurl paste path covered 100% of a real corpus. All 844 lines used only [Asserts] (75) and [Captures] (15) — no [Options], [Query], [FormParams] or [Cookies] — with seven ordinary predicates (==, exists, not exists, matches, count ==, >=, isString), every one passing through untouched. The strongest evidence yet for ADR-0004, and the kind of claim that gets doubted later.
  • proef macros printing sentences (#29) is load-bearing. The porting plan gated its prerequisite phase on it, purely to author 14 files of new prose.
  • --rerun (main.rs:120-122, re-run only the last run’s failures) fits conversion iteration exactly.

Suggested order (historical — every item now resolved or parked)

The order was M1 → M2 (adoption becomes provable) → E1 (dissolves E2). E1 shipped as [run] exclusive-tags; M1 closed against ADR-0018; M2 is deferred on a named trigger; M3 closed as overtaken. The two documentation items, C1 and C3, shipped in #43. E3, E4, D1 and D2 remain record-only — none blocks anyone today.


Closed — docs drift (2026-08-11)

Every item in this section shipped; the table above records which PR each landed in. Two did not reproduce when re-checked, and are recorded here rather than dropped, so the next reader does not spend the same time on them:

  • A2 — CONFIG.md was said to claim [env.<name>.run] overrides any section. It carries no such claim today: its precedence text names jobs specifically, which is what RunOverride actually allows.
  • B12 — the CHANGELOG’s 0.5.2 entry was said to lack a line about the directory-valued-phase hard error. It has one, first bullet under Fixed.

One half of A5 was deliberately not acted on: TECH-SPEC §11’s run-dir inventory lists the files a run generates, and the [run] setup/teardown features are inputs named by config, not run-dir output. The reviewer called this half “defensible-but- interpretive” and it is; report.html, which the inventory genuinely omitted, was added.


Open — maintainability and CI

IDFinding
B10The canary would chase a hurl prerelease (no semver filter) — shipped: the index parse (latest_stable_in_index) skips - versions, unit-pinned; build metadata needs no rule, crates.io refuses versions differing only by +meta
P12The matcher re-tokenizes per (step, pattern) pair on every bind (performance)
P13No World snapshot/restore proptest (moot — that API was removed 2026-07-29, ADR-0005 errata; the store’s live invariant is property-tested, no_guarded_secret_ever_enters_the_global_store); no CI workflow runs llvm-cov — the local half shipped 2026-09-07 (#173): just cover/cover-html/cover-lcov; the CI job stays a maintainer’s cadence/cost call and must be a ratchet, never a threshold (TESTING-STRATEGY §3)
Q1structured payloads unreachable (premise false 2026-08-23: they parse, validate through the engine seam, lower and skip the hurl emitter — pinned by tests; EngineLowering was a review’s name, never a symbol) — what survives: no registered engine claims a structured kind, so the path runs only under test fixtures
Q6html.rs re-derives the emitter slug; four file_stem() sites (closed 2026-09-02 — the count was exactly right, four production sites, and the fix is structural rather than descriptive: emit::feature_stem and emit::artifact_slug are now the one definition of each, called by the emitter’s own caller, the dispatcher’s spec naming, the report’s anchors/artifact links, and the editor analysis. The other premise had gone stale the other way: ScenarioOutcome.artifact_slug has carried the emitter’s naming to runtime consumers since round 19, so “the schema carries no slug” no longer forced anyone to re-derive)

Open — deferred during the v0.6.0–v0.8.0 correctness series

Found while fixing the above; each was validated and consciously left out of scope.

  • proef-harness PROEF_BIN/PROEF_HARNESS_SUITE — fixed in #19, but the same reader is now duplicated in proef-cli and proef-harness. Justified today (a binary crate cannot be depended on; these are the only two env::var callers in the tree). Tripwire: at a third caller, promote it to a shared crate.

  • Capture-name charset is narrower than hurl’s grammar, so an out-of-charset name is silently omitted from .map.json. Closed 2026-09-11: aligned with hurl_core’s key_string_text — any char::is_alphanumeric (Unicode, not ASCII) plus _ - . [ ] @ $. So user.id, items[0], @type, total$ and précis all parse as captures in hurl and were all absent from the sidecar. A leading [ stays refused because hurl refuses it too; {/} stay out because a templated name has no statically knowable text.

  • A # comment inside a [Captures] run — fixed in #16; the one/two-letter-method gap it exposed remains (is_method_line requires three characters, hurl’s grammar does not). Closed 2026-09-11, and the measurement found a second error in the opposite direction: the predicate also allowed -, which hurl_core’s method (read_while(is_ascii_alphabetic), non-empty, uppercase) does not. Too narrow on length and too wide on charset, each masking the other, which is how both survived from 0.1.0. The failure is a phantom row, not only a missing one: with the run left open across a short method, a header of the next entry reaches .map.json as a capture nobody wrote — the first version of the regression test missed exactly this, because a response line closed the run anyway and it passed against the defect.

  • key_line_spans’ flow-style undercount is guarded by convention, not types. Two callers guard it independently; a third would have to remember. Cheap hardening: have the primitive return a reliability flag. Closed 2026-09-11, one step past the prescription. A flag can be ignored; the scan is instead private behind a KeyLines value whose only accessors are paired_with(parsed) — the spans, and only when the counts agree — and sole() for a key that occurs at most once. There is no path to a positional list that does not state the count it expects, so the third caller has nothing to remember. Both existing guards became the call itself, and spans_reliable is gone.

  • Cross-scenario ${fake:*} coincidence — two scenarios can still draw the same value. Documented as a known limitation in AUTHORING/CHANGELOG/TECH-SPEC.

  • No corpus tier for engineered robustness fixtures. tests/ has zero custom-method entries and zero fenced blocks, so that bug class is pinned only by unit tests on private functions.

  • fmt’s tie-break (equal CRLF/LF → LF) now applies only to the trailing newline of a file that lacked one — per-line endings are preserved (#33). Lone-\r files are still unhandled: the splitter keys on \n, so a classic-Mac file is one long line.

  • normalize_pack keeps the skeleton verbatim by construction at each push, not by the algorithm’s shape. The “hurl blocks only” promise has broken three times (#18 line endings, #33 mixed endings, #40 trailing whitespace), each caught by an example pinning that one instance. #44 added properties — skeleton-only text round-trips byte-for-byte, and formatting is a fixed point — so a fourth over-reach now fails CI instead of shipping. The structural version would locate each block’s byte span and splice the canonicalized body back into the original text, making “bytes outside a span are never visited” a property of the shape. Not worth the rewrite for a small textual formatter; revisit if a fourth normalization rule is ever added to that loop.

  • The stdout latch’s single-reader test isolation is safe under the mandated nextest (one process per test) but is a convention, not an enforced invariant.

  • A disk filling mid-run still truncates the human console report without reaching the exit code (closed: the console latch shipped in the 2026-09-02 series (#160), to its own written design; the record’s own writer got the same latch in #168, and both reach exit 3 through escalate_environment_failures).

  • Absent-secret fallthrough (“an unset PROEF_SECRET_<NAME> still reads the store”) is load-bearing and pinned only by an integration test, not a unit test. Closed 2026-09-11: resolve_all carries unit tests for the fallthrough, for the override winning over a stored value, and for the neither-source error naming both remedies — each checked against a mutation that breaks it. The PROEF_KEY override supplies the key, so nothing touches a key file.

    Recorded because it cost a rewrite: the override test first claimed to prove the from_store.is_empty() early return by using a corrupt store, and deleting that return left the test green. load_store’s error reaches the caller only through names that needed the store, and a fully env-supplied run has none — so the early return is an IO saving, not an observable behaviour, and the test’s stated mechanism was not the one making it pass. Kept as a separate test that says so.

Versioning & release procedure

This document is the versioning policy and the release runbook. The README carries a summary; this file wins on detail.

Versioning policy

Scheme: SemVer 2.0.0. Pre-1.0 semantics, applied strictly:

  • MINOR (0.X.0) — any breaking change, or a coherent feature series/milestone.
  • PATCH (0.x.Y) — fixes and purely additive changes that break nothing below.

What counts as breaking (these are the public contracts, per the ADRs):

SurfaceBreaking examplesNon-breaking examples
CLI + exit codes (ADR-0009)removing/renaming a flag; changing an exit-code meaningnew flag; new subcommand
Pack schema (ADR-0004)removing a key; changing key semanticsnew optional key
Event wire schema (ADR-0008)removing/renaming a field or variant; changing schema semanticsnew variant; new field with a default (additive-only rule)
Canonical artifact format (ADR-0010)any change to emitted bytes (snapshot-locked)— (changes are inherently breaking; bump minor)
Engine seam (ADR-0002)changing EngineFactory/EngineSession/StepBatch/ScenarioCtx shapesnew defaulted trait method
Config fileremoving/renaming a proef.toml keynew optional key

1.0.0 is declared when the pack schema, CLI grammar, event schema, and exit codes are stable enough to promise MAJOR-only breakage. Until then, downstream consumers should pin minor versions.

Single source of truth: [workspace.package] version in the root Cargo.toml. Every crate inherits it (version.workspace = true); the workspace releases as one set, always. Never version a crate individually.

Orthogonal versions, not to confuse with the crate version:

  • The event schema version is the schema field in run_started (EVENT_SCHEMA_VERSION). It only moves on a semantic break of the stream — additive variants/fields do not bump it.
  • The hurl pins (=8.0.1) never move as a side effect of a release. Upgrades go exclusively through the canary + runbook (IMPLEMENTATION-PLAN §7, ADR-0003).
  • Toolchain policy: the pin tracks latest stable Rust but adopts a new minor only at its x.y.1 point release, ~3-4 weeks after x.y.0 (tools and third-party crates track latest immediately; exact pins like hurl outrank everything). A reviewer reading “latest stable” as “bump on release day” prompted writing this down (R18-2).
  • MSRV is the toolchain pinned in rust-toolchain.toml; it may rise in any MINOR release pre-1.0 and is not a separate contract yet.

Tags: annotated vX.Y.Z on main, linear history. Cadence: release when a milestone or a coherent series lands — not on a calendar.

CHANGELOG rules

Keep a Changelog 1.1.0:

  • ## [Unreleased] always exists at the top; every landed change adds a line there in the same commit series that lands it.
  • On release, Unreleased content moves under ## [X.Y.Z] - YYYY-MM-DD (with a short parenthetical theme) and a fresh empty Unreleased is left behind.

Release runbook

From a clean, green main (all gates local + CI).

main is protected — the release commit goes through a pull request, and the tag is pushed only after it merges. Do not git push origin main, and do not tag before the merge. git push --follow-tags is not atomic: git pushes refs independently, so a protected-branch rejection stops the branch while the tag still lands — and a tag is exactly what release.yml triggers on. That combination starts a release build from a commit that is not on main. It happened cutting 0.10.0; the run was cancelled and the tag deleted before anything published, but the recovery is avoidable and this ordering avoids it.

# 1. On a release branch, cut the changelog: move [Unreleased] → [X.Y.Z] - date
#    with a short parenthetical theme, and leave a fresh empty [Unreleased].
#    (There is no link-reference section at the bottom of CHANGELOG.md — nothing
#    to update there.)
git switch -c release/vX.Y.Z
# 2. Bump the version in the root Cargo.toml — BOTH places:
#      [workspace.package] version = "X.Y.Z"       (the crates' own version)
#      [workspace.dependencies] proef-core / proef-engine-hurl / proef-lsp
#        version = "X.Y.Z"
#        (the inter-crate pins — belt-and-suspenders for independent crates.io
#         publish; a stale pin no longer satisfies the bumped version and fails
#         resolution, so these move in lockstep with the line above).
cargo build --workspace                            # refreshes Cargo.lock versions
# fuzz/ is a separate workspace with its own committed lock, and the gates job
# checks it with --locked: refresh it too or that gate goes red on the release
# commit.
cargo check --manifest-path fuzz/Cargo.toml --all-targets
# 3. Full gates — the same set CI runs, so a green local pass predicts a green PR:
cargo nextest run && cargo test --doc
cargo clippy --all-targets --all-features -- -D warnings && cargo fmt --all --check
RUSTDOCFLAGS="-D warnings" cargo doc --no-deps --all-features --workspace
cargo deny check && cargo audit && cargo machete
cargo run -p xtask -- docs-check     # the gate a docs-touching release commit trips
zizmor .github/workflows/
# 4. Commit and open the release PR (no tag yet):
git commit -am "release: vX.Y.Z"
git push -u origin release/vX.Y.Z
gh pr create --base main --title "release: vX.Y.Z"

# 5. After CI is green and the PR is MERGED, tag the *merged* commit and push
#    only the tag. The squash merge creates a new commit, so tagging the branch
#    would leave the tag off `main`'s history.
git switch main && git pull --ff-only
git describe --tags --exact-match HEAD 2>/dev/null && echo "already tagged — stop"
git tag -a vX.Y.Z -m "proef X.Y.Z"
git push origin vX.Y.Z            # this, and only this, starts release.yml

If the release commit was made on main locally before branching, git pull --ff-only refuses afterwards: the squash merge superseded it. Confirm the merged commit carries the version bump, check git diff --quiet HEAD origin/main, then git reset --hard origin/main.

The tag push triggers .github/workflows/release.yml, which:

  1. builds release binaries for five targets (macOS arm64/x86_64, Linux arm64/x86_64-gnu, Windows x86_64-msvc — the Windows zip bundles the vcpkg DLLs; macOS links the SDK’s system libxml2 and vendors OpenSSL, so shipped binaries need no Homebrew), via cargo auditable (binaries stay scannable) with no cache restore (cache poisoning must not reach published artifacts), attesting SLSA build provenance per artifact (the repo is public, so this runs unconditionally);
  2. publishes the GitHub Release with the version’s CHANGELOG section and all five archives (asset names must stay in sync with the binstall metadata in the proef package manifest);
  3. regenerates Formula/proef.rb in the emrecdr/homebrew-proef tap (deploy-key auth via the HOMEBREW_TAP_DEPLOY_KEY repo secret) — only when the tag is newer than the version the tap already carries. That step is gated on nothing but “a tag was pushed” and rewrites the formula whole, so a tag pushed late or out of order would downgrade every brew upgrade; it now skips green instead, leaving the tap alone while the release still publishes. The formula installs the binary, its man page and the bash/zsh/fish completions; from 0.16.0 until 0.18.0 it installed only the binary, so Homebrew users silently got neither the man page nor completion while every other channel did. The render step now checks the archive for each file the formula claims to install, because nothing else connects the two.

A tag runs the workflow as it existed at the tagged commit, not as it exists on main — the ordinary push-event rule, and the one that decides what backfilling a missing tag actually does. Patching release.yml therefore protects future tags only: a tag cut on a commit older than a fix runs the pipeline without it. This is not theoretical here. Backfilling v0.15.0 (commit dated 2026-08-25) would have run that commit’s unguarded tap job and walked the published formula from 0.17.0 back to 0.15.0, defeating the forward-only guard added in 0.18 — which lives on later commits and could not apply. The backfill was done with gh workflow disable release.yml around the push for exactly that reason, then the Release created by hand. Disable the workflow before pushing any tag whose commit predates a release-pipeline fix.

workflow_dispatch runs build+attest only — a full matrix smoke without publishing. crates.io publication remains a deliberate manual cargo publish per crate in dependency order (core → engine-hurl → lsp → proef — proef-lsp before proef, which depends on it non-optionally) and is not automated.

Manual because it is the one step nothing undoes: a published version can be yanked, never replaced or re-uploaded. So publish from the tag, not from a working tree that merely resembles it:

git describe --tags --exact-match HEAD     # must print vX.Y.Z
git status --short                         # must be empty
cargo publish -p proef-core --dry-run --locked
cargo publish -p proef-core --locked
cargo publish -p proef-engine-hurl --locked
cargo publish -p proef-lsp --locked
cargo publish -p proef --locked

--locked throughout, so what ships is what the committed lockfile resolves. Only these four go: [workspace.package] publish = false is the default and each publishable crate overrides it, so proef-fixture, proef-harness and xtask are excluded by construction rather than by remembering to skip them. Each command waits for the registry before returning, which is what makes the next one resolvable.

The registry does not carry every tag. 0.15.0, 0.16.0 and 0.17.0 were tagged and released on GitHub but never published, so crates.io goes 0.14.0 → 0.18.0 (published 2026-09-09, from the tag, all four crates). Cargo resolves version requirements rather than sequences, so the gap costs a consumer nothing — it is recorded here so that a reader comparing git tag against the registry does not read it as a failed upload.

History

  • v0.1.0 — initial release (fresh history baseline, 2026-07-29)
  • v0.2.0 — deep-review correctness blockers (duplicate-request, body corruption, delay budget), panic containment, the output contract, and the author guides — breaking: --output json stream split, empty selections exit 2, event schema grew additively
  • v0.2.1 — review P0 (header grammar, pipe, filters, name dedup, secrets perms) + failure UX (hurl expected/actual, true error-line anchoring)
  • v0.3.0 — data-safety blockers (asset copy, run rotation, zero-entry false green), Then-step visibility with exact attribution, UserInput taxonomy (user mistakes exit 2), option caps + repeat budget, atomic locked stores — breaking: proef-core API pruned, when: skips on literal false, zero-entry packs fail validation
  • v0.3.1 — secret hardening: secret rm, PROEF_KEY CI override, the saveAs-vs-secret promotion guard, doctor store/key health, corrupt-store recovery, warned-step reasons on the console
  • v0.4.0 — external config & environments (proef.toml [url]/[vars]/ [env.<name>], ${url:}/${vars:}, --env/PROEF_ENV, ADR-0012), default suite path, and the competitive-review breadth pass — breaking: the pack root key templates: became macros: with no alias (ADR-0004 amendment)
  • v0.5.0 — the proef-lsp language server: diagnostics, completion, go-to-definition and references over the sans-IO core (ADR-0017)
  • v0.5.1 — LSP correctness: process-leak, malformed-request crash, broken-pack degradation, root-at-suite, overlay keying; use:/match: go-to-definition
  • v0.5.2 — CLI correctness: diff step-collision, truncated-run gate, setup double-run, the first EPIPE guard, overflow hardening, bare-filename path resolution, exit-130 documentation
  • v0.5.3 — closed-pipe safety: every remaining raw eprintln! in proef-cli routed through the EPIPE-safe guard (with a source-scanning drift test), and proef-lsp’s panic-recovery notice no longer kills the server it just rescued
  • v0.6.0 — first-run UX & run-record correctness: proef init, a did-you-mean for unset config variables, a next-command nudge; one run_started/run_finished pair per record with suite-only totals, truncated-record banners in report/explain, a real worker slot index — breaking: a scenario with no steps is now an error
  • v0.7.0 — record & artifact integrity: run_finished is the record’s last line again (a watchdog-abandoned scenario no longer appends past it), ${fake:…} values no longer repeat across a scenario’s steps, .map.json stops listing captures that were never made and stops dropping real ones, and a whitespace-only expect: is rejected instead of emitting an inverted span — breaking: proef_core::resolve::resolve takes a caller-owned occurrence counter and Resolution::fakes is gone
  • v0.8.0 — CLI output & exit integrity: an unreadable PROEF_KEY/PROEF_ENV/ PROEF_SECRET_<NAME> is a loud user error instead of reading as unset, the run.log tee no longer duplicates bytes on a short write, proef fmt keeps a file’s own line endings, report -o writes artifact links that resolve, and diff stops inventing flakiness for a step with no baseline — breaking: a failed stdout write now exits 3 where it exited 0, and a malformed environment variable exits 2 where it was silently ignored
  • v0.9.0 — tool-surface integrity & authoring guidance: values interpolated into LSP snippets, GitHub annotations and job-summary tables are escaped, proef fmt refuses a file that is not a pack and stops trimming the YAML skeleton, --sarif carries startLine so annotations land, --watch retriggers on proef.toml, --dry-run’s nudge echoes the run that was actually validated, a templated retry: stops under-counting the batch budget, --output json reports the real exit, a truncated record counts its warned scenarios, a failing run says when the scaffold’s routes are still placeholders, macros prints the sentence an author needs, proef lsp adopts the client’s workspace root, and AUTHORING documents docstring placeholders and the validation-catalogue pattern — breaking: proef secret set --value was removed in favour of --stdin (a secret in argv is visible to ps), and proef macros --output json’s pattern field changed from a boolean to string|null
  • v0.10.0 — named hurl fragments (ADR-0018): a step may ref: one # @proef <name> entry of a real .hurl file, values supplied by bind: at pack/macro/step scope, so the same bytes run under stock hurl and under proef; [run] fragments names the scanned root, a ref: step records the fragment it ran as file.hurl#name everywhere a failure is reported, and the editor completes bind: keys and jumps from ref: to the annotation — breaking: pack::load takes a &FragmentCorpus, PackSet::fragments is an Arc, LoweredScenario::secrets is a map, and LoweredStep/StepOutcome/ Event::StepFinished carry fragment
  • v0.11.0 — the adoption response: ADR-0007’s value caps reach fragment text (byte-identical [Options] exited 2 inline and 0 behind a ref:, then ran), proef fragments lists the corpus and names both ways a fragment dies with a --check gate, a bind: key nothing reads is refused with did-you-mean, doctor reports the corpus, init scaffolds both body forms, --config names the proef.toml to read, and [run] exclusive-tags runs a scenario with the pool to itself — breaking: FragmentScanner returns ScannedFile, AnalyzeCtx takes the corpus rather than building one per call, StepKindSpec carries an options recogniser, and ScenarioSpec carries exclusive
  • v0.11.1 — the gaps 0.11.0 shipped with: --config reaches proef lsp and --watch (it was honoured by the runner alone, so the editor reported every ref: as unknown in exactly the layout the flag exists for), proef fragments counts [run] setup/teardown usage instead of calling a phase-only fragment unreachable and failing --check, one predicate answers “is this a fragment file?” where three disagreed, and --junit/--sarif/report -o create the directories their paths name — as artifacts -o and the run directory already did, and as pytest, jest-junit, cargo-nextest and the embedded hurl all do
  • v0.12.0 — one path rule, and a watcher that stops lying: a path written in proef.toml resolves against the config, a path typed on the command line against the working directory, with no exceptions — which finally inventoried .proef-state.json, .proef-secrets.json and the run records, all three cwd-anchored and unlisted. --watch rereads the config it retriggers on; stops feeding itself when runs-dir changes mid-loop (one edit produced 39 runs in 12 seconds against a live API); and matches a relatively-typed or symlinked --config, which also restored proef lsp --config go-to-definition across the fragment corpus. doctor fails a proef.toml that will not parse instead of reporting on invented defaults, --config is honoured or refused by every subcommand, and [run] exclusive-tags validates itself in both paths — breaking: the secret store, the World and the run records move with the config rather than the shell, which reaches anyone who ran proef from a subdirectory
  • v0.13.0 — a record that travels, and a secret that stays one: nothing proef records names the machine that produced it (one naming boundary, the dual of the path rule — breaking: artifact bytes change for path-less runs), and a secret reflected base64/hex/percent/JSON-escape-encoded is redacted like its raw form (live leak reproduced, then closed; ADR-0005 amended). The fragment corpus read is bounded (601 MB → 15 MB on the measured pathological input), fuzzing actually reaches the fragment rules (probe-verified), [run] keep-runs makes retention expressible, diff takes a record path for the CI-baseline flow, the bundled libcurl gets a CVE floor no advisory scanner would catch, and a hung test is a five-minute failure instead of a five-day zombie
  • v0.14.0 — proef at CI scale: --max-fail N stops a run honestly (the never-run tail records as skipped, the record is a cancelled run diff refuses to certify), --rerun continues a cancelled run instead of a false green, proef flaky folds the retained history into verdicts (flapping by transition-count, passes-only-on-retry, broken-not-flaky) completing the detect→quarantine→resolve loop the @quarantine tag already anchored, and --shard I/N partitions a matrix by a frozen identity hash so adding a scenario never re-buckets the others — plus the reverse docs gate: every subcommand must be documented, enforced rather than noticed
  • v0.15.0 — validation rounds 17–18 + the Robot Framework capability audit: @skip/@skip:reason and @quarantine visible in every sink (ADR-0019), tag globs + per-tag report verdicts + [tag-links], explicit run metadata (--meta/[meta], ADR-0020), the rerun overlay (one JUnit and one report covering the whole suite), --console dotted|quiet, --shuffle seeded by the run id, reproduce_hint into the record — breaking: quarantined failures reach JUnit as skipped-with-message, --shard re-deals (the hash gained fmix64), tag atoms glob, JUnit identity is classname+name. Its tag was missing for two weeks (found 2026-09-09): the release commit landed 2026-08-25 but v0.15.0 was never pushed, and release.yml starts on the tag alone — so the pipeline never ran and 0.15.0 had no GitHub Release, binaries or attestations. Backfilled 2026-09-09 with the workflow disabled for the push: the tag now points at the release commit and the Release carries the changelog section, marked not-latest, with no archives — the only release without them. Not repaired by simply pushing the tag; the runbook above says which workflow a tag actually runs
  • v0.16.0 — the surfaces tell the truth: an eight-wave improvement programme (#112–#142) plus the round that found what it missed (#143–#150). CI-sink conformance (JUnit detail into element content, an XML-1.0 control-character boundary, real limits on the GitHub summary and annotations), a triageable and linkable HTML report, console colour, shell completions and a man page in every archive, a project-aware doctor, explain/diff/doctor --format json, --console failed, flaky --by, proef schema config, and the LSP wave — document symbols, hover, quick-fix code actions off a structured Diag::fix, one analysis per edit rather than per keystroke, and a panic guard on both message-loop entry points. Pack validation became linear in the macro count (65× at 3200 macros) and the last unfuzzed parser gained a target — breaking: --output split by meaning into --format (which format) and -o/--output (which path), World::set_global returns a #[must_use] bool, ConsoleReporter::new takes a color flag
  • v0.17.0 — the environment a suite runs in, and the guards that keep its claims true: [http] gained the keys that describe an environment rather than a request (TLS insecure, proxy, mTLS cert/key, max-redirs, user-agent, cookie-store = false), --ctrf renders the run off the same fold as JUnit, --shard-weights balances a matrix by measured duration from one shared timings.json, and the HTML report answers “what is slowest”. The hurl-coverage audit closed the two defects a ref: fragment could not work around (#164–#166): a file,…; body resolves beside the file that wrote the reference, and two features’ same-named assets stop overwriting each other. A --run-id record is findable again (ADR-0021), a disk filling mid-run reaches the exit code, and the 23 diagnostic codes that had no test got one — breaking: emit::file_references became Artifact::assets carrying each reference with the source that wrote it, emit::asset_root is new, HttpDefaults gained eight fields and lost Copy, and the canonical artifact format moved (an artifact that reads a file now names its --file-root in the replay line)
  • v0.18.0 — the CI-consumer surfaces, run to exhaustion (#168–#179): output proef could not deliver never looks like success. A run-record write failure latches into exit 3 through one fold (escalate_environment_failures, beside the JUnit/CTRF and GitHub-summary failures), SIGTERM/SIGHUP take the graceful cancel so a CI job timeout leaves a complete record and its reports, and a custom --run-id no longer collapses the JUnit identity onto the nil uuid. Asset staging resolves beside the file the parser read wherever you cd from, with --sarif lines from the carried source and the symlink and case-insensitive staging edges closed. The ADR-0007 budget family is closed over its inputs and bounded as a product (a four-hour batch ceiling, [http] timeout-ms = 0 refused), every RunSummary sink routes identities through the masker, and proef flaky gained the 2026 statistical guards — a sample floor, hysteresis, an environment-outage guard, and an input-fingerprint equivalence class — breaking: proef_lsp::RootResolver returns a ResolvedRoot, timings::render takes a &Redactions, and flaky’s new verdict is renamed insufficient-data with its default sample floor rising from 2 to 10
  • v0.19.0 — the checks that could not see what they claimed to cover (#186–#191). The Homebrew formula installs the man page and the shell completions again — broken since 0.16.0 because the formula is a heredoc in release.yml and the archive is staged in another job, so nothing tied the two together; the render step now fails if the archive lacks a file the formula installs, and this tag is the first to carry it. --dry-run refuses a file,…; asset that is not there, through staging’s own checker rather than a second walk, so the gate CI runs before standing an environment up is no longer blind to a defect that is entirely static. The asset scan moved behind the engine seam: it was a scan for the literal "file," in proef-core that the ADR-0002 guard structurally could not classify — so it was never sanctioned and never reported missing, and that ADR’s “thirteen literals” measurement is corrected to fourteen — while reading hurl’s own AST also stops file, inside a JSON body counting as an asset. A fragment’s assets stage from where its file was read rather than from where its recorded name points, the fragment-side twin of what LoadedFeature::read_from already does. --rerun reads its base record once instead of twice, and the one doc check that only reads files moved into the half of the gate that only reads files — breaking: proef_core::emit::emit takes the registered step kinds and StepKindSpec gains an assets hook, replacing emit::file_refs_in

Contributing to proef

Small, focused PRs against main. The corpus in docs/ is the source of truth — CLAUDE.md is the working summary, docs/TECH-SPEC.md and the ADRs win on conflict.

Setup

The toolchain is pinned by rust-toolchain.toml (latest stable; rustup picks it up automatically). One-time tools:

cargo install cargo-nextest cargo-deny cargo-audit cargo-insta just
# plus, for the full CI surface locally:
cargo install cargo-fuzz cargo-public-api cargo-machete cargo-llvm-cov   # llvm-cov: `just cover`

Native build prerequisites (only proef-engine-hurl needs them): Debian/Ubuntu apt install build-essential pkg-config libssl-dev libcurl4-openssl-dev libxml2-dev libclang-dev; macOS: Xcode CLT.

The gates (green before every commit)

cargo nextest run                     # all tests
cargo test --doc                      # doctests (nextest skips them)
cargo clippy --all-targets --all-features -- -D warnings
cargo fmt --all --check
RUSTDOCFLAGS="-D warnings" cargo doc --no-deps --all-features --workspace
cargo deny check
cargo run -p xtask -- docs-check      # indexes ↔ reality
cargo run -p xtask -- public-api      # proef-core API surface (nightly rustdoc)
cargo check --manifest-path fuzz/Cargo.toml --all-targets --locked   # fuzz/ is its own workspace
cargo machete                         # unused dependencies

just gates runs the set above. CI additionally runs the #[ignore]d complexity guard alone (just perf — TESTING-STRATEGY §7), zizmor, a proef doctor smoke, a fuzz smoke, the hurl canary, a Windows gate, the docs-site build, and — on a pull request — the changelog self-recording check; cargo audit and the full fuzz run nightly.

Rules that are easy to trip over

  • hurl pins are exact (=8.0.1, built --locked). Never bump them in a PR — upgrades go through the canary + runbook (ADR-0003).
  • Snapshots are deliberate. Artifact bytes, sidecars, diagnostics, and event streams are insta-locked; cargo insta review each diff and be able to say why it changed. Never blind-accept.
  • proef-core API is snapshot-locked (crates/proef-core/public-api.txt). An intended surface change regenerates it: PROEF_PUBLIC_API_UPDATE=1 cargo run -p xtask -- public-api.
  • Core purity: proef-core does no IO and reads no clocks/env/randomness — inject values instead. This keeps every snapshot deterministic.
  • New architectural decision → new ADR (docs/adr/ADR-00NN-*.md, next number, same format) in the same PR. Diverging from an ADR without a superseding one is a bug.
  • New diagnostic → index it in docs/DIAGNOSTICS.md, prefer a seeded case under tests/errors/<area>__<name>/ (dry-running that corpus fails by design).
  • YAML is serde_norway, datetime is jiff, and reqwest/async-trait/ a tokio runtime are banned (see CLAUDE.md for the full list and why).
  • No raw print macros in proef-cli. Use crate::render::outln! for stdout and crate::render::errln! for stderr. println!/eprintln! panic when the write fails, and a closed pipe (proef … | head) surfaces as EPIPE rather than a signal — so a raw macro aborts with 101, outside the typed 0/1/2/3 exit contract (ADR-0009). A source-scanning test enforces this.

Testing

docs/TESTING-STRATEGY.md is normative. In short: everything is device- and network-free except the fixture integration suite (cargo run -p xtask -- fixture runs the dev API server standalone). Assert attempt counts and normalized event order, never wall-clock. just cover measures line coverage on demand (cargo-llvm-cov; advisory, never a threshold — TESTING-STRATEGY §3).

Commit messages

Conventional prefixes (fix:, feat:, docs:, refactor:, release:), imperative subject, body explains why. No AI-attribution footers.

Security policy

Reporting a vulnerability

Use GitHub’s private vulnerability reporting on this repository (Security → Report a vulnerability). Please do not open public issues for security reports. You will get an acknowledgment within a week; fixes ship as patch releases with a CHANGELOG entry.

Supported versions

The latest released 0.x version. Pre-1.0, fixes are not backported.

Threat model

proef is a test tool with a deliberately modest threat model: it protects secret material at rest and keeps it out of every output, and it does not attempt to defend a compromised host.

What proef guarantees:

  • The secret store (.proef-secrets.json) holds only XChaCha20-Poly1305 ciphertext (enc:v1: envelope) — safe to commit and share.
  • Secret values never appear in any sink: artifacts carry {{name}} placeholders, events/logs/reports are value-redacted at the sink boundary (property-tested) — the event stream through one exhaustive apply_event, and the CI sinks that render from the run summary (JUnit, CTRF, TAP, timings.json, the GitHub summary and annotations) through its twin apply_outcome, so a new text field cannot ship unmasked — and a saveAs: global capture whose value equals a known secret is refused rather than persisted to the plaintext .proef-state.json.
  • Sensitive files (.proef-secrets.json, the key file, .proef-state.json) are created 0600, private from the first byte. proef doctor warns when permissions have drifted.
  • Request file bodies are confined: every file,…; asset is staged from beside the source that names it into the run’s per-scenario asset root, which is the engine’s context_dir sandbox. A reference must be a plain relative path (no leading /, no ..); a symlink already sitting at a staging destination is replaced rather than written through; and two references that are one file to a case-insensitive filesystem are refused rather than last-writer-won.
  • A fragment corpus is read, never written (ADR-0018). Pointing [run] fragments at .hurl files somebody else owns is one-directional: proef fmt refuses them in both discovery branches, and the declared root is the confinement boundary — nothing outside it is scanned. Files come back byte-identical, which an integration test asserts.
  • A renamed secret is still never materialized. bind: { token: "${secret:x}" } lets a foreign corpus keep its own variable name; the value still travels via insert_secret and never enters the artifact. Mixing a secret into a larger bound value is refused (lower::secret_in_composite_bind) rather than quietly written out, because injecting the joined string would require putting it in the artifact.
  • TLS verification is on unless a profile says otherwise, and saying so is loud. [http] insecure = true exists because staging environments really do present self-signed certificates, but a suite that goes green without verifying one has not proved what a green suite normally proves. Every run with it active prints a warning naming the profile that set it. The run record deliberately carries no config, so that warning is the whole audit trail — which is why it cannot be suppressed.
  • mTLS credentials are file paths, not values. [http] client-cert / client-key name files; proef reads no key material into its own memory and writes none into any artifact. A client-key without a client-cert is exit 2 rather than a silent pass-through: libcurl would accept the pair and then present nothing, so the failure would surface at the server as an authentication error naming nothing about the cause.
  • Release binaries are built with cargo auditable (dependency trees stay scannable) on cache-isolated CI runners.

What proef does not defend against:

  • A compromised host or user account: the key file lives on disk, decrypted values live in process memory (no zeroize — hurl holds its own copies), and PROEF_KEY/PROEF_SECRET_* are readable from the process environment.
  • Malicious suites: packs execute arbitrary HTTP requests by design; run suites you trust.
  • Credentials written into proef.toml. There is deliberately no [http] user or netrc key — a password belongs in the secret store, where it is encrypted at rest and masked out of every sink. [http] proxy is the one edge: a proxy URL embedding credentials is plaintext in a file you probably commit, and proef cannot mask a value it was never told is a secret.

If your environment needs more than this, inject secrets per run via PROEF_SECRET_<NAME> from a real secret manager and skip the store entirely.

Changelog

All notable changes to this project will be documented in this file. The format is based on Keep a Changelog; versioning follows SemVer (policy in docs/RELEASING.md).

Each release groups its entries under one heading per kind, in the order Added · Changed · Fixed, with Breaking / Internal / Documentation after them where a release used those. Three shipped releases carried the same heading two or more times — a release cuts by moving [Unreleased] wholesale (RELEASING.md), so whatever shape it had at the time shipped verbatim. Regrouping preserved every entry and its order within its kind.

[Unreleased]

Fixed

  • The sidecar’s two scanners now use hurl’s own grammar rather than an approximation of it. Both were read off hurl_core’s parser and corrected against it, and both errors cost rows in .map.json — a normative artifact whose contract is that no legitimate row is dropped and no invented one appears (ADR-0010).

    is_method_line demanded three characters while hurl_core’s method parser takes one or more ASCII uppercase letters, so a short method opened an entry proef’s capture scan did not see: the previous entry’s [Captures] run stayed open across the boundary, and a header of the next entry (X-Trace: abc) was recorded as a capture nobody wrote. The same predicate allowed -, which hurl’s grammar does not, so a dashed uppercase word could end a capture run on a line hurl would refuse to parse as a request — the two errors pulled in opposite directions and hid each other, which is how both survived from 0.1.0.

    A capture name was matched against [A-Za-z0-9_-] while hurl’s key_string_text admits any char::is_alphanumeric — Unicode, not ASCII — plus _ - . [ ] @ $. So user.id, items[0], @type, total$ and précis all parse as captures and were all silently missing from the sidecar. A leading [ stays refused, matching hurl, and {/} stay out deliberately: a name written as a template has no statically knowable text to write a row for.

Internal

  • The secret-resolution order is pinned where it is decided. “An unset PROEF_SECRET_<NAME> still reads the store” is load-bearing — every run that keeps its secrets in the committed store depends on it — and was asserted only by an integration test that stands up a fixture server, executes a suite and checks exit 0. That test does catch a regression, indirectly and by way of an exit code. secretstore::resolve_all now carries unit tests for both directions of the precedence and for the neither-source error, each verified against a mutation that breaks it.

    One of those tests had to be rewritten first. It claimed to prove the from_store.is_empty() early return by pointing at a corrupt store, and deleting that return left it green: load_store’s error only reaches the caller through names that needed the store, and there were none. The early return is an IO saving, not an observable behaviour. The corrupt-store case is kept as its own test, saying that.

  • A located-lines undercount can no longer reach a caller looking complete. locate::key_line_spans returned a bare Vec<Span>, and it sees only block-style key: lines — a flow-style - {use: base} item is valid YAML, parses to a real step, and contributes no line. So the list can be shorter than the items it describes, and pairing them positionally attributes every span after the gap to the wrong item: a go-to-definition landing on the neighbouring line, a diagnostic pointing at it.

    Both callers already knew, and each had written its own length comparison in its own words from a prose warning. Both were correct; neither was enforced, and a third caller would have had to rediscover the hazard and the remedy together. The scan is now private behind a KeyLines value whose only accessors are paired_with(parsed) — which yields the spans only when the counts agree — and sole(), for a key like match: that occurs at most once and has no sequence to pair against. No behaviour changes; what changes is that the guard is the only way through.

Documentation

  • 0.19.0 is recorded where the corpus says it should be. The release History in RELEASING.md, the milestone Status in CLAUDE.md, the corpus index’s “through vX.Y.Z” line, and OPEN-FINDINGS’ own re-check date. This is the set that drifted after 0.15.0–0.17.0 — three tags with no History entry, concealed by a fourth filed out of order — so it is done in the same session as the tag rather than left for the next reader to discover.

[0.19.0] - 2026-09-11 (the checks that could not see what they claimed to cover)

Fixed

  • The Homebrew formula installs the man page and the shell completions. Its def install was bin.install "proef" and nothing else, so from 0.16.0 — the release that started shipping proef.1 and five completions/ files in every archive — until 0.18.0, brew install proef gave no man proef and no tab completion, while binstall and a direct download gave both. The formula is a heredoc inside release.yml and the archive is staged in a different job, so nothing tied the two together and no gate could see the gap; the render step now fails if the archive lacks a file the formula installs, and the formula’s own test do asserts the man page and completion landed. Takes effect on the next tag: a tag runs the workflow from its own commit.

  • --dry-run refuses a file,…; asset that is not there. A suite whose asset had been deleted reported dry-run OK, and the failure arrived later from a different command, against a live backend — from the one gate CI runs before standing an environment up. Whether an asset resolves is statically knowable, so it is answered there now. The checker is staging’s own (assets::resolve_assets, split out of stage_assets) rather than a second walk over the same artifacts, so validation and the run cannot disagree; the message and the diagnostic code are the ones a run already gave.

  • file, inside a JSON or assertion body is no longer mistaken for a file asset. The emitter found the files an artifact reads by scanning its text for the literal file, and a closing ;, so a request body containing that substring — {"note": "see file,notes.txt; for details"} — produced a phantom asset, and staging then failed the run over a file the request never reads. The claiming engine now reads its own AST, where a body reference and six characters of prose are different things.

Breaking

  • proef_core::emit::emit takes the registered step kinds, and StepKindSpec gains an assets hook. Asset recognition was hurl’s body grammar living in proef-core: emit::file_refs_in scanned for the literal "file,", which ADR-0002’s amendment forbids and — worse — which the guard pinning that amendment could not see. engine_grammar_kind classifies fences, HTTP, [Section] headers, method lines and key: value options; a body constructor is none of those, so the literal was never sanctioned and never reported missing. The ADR’s own measurement said thirteen literals; it was fourteen.

    The scan moves behind the seam as StepKindSpec::assets, the fourth engine-contributed hook beside validate, fragments and options, and the guard gains a body arm so the shape is classifiable whether or not anything currently uses it. emit() takes &[StepKindSpec] to reach it; FrontEnd carries kinds beside the kind_to_engine table it is built with, which registry already documents as a pair that must not be re-derived separately. emit::file_refs_in is gone.

Internal

  • A fragment’s assets stage from where its file was read, not from where its name points. AssetRoots::source_dir rebuilt a fragment’s directory by splitting file.hurl#name and joining the file half onto the project root — the naming boundary run backwards, without the canonicalize fallback that boundary carries precisely because a lexical-only version already shipped a bug (a suite reached through a symlink silently failed to match, R11-9). The two agreed only because both were seeded from config.root() and discovery walked from that same root, so only the lexical case was ever exercised, and nothing made them stay inverses. The corpus reader now records the directory it read each file from (front::CorpusDirs, carried on FrontEnd beside kinds), and staging looks it up — the fragment-side twin of what LoadedFeature::read_from already does for features, so both halves of the naming boundary are one-way in the same way. AssetRoots loses its project field and build_specs its project_root argument: with nothing to recompute, the project root is no longer staging’s business.

  • --rerun reads its base record once. It called record::read_events for the JUnit overlay and then record::rerun_candidates, which read and deserialized the same events.jsonl a second time — two full passes bounded only by the 256 MiB record ceiling, over a file another process may still be writing. rerun_candidates now takes the &[Event] its caller already holds, which is the rule read_record’s own documentation had already stated for exactly this case. The read error is handled once as well: the first call swallowed it with .ok() and the second rediscovered it a line later.

  • The one doc check that reads only files now runs in the half that reads files. no_current_behaviour_doc_spells_a_format_as_an_output_path lived in tests/docs.rs, whose stated charter is the checks needing a built binary to ask clap — this one only scans markdown, so it never ran in the fast doc-only CI step. It is now xtask docs-check’s check_output_path_spelling, reusing living_docs() instead of carrying a second directory walk. Its allowlist-shrink guard got stricter on the way: it counted ADRs into the same total, so a renamed entry could be masked by docs/adr being larger than the shortfall — which is the one failure that guard exists to catch. All three paths were checked by mutation: a stale spelling planted in an allowlisted doc, one planted in an ADR, and an allowlisted doc renamed away.

Documentation

  • The worklist stops contradicting what shipped. Three entries in OPEN-FINDINGS still called CTRF declined or its trigger unfired — the 2026-08-31 external re-test, the RF audit’s deferred list, and R3-5 under “deferred, with the trigger named” — for the eight days after --ctrf actually shipped (#160). R3-9, four bullets below R3-5 in that same list, was annotated the moment it shipped — the convention the three missed. Two more claims had outlived their facts: the shipped-changelog duplicate headers (no release carries one now, and check_changelog_kinds fails if one returns) and the machine-side note about Homebrew’s Rust shadowing rustup. Filed at the same time: a_second_interrupt_hard_exits_with_130 failed once on Linux CI and passed on a re-run of the same commit, so the evidence, the mechanism and the fix shape are written down instead of left to the next re-run. And the stance that a scenario-level @retry is deliberately absent — retry-until-green hides a one-in-four defect 99.6% of the time — is stated in TESTING-STRATEGY §5, which the worklist asked for and nobody had written.

  • The runbook records that the registry skips three versions. 0.15.0–0.17.0 were tagged and GitHub-released but never published, so crates.io moves 0.14.0 → 0.18.0. Noted in RELEASING.md so the gap does not read as a failed upload. The long-standing homepage question in OPEN-FINDINGS is also resolved: the field reached the registry with 0.18.0, exactly as that entry predicted; documentation remains unset and still open.

  • The release history records every release again. RELEASING.md’s History section carried no entry for v0.16.0 or v0.17.0 and filed v0.15.0 between v0.13.0 and v0.14.0; the order is repaired and all three versions are present, v0.18.0 included. The corpus also stops calling the 0.18 series unreleased, and an IMPROVEMENT-PLAN pointer into CHANGELOG [Unreleased] now names the releases that actually carried the work — [Unreleased] has been cut several times since that sentence was written.

[0.18.0] - 2026-09-09 (the CI-consumer surfaces: output proef could not deliver never looks like success)

Added

  • SIGTERM and SIGHUP now take the graceful path (ctrlc’s termination feature): a CI job timeout or docker stop cancels the run — in-flight batches finish, the rest record as skipped, teardown runs, the reports are written, and the record closes with a cancelled run_finished — where it used to kill the process mid-write and leave a truncated record with no tail. A second signal still hard-exits 130 (the handler carries no signal identity, so the code is 130 for every second signal). Pinned by sigterm_cancels_gracefully_and_the_record_completes and — for the first time anywhere — an exit-130 assertion, a_second_interrupt_hard_exits_with_130.

  • test --format json and explain --format json now report warned and cancelled. A warned scenario (an optional: step failed, or a saveAs: global promotion was refused) folded into passed, and cancelled — in the record’s run_finished — was surfaced by neither, so a script could not tell a spotless run from one with warnings, nor a complete run from a cancelled one, and the two JSON surfaces disagreed on how to say “did not finish” (0.18 survey). Both keys are additive and always present. warned also becomes visible in JUnit (a <system-out> note, the status stays success since JUnit has no warned) and CTRF (an extra.warned flag) — it was previously visible only in the HTML report.

  • A tag that looks like a reserved one but is not exactly it now warns (proef::tags::reserved_tag_typo). @quarantined, @skipped, @Skip matched no reserved tag and silently did nothing — a scenario the author believed was quarantined gated the build. The warning names the spelling it likely meant, tuned to catch the real typos without firing on legitimate short tags (ship, slip, step).

  • proef flaky gains the 2026-field statistical guards (0.18 survey §6), each a pure fold over the JSONL history already retained — no new state, no gating mode (advisory stays the design):

    • A minimum-sample floor (--min-samples / [flaky] min-samples, default 10): below it a scenario is insufficient-data rather than classified, because a verdict on thin data is worse than none.
    • Hysteresis (--recovery-runs / [flaky] recovery-runs, default 5): a flapping or latent scenario holds its flag until it earns a trailing clean run, so it cannot oscillate flaky↔healthy between adjacent runs.
    • An environment-outage guard (--outage-rate / [flaky] outage-rate, default 0.8): a run where over this share of suite scenarios failed is an environment incident, not evidence about any one scenario, and is excluded — so a single fixture or staging outage cannot mark the whole suite broken.
    • An input-fingerprint equivalence class — the default key. Each run writes an inputs.json sidecar carrying a hash of what it executes (feature sources + loaded macros/fragments + the resolved ${url:…}/${vars:…} scope), so a pack, feature, or proef.toml edit correctly ends the comparison window instead of silently mixing runs of different inputs. It is a proef-computed fact about proef’s own inputs, not harvested from the environment (ADR-0020 unchanged — git-commit grouping stays handed-over via --meta commit=… and proef flaky --by commit). broken≠flaky, transition-counting, and the quarantine lifecycle were already present and are unchanged.

Fixed

  • A run-record write that fails now reaches the exit code. The JSONL reporter deliberately swallows write results (a reporter cannot report its own channel dying), and events.jsonl was handed a bare File — so a disk filling mid-run truncated the record while the run still exited by its verdict, the exact class the v0.6–v0.8 series closed for the console. The record’s writer now latches its first failure (one stderr line, run continues) and the exit funnel turns it into a system error, the same shape as the stdout latch and the JUnit-write fold — unified in one pinned function, escalate_environment_failures. run.log’s mirror keeps its own contract (creation is warn-and-continue, so a mid-run failure warns once and leaves the verdict alone — previously it was silent).

  • The GitHub step summary can fail again. It was the only CI sink that couldn’t: a failed open or write vanished while JUnit and CTRF failures re-classify the exit — so the page a reviewer actually reads could be missing on a green exit. write_github_summary now returns the error and the caller folds it into the same reports_failed path as its siblings.

  • A custom --run-id no longer collapses the JUnit report identity onto the nil uuid. ADR-0021 made non-uuid run ids first-class, but the report uuid was parse_str(...).unwrap_or(nil) — every --run-id ci run emitted 00000000-…, colliding in any consumer keyed on it. A non-uuid id now derives a stable UUIDv5 from its bytes (a uuid id passes through verbatim).

  • The interrupt window and the interrupt’s own words. The handler is installed at the top of execute — before the front end, the run dir and the record exist — so no startup window takes the process default any more. Its installation failure is a printed warning (it was silently ignored, unlike --watch’s handler). The second-signal path no longer prints before exiting: the print took stderr’s lock, which a worker blocked on a full pipe can hold, wedging the escape hatch behind the very stall it exists to escape. And the teardown notice said “Ctrl-C again to skip” when a second interrupt actually hard-exits dropping every report — it now says what happens.

  • Asset staging no longer depends on the working directory. A feature’s file,…; assets were resolved by joining its portable name against the cwd — but a name’s anchor (the project root, or the caller’s own typed spelling) is not in the string, so a typed-absolute or config-written suite path run from any subdirectory failed staging with exit 2, blaming the author for a correct file (the feature-side twin of OPEN-FINDINGS H5). The resolved discovery path now travels beside the name (LoadedFeature::read_from) and staging resolves beside the file the parser actually read — the H5 prescription, applied to the feature side. Reproduced before the fix and re-verified after, from a subdirectory, against the reference corpus; a new integration test pins a project under a path with spaces and non-ASCII segments, which nothing in the suite had ever exercised.

  • --sarif line numbers survive a cd, and byte-match the parser. The SARIF writer re-read each source from disk by its portable name to count lines — from any subdirectory every read failed and startLine silently vanished, annotating nothing; the re-read could also disagree with the span by exactly the parser’s normalization. Lines now come from the diagnostic’s own carried source text — the same normalized bytes the span indexes. (On Windows, an absolute out-of-project uri also spells its separators as a URI requires.)

  • Staging’s two symlink edges. An existing symlink at a staging destination was written through — fs::copy follows links, so the bytes landed wherever it pointed, outside the root built to contain them; it is now replaced. A source symlink stays followed, deliberately: stock hurl follows it too, and refusing would break the dual-runner rule (the module doc now says so).

  • Asset names that are one file to the filesystem are refused. The duplicate-name guard keyed on the raw reference string, so Data.json and data.json — one file on macOS and Windows — silently last-writer-won, the very overwrite the per-scenario root was built to end. The check now runs on the canonical path the copy actually landed on, which is exact on every platform: a case-sensitive volume keeps both files legitimately, and nothing fires.

  • Artifact slugs cap at 120 bytes. The slug flattens the feature’s whole directory path into one filename component, and assets/<slug>/ repeats it as a directory — so path depth became filename length, and a deep tree or a long scenario name (multi-byte scripts at a quarter of the visible characters) sailed past NAME_MAX and failed the write. Over the cap, the tail is a hash of the whole uncapped slug, so two names differing only past the cut still name two artifacts; every slug the existing corpus has is under the cap and unchanged byte-for-byte.

  • The ADR-0007 budget family is closed over its inputs, and bounded as a product. [Options] max-time: was read by the budget calculator (as the entry’s timeout) while invisible to the lint — max-time: 100000h was lint-clean and produced a multi-year watchdog budget; it now carries the duration cap, and a test pins the rule the hole broke (every option the budget reads must be one the lint can see). retry-interval: — the one uncapped multiplicand — carries the cap too. And because individually capped values still compose into an unbounded product (retry: 10_000 × a 30 s timeout is ~83 lint-clean hours, saturating to Duration::MAX, whose deadline addition panicked as a phantom “scenario thread panicked” fault), the computed batch budget now clamps to an absolute four-hour ceiling and the dispatcher’s deadline arithmetic can no longer overflow. ADR-0007 carries the amendment.

  • [http] timeout-ms = 0 is refused. libcurl reads zero as no timeout, so the value opted a suite into exactly the unbounded hang the default exists to defend against — while reading like “immediately”. Exit 2, in whichever table it appears.

  • Every sink that renders run values now routes identities through the secret masker. The event stream masks scenario, file, tags and the skip reason under an explicit no-exemptions rule (“a field exempted because it can’t contain one is how that stops being true later”), and five sinks bypassed it for the same fields (0.18 survey): the GitHub annotation title=/file= lines (written to CI stdout), TAP’s skip reason and scenario name, CTRF’s name/suite/filePath/tags, JUnit’s suite/testcase identity and file attribute, and timings.json — the one sink that took no Redactions at all, in the file whose documented workflow is being archived and shared across a CI matrix. Structural mitigations (secrets lower to {{name}}; the engine pre-redacts details) made a live leak unlikely, but the boundary rule was unenforced; a per-sink leak test now pins each, and a whole-run sweep asserts a reflected secret reaches no file any sink writes.

  • proef lsp honours --env. The global flag was parsed and then silently dropped for lsp, so proef lsp --env staging analysed the default profile while runs used staging — the editor/runner drift R10-1 closed for --config. And the workspace-root re-resolution (for an editor launched outside the project) re-loaded the config to find the root but dropped the ${url:…}/${vars:…} scope it had computed, analysing the right tree against the wrong directory’s config; the scope now travels with the root it belongs to.

  • A CTRF report cannot gain a key the spec would reject. CTRF §4.4 makes consumers reject any key outside the defined set (unless under extra), and the spec moved five times in 2026 — so an additive field is a hard break. A test pins the exact allowed key sets.

  • De-flaked three tests (0.18 survey): the abandoned-scenario record-gate test waited on a 500 ms blind sleep that passed vacuously on a loaded runner — it now polls a drop latch set strictly after the worker’s final emit attempt, so it tests the dropped event on every machine; the bounded-runtime smoke test’s wall-clock assertion is widened and documented as the “generous upper bound” class TESTING-STRATEGY §7 sanctions (distinct from the #[ignore]d ratio guard); and a watch test’s fixed shared temp path (temp_dir()/proef-watch-alias-test + remove_dir_all) — the one cross-process race nextest cannot cover — moved to a unique tempdir.

  • proef doctor no longer prints fourteen literal spaces mid-sentence (a lost line continuation in the “hurl not on PATH” note).

Breaking

  • Library: proef_lsp::RootResolver now returns a ResolvedRoot (root + disk + config_vars) instead of a (PathBuf, Box<dyn SourceProvider>) tuple, so the re-resolved config scope reaches the server. proef-cli’s lsp::run takes the --env value. timings::render takes a &Redactions.

  • proef flaky’s new verdict is renamed insufficient-data (its --format json verdict key and human label), matching the 2026 vocabulary and the new sample-floor meaning — a MINOR break for a consumer keyed on the old spelling. The default per-scenario floor also rises from 2 to 10 runs, so a scenario with fewer than 10 runs now reads insufficient-data where it previously received a verdict (--min-samples 2 restores the old behaviour).

Internal

  • Post-0.18 /simplify cleanup — duplication the wave programme left behind, collapsed with no behaviour change (outputs byte-identical, no public API moved):

    • The input fingerprint’s FNV-1a loop and fake’s were the same loop and constants twice; now one fingerprint::fnv1a_with primitive, with fake::fnv1a a thin alias at the canonical offset basis.
    • proef test and proef --watch duplicated the whole two-stage-interrupt skeleton (the once-latch, the second-signal hard-exit, the stderr-lock rule); now one install_two_stage_interrupt taking the divergent first-signal action as a closure.
    • proef flaky recomputed each scenario’s verdict at ~8 sites — twice per comparison inside the sort; now classified once into a stored field, and render_table no longer threads the thresholds through to recompute it.
    • The reserved-tag typo warning derives its edit-distance threshold from each reserved word’s own length instead of hardcoding quarantine, so a future reserved tag earns fuzzy protection automatically, and it builds its diagnostic once rather than twice.
    • is_outage counts without a throwaway Vec and drops a dead precision-loss suppression; JUnit redacts a scenario’s file once, not twice; emit::cap_slug drops a redundant rebinding.
  • Redaction centralized at one exhaustive boundary (ADR-0005 hardening, no behaviour change on clean output). The CI sinks (JUnit, CTRF, TAP, timings, the GitHub summary) render from RunSummary, not the event stream, and each masked its identity and failure strings field by field — correct today, but a new field or sink could slip past unmasked. A new Redactions::apply_outcome mirrors the event stream’s exhaustive apply_event: it destructures ScenarioOutcome/StepOutcome with no .., so a new text field fails to compile until it is masked, and each sink now redacts an outcome once instead of the ~10 scattered apply calls it used to sprinkle (a scenario’s fault message, which reaches only this path, is masked with the rest). Additive to the library surface (pub fn Redactions::apply_outcome).

  • Redaction masking deduplicated to one primitive (a follow-up /simplify pass, no behaviour change). apply_outcome/apply_step_outcome were written as siblings of apply_event but re-spelled its Arc<str> masking idiom inline and dropped its clean-field optimization; a shared mask_arc/mask_step_ref now backs all four maskers, so a clean field reuses its Arc instead of reallocating (and apply_step_finished inherits the same win). timings reverts to masking just the two identity fields it renders, rather than cloning the whole outcome graph to read them.

  • Coverage is measurable on demand, and deliberately not a gate (the local half of P13, 0.18 survey). just cover / cover-html / cover-lcov run cargo-llvm-cov over the workspace (~90 % line coverage of the unit + integration suites today, xtask aside); TESTING-STRATEGY §3 records the policy any CI half must follow — a ratchet that fails only on a drop, never a fixed threshold. The gating CI adoptions the survey also listed (a cargo-mutants job, a coverage-service job, immutable releases) are a maintainer’s cadence/cost call and stay open in OPEN-FINDINGS.

  • The Homebrew tap only moves forward. The release workflow’s tap job is gated on nothing but “a tag was pushed” and rewrites Formula/proef.rb whole, so a tag pushed late or out of order would regenerate the formula for an older release and downgrade every brew upgrade. Not hypothetical: v0.15.0 was released and never tagged, so backfilling that tag would have walked the tap from 0.17.0 back to 0.15.0. The job now compares the tag against the version the tap carries and skips green when it is not newer — green, because publishing an old release’s binaries is legitimate and the correct outcome there is an untouched tap. It guards future tags only: a tag runs the workflow from its own commit, so one cut before this fix still runs the unguarded job, and RELEASING now says to disable the workflow around such a push. v0.15.0 was backfilled that way on 2026-09-09 — tag, and a not-latest Release carrying the changelog section without archives.

[0.17.0] - 2026-09-06 (the environment a suite runs in, and the guards that keep its claims true)

Added

  • [http] cookie-store = false runs the whole suite cookie-less — hurl 8.0’s --no-cookie-store, surfaced through the table built for exactly this class of setting. No Set-Cookie is retained and none is replayed, which is how a stateless API is proven stateless: the fixture-backed test is green only because its steps assert the 403 a missing session cookie earns.

    This is the one [http] key with no per-entry [Options] spelling at all (OptionKind has no cookie variant — verified against the enum), so run-wide is not a compromise but the only place it can be said. With the store off, the engine also skips both halves of the batch-split cookie round-trip: hurl reads a cookie_input_file only when enabling the engine, so injecting one would be silently ignored — and there is nothing to write. hurl’s own FIXME (a handle once given cookie storage cannot lose it) never reaches proef, because run_entries builds its client per call (TECH-SPEC §5) — a handle never transitions on → off.

    Breaking (library): HttpDefaults gains the cookie_store field, so a struct-literal construction needs the new line (..Default::default() sites are untouched, and an absent [http] cookie-store key changes nothing).

  • --ctrf <path> — the run’s verdicts as a CTRF report. CTRF (https://ctrf.io) is the emerging JSON successor to JUnit XML for CI dashboards, and it models in the schema what JUnit can only smuggle through extensions — which is exactly the data proef already tracks: a pass-after-retry carries flaky, retries, and retryAttempts listing the real failed attempts with their (redacted) messages; every test carries its tags and file path. One serializer off the same fold as JUnit, so the two files cannot disagree — most visibly for a quarantined failure, which both report as skipped with a message (ADR-0019), because a dashboard reading “failed” beside exit 0 would contradict itself. A User/System fault stays failed even under a quarantine tag: quarantine is for flaky tests, not broken input.

    The R12-3 contract applies from day one: a [run] setup abort still writes the file, carrying the setup scenario itself — a job gating on the report must never see no file at all. The schema’s required wall-clock start/stop are measured at the CLI edge like every other clock read (ADR-0015); the sans-IO core and the JSONL record are untouched — the record remains the only record (ADR-0008).

  • The HTML report answers “what is slowest”. After “what failed”, it is the question a test report is most often asked, and the page could not answer it: the timeline showed that workers were busy, never which scenarios to attack. Every number needed was already in the fold.

    A ranked section, slowest first, each row linking to its own block, with the heading reporting the share of run time the listed scenarios account for — “3 of 40 · 71% of run time” is a decision, where a column of durations is homework. Capped at eight: a ranking long enough to scroll has stopped answering the question.

    Cost is the sum of a scenario’s step durations, the same definition timings.json uses for shard weights — one notion of what a scenario costs across the whole tool. Not the wall-clock span, which includes time waiting for a worker: a property of how the run was scheduled, and not something the reader can go and fix.

    Absent when there is nothing to rank — fewer than two timed scenarios, or a record with no injected durations at all.

  • --shard-weights balances a shard matrix by measured duration. --shard assigns by a frozen hash, which guarantees that adding one scenario never re-buckets the others but cannot balance by time — and a CI matrix finishes when its slowest shard does, so a count-split routinely leaves runners idle. Every run that reaches its suite now writes a small timings.json into its run directory; CI archives that one file and each matrix job points --shard-weights at the same copy.

    The obvious design is silently wrong, and the module says so at length. proef already retains records carrying every step’s duration, so “weight by the newest local record” looks free. But matrix jobs run on different machines, each with its own (usually empty) runs-dir — every job would compute a different weight table, therefore a different assignment, and scenarios would run twice or not at all while the suite reported green. Nothing about that announces itself. One named file shared by every job is what makes the split a pure function of (selected scenarios, that file).

    Two rules place scenarios and they partition rather than compete: a scenario the file mentions goes through longest-processing-time-first placement, and one it does not mention falls back to the frozen hash. So a test added after the timings were captured still runs exactly once. That is pinned by a test that runs a whole three-way matrix — with a weights file covering only five of nine scenarios, so both rules are exercised at once — and asserts set equality both ways; mutating the placement by one bucket drops two scenarios and the test names them.

    The weight is the sum of a scenario’s step durations, not its wall-clock span. The span includes time spent waiting for a worker, which is a property of the run’s scheduling rather than of the scenario, and feeding it back would let one crowded run’s queueing distort the next split.

    What this gives up is exactly what hash mode was chosen for: a balanced split is not stable under insertion. That is what balancing means, which is why the flag is opt-in. A missing or malformed weights file is exit 2 — falling back silently would hand back the unbalanced split the flag was passed to avoid.

  • The editor tells proef’s two variable tiers apart. A pack’s hurl: | block is the centre of the authoring experience and, to every editor, a plain YAML scalar — inside which ${…} (resolved at lower time, by proef, before any request exists) and {{…}} (resolved at run time, by hurl) look identical. That distinction is ADR-0005’s whole model and the thing authors most often get wrong, and no generic grammar can see it: a YAML highlighter sees a string, and a hurl highlighter never runs because the block is not a file. proef is the only party that knows.

    The server now answers textDocument/semanticTokens/full, lighting ${…} as macro — a substitution performed before execution, which is what a macro is — and {{…}} as variable. Both are coloured differently by every mainstream theme, so it works without anyone configuring anything. The $${ escape stays dark, because telling an author proef will substitute text it will in fact leave alone is worse than no highlighting.

    The ${…} scan is proef_core::resolve::reference_spans, walking the same first_reference the resolver itself uses — a second implementation of the escape rule would drift, and the drift would show as an editor confidently colouring literal text. The {{…}} scan lives in proef-lsp rather than core, because that spelling is the engine’s and ADR-0002’s amendment is that engine syntax does not accumulate in the core.

    Collapsing the seven-arm request dispatch behind a local macro came with it: the chain crossed clippy’s line limit the moment an eighth feature landed, and the honest fix was to stop repeating an identical frame seven times rather than to suppress the lint that noticed.

  • The linear-validation claim is now a test, not a sentence. #138 made pack validation linear and recorded the result as a shape: “the curve changed shape — 4× per doubling before, ~2× after”. That number lived only in the changelog, where nothing could re-run it — so a future span locator scanning the whole pack file again would have restored the quadratic behaviour silently, a regression that costs seconds rather than correctness and which no gate measured.

    The guard asserts the ratio between 1000 and 2000 macros, because the claim is a ratio. It observes ~2.05× against a bound of 3.0; mutating locate::MacroIndex to re-index per lookup — the exact pre-#138 shape — measures 4.01×, matching the changelog’s own prediction of 4× and turning a 0.4-second test into a 73-second one. The failure message names the cause rather than reporting a number.

    A ratio rather than a benchmark, for a reason now written into TESTING-STRATEGY.md §7: load on a shared runner inflates both measurements together and cancels, where an absolute threshold has to be loosened until it means nothing. iai-callgrind would be the better CI gate — instruction counts ignore runner noise entirely — but it needs valgrind, so it would be a gate the maintainer cannot reproduce on macOS; criterion and divan sit in the same noise regime as this test while adding a dependency tree to a workspace that audits every edge. No new dependency was added.

  • Every diagnostic code is now named by a test, and a guard keeps it that way. DIAGNOSTICS.md calls codes “a contract: they never change meaning”. Twenty-three of seventy-five had nothing holding them to it — reachable in production, documented, exercised by nothing at all: not a seeded corpus directory, not a unit test, not even an assertion on their message text. They existed only at their definition site.

    The catalogue itself was found exactly honest — 75 codes defined, 75 documented, and its corpus column matched disk in both directions with zero drift. The gap was never documentation; it was that a documented promise had no enforcement.

    Nineteen new tests close it, each reaching its code through a real path rather than constructing the diagnostic directly. Two of them exercise guards that are unreachable in normal operation and were therefore the most valuable to test: lower::kind_unrouted fires only when the engine registry and pack validation disagree, so the test makes them disagree on purpose; and lower::expansion_too_deep sits behind pack validation’s identical depth limit, so the test bypasses validation with load_collecting — the only way to hand lowering a graph validation would have stopped, and therefore the only way to prove the second line of defence is still there.

    Two codes are exempted by name, with reasons recorded in the guard: source::unreadable and config::unreadable need a file the process may stat but not read, a permissions state CI runners do not reproduce because they run as root. The guard also checks its own exemption list, failing if an exempted code is deleted or renamed — an exemption that outlives its code silently excuses nothing.

    The guard joins the four in source_guards.rs and is mutation-verified: rewriting one test to match a code by suffix instead of naming it turns the guard red, which is the point — a test that matches the prose pins the wording, and only one that names the code pins the contract.

  • [http] now carries the settings that describe an environment: TLS, proxy and mTLS. The table exposed two of hurl’s runner options — timeout-ms and follow-location — while the embedded engine has supported the rest all along; TECH-SPEC.md:235 even listed insecure among what RunnerOptions carries. So a suite that had to run against staging’s self-signed certificate, or through a corporate proxy, or against an mTLS-protected API, could not say so anywhere: the only route was repeating an [Options] block inside every macro’s raw hurl, which defeats environment profiles exactly where they are most useful, since these settings are the difference between environments.

    Eight new keys — insecure, proxy, no-proxy, cacert, client-cert, client-key, max-redirs, user-agent — each merging field-wise through the existing [http] < [env.<name>.http] chain, so a staging profile turns verification off without production inheriting it. No new concept: only more of one that already worked.

    Three deliberate edges. insecure = true warns on every run, naming the profile that set it — a suite that goes green without verifying a certificate has not proved what a green suite normally proves, and since the run record carries no config by design, the warning is the entire audit trail. A client-key without a client-cert is exit 2 rather than a pass-through: libcurl accepts the pair and then presents nothing, so the failure would otherwise surface at the server as an authentication error naming nothing about the cause. And credentials are excluded on purpose — there is no user or netrc key, because a password belongs in the secret store where it is encrypted at rest and masked out of every sink.

    The three path-valued keys resolve against proef.toml, the one-path rule every other config path follows; core still reads no filesystem and receives them already resolved (ADR-0012). Each option is applied to hurl’s builder only when actually set, so a project with no [http] table runs byte-identically to one built before the keys existed — pinned by a test. Per-entry [Options] still override all of them except user-agent, for which hurl has no per-entry option at all; that exception is documented rather than papered over.

    Breaking (library): proef_core::engine::HttpDefaults gains eight fields and loses Copy — it now carries Strings. Default stays hand-written, and the reason is now stated in the type: a derive would make timeout_ms zero, which libcurl reads as no timeout at all, silently converting ADR-0007’s budget into an unbounded wait at every existing default() call site.

Changed

  • The toolchain policy is stated in the spec that rust-toolchain.toml cites. R18-2 corrected the policy to latest stable, adopted at its x.y.1 point release, and the correction reached RELEASING.md and CLAUDE.md while TECH-SPEC §15 — named by rust-toolchain.toml as its authority — still said “always latest stable”. R18-2’s own conclusion was that an unwritten policy contradicting the written one is a docs defect; fixing it in two files and leaving the source of truth contradicting itself reproduced the defect one level down. Now consistent across all four.

  • An artifact is named by its feature’s path, not its stem — two scenarios can no longer claim one file. Slugs were {stem}--{scenario}, dropping the directory, so features/x.feature and features/sub/x.feature each with a same name scenario both produced x--same-name: the second artifact silently overwrote the first while the CLI reported writing two. Silent loss of the hand-off ADR-0010 calls a contract — and the project already treats same-named scenarios across files as real, which is what --scenario-file exists for. The same slug drives the HTML report’s anchors and artifact links, reproduce: lines, and harness trial names, so all of them move together off the one helper.

    Names are now features-sub-x--same-name. Derived from the path rather than disambiguated on collision, deliberately: a counter or hash appended only when two names clash would make one scenario’s artifact name depend on whether some other file exists, so adding a feature would rename an unrelated artifact — the instability --shard’s frozen hash exists to avoid. The path fed in is the portable suite-relative name the record carries, never a path off the running machine.

    Breaking, and quietly so for library callers: emit::artifact_slug keeps its (&str, &str) -> String signature while its first argument changes meaning from stem to feature path, so the API gate cannot see it — passing a stem still compiles and now yields a different name. emit::emit’s second parameter changes the same way, and emit::feature_stem is removed (it had no remaining consumer). Artifact filenames and report anchors change for every suite; the snapshot corpus was regenerated under the new names and reviewed.

    The unification that made that a one-line change came first: the stem expression (file_stem, falling back to "feature") had existed four times across both crates — the emitter’s caller, the dispatcher’s spec naming, the HTML report’s anchors, the editor’s analysis — and the stem--scenario composition twice, with the report’s links to artifact files resolving only because both sides happened to derive the same name. Worklist item Q6 called the four sites a future-drift risk; collapsing them to one helper is what let the collision above be fixed in a single place instead of four. In the same pass, Q2 (the editor’s per-request walks) was found already closed by the #146 analysis cache, and its entry now says so with the evidence.

  • “What a scenario costs” is defined once, as ScenarioOutcome::cost. The sum of a scenario’s step durations was computed in three places on the same type — JUnit’s per-suite time, JUnit’s per-case time, and the new timings.json weights — plus a fourth over the record-fold shape in the HTML report. Four surfaces free to drift apart about a number they are supposed to agree on, and the argument for summing steps rather than taking a wall-clock span was written out twice.

    Now a method on the type that owns the steps, with the rationale stated there and referenced from the rest. The one behaviour change is a fidelity gain: the weights file used to truncate each step to whole milliseconds before summing and now truncates the sum, so its numbers agree with the times JUnit has always reported. Additive to the library surface.

  • A run whose setup aborted no longer leaves shard weights behind. timings.json was written from inside the CI-report block, which a setup abort also reaches — with the setup phase’s summary. The file that came out named setup scenarios, and a weights file naming them is worse than no file: those identities never appear in a suite run, so they absorb bucket load on behalf of scenarios that never run and skew the very split --shard-weights exists to balance, silently. The write moved to the one site where the summary is the suite’s, pinned by a test that reproduces the old file.

  • lower.rs stops threading the same three values through twelve functions. out, refs and sinks travelled as separate parameters everywhere, and five functions — expand_macro, expand_step, expand_ref_step, expand_payload_step, finish_step — carried 8 to 11 parameters each behind individual arity suppressions. Adding one piece of lowering state meant editing five signatures and five call sites, which is the shape of change that drops a parameter at one site.

    Two bundles, both of them types that were already implied by the code: Emit { out, refs, sinks } (the mutable outputs, always passed together and never independently), StepScope { step_ref, ctx, at } (what stays fixed for one authored step however deep expansion recurses), and a small Finished for the four values that describe a step being completed.

    What was not done matters as much. The obvious refactor — hoist the state into a self and make the five methods — would have broken the reason they are parameters at all: resolve_in and friends take them explicitly so they remain callable while other state is mutably borrowed, and a method on &mut self cannot be called while self is borrowed elsewhere. The threading discipline is load-bearing, so it stays; only the arity changes.

    Arity suppressions across the workspace: 13 → 6, with lower.rs at zero. No behaviour change, and the 241 core tests say so.

  • ADR-0002 now names the core’s hurl entry grammar, and a guard keeps it closed. “Adding an engine leaves proef-core diff-empty” was true of engine-types and never of engine-syntax: the core does text surgery on entries — splicing [Options] in, merging an expect: block’s asserts into the previous entry — so it has to find an entry boundary in text hurl will later parse. The worklist carried the gap for two rounds as “~290 lines of hurl grammar in core”, a figure that counted #[cfg(test)] fixtures.

    Measured: twelve literals across four files. Seven the core writes, four it recognises to find a boundary, and one it quotes — a hurl snippet inside a did-you-mean help string in bind.rs, which generates nothing and parses nothing but drifts like any other copy. The four boundary recognisers are already one shared pub(crate) set. proef’s own pack keys are shaped like option lines and are excluded by name rather than listed as sanctioned rows.

    The guard lexes whole files. The first version scanned line by line and so could not see a literal that spans lines — which is where a larger piece of engine syntax would naturally be written, and where the one entry above that nobody had counted was in fact sitting.

    The amendment sanctions that set and closes it. Deferred with a named trigger — a second engine being scheduled — is moving the written half behind the seam, where the reading half already lives: StepKindSpec::options exists precisely so an engine’s option spellings stay out of the core, and it covers recognising them only, so retry:, retry-interval:, delay: and variable: are still core literals. Until a second engine exists that migration relocates seven literals that exactly one implementation will ever supply, at the cost of a public-API break.

    crates/proef-cli/tests/source_guards.rs (renamed from stderr_hygiene.rs, which had not been only about stderr for two rules now) pins the set: a new token, or an existing one spreading to another core module, fails the test and names both remedies. A claim of this shape decays the moment it is only prose — this one already had, by an order of magnitude, in the direction that made it look worse than it is.

Fixed

  • A file,…; body in a ref: fragment resolves where its author put it. hurl resolves a file body against the directory of the file that wrote the reference — its --file-root default, and the same rule Karate, pytest and Jest use for fixtures. proef resolved every asset against the feature, and a fragment lives in another tree entirely ([run] fragments), so the same bytes passed under stock hurl and failed under proef, as exit 2, blaming the author for a path that was correct. Nothing worked around it: moving the file beside the feature breaks the standalone run ADR-0018 exists to guarantee, a reaching ../ path is refused by hurl’s own sandbox, and the advice that refusal prints — “check –file-root option” — names a flag proef does not expose.

    Each asset is now staged from beside the source that referenced it, feature or fragment, into that scenario’s own asset root, which is what the engine gets as its context dir. Staging rather than two roots because hurl offers one context dir per run of entries and no per-entry override, while a single batch may mix both body forms — measured, not assumed: an inline step and a ref: step in one macro lower to one batch. Copying fixtures into the build output is the standard answer to exactly this, and it adds no copy operation: the record already copied these files once per scenario, just into a shared directory instead of the right one. What it does change is the footprint — an asset N scenarios read is now N files in the run record rather than one, which is the same fact as the collision below, seen from the disk’s side rather than the reader’s.

  • Two scenarios’ assets no longer overwrite each other. Staging was flat and keyed by the asset’s bare name, so two features that each keep a data.json beside them staged to one file — last writer wins, with “0 warning(s)” — and the loser’s artifact replayed against the other’s bytes. artifact_slug already refuses that trade for the .hurl text, deriving from the feature’s whole path so two same-named scenarios cannot collide; the files it reads now get the same treatment. An artifact that reads a file says so in its replay line (--file-root assets/<slug>); one that does not is byte-identical to before. Two sources claiming one name inside a single scenario — the case a per-scenario root cannot separate — is refused rather than narrowed.

    A missing asset is also an error now instead of a silent skip. It had to become one: the staged root is what the engine reads, so a file that quietly failed to arrive is no longer an incomplete record but a request reading nothing.

    [Options] output: resolves through the same root, so the root is created for every scenario rather than by the staging loop — which never runs for a scenario that reads no file body. A response written that way now lands inside the run record, where a run’s outputs belong, instead of in the feature’s own directory.

    Breaking (library): emit::file_references is replaced by Artifact::assets, a Vec<AssetRef> carrying each reference with the source that wrote it — the provenance a whole-artifact text scan destroys, and the whole reason the bug was expressible. emit::asset_root names the staging directory for the three call sites that must agree on it, and pack::split_qualified is now the one reader of the file.hurl#name form Fragment::qualified writes — there were two, resolving a ref: and a use:, and staging assets was about to make a third in another crate. New diagnostic: proef::run::asset_unstageable.

  • A --run-id record is findable again (ADR-0021). --run-id pr-1234 writes a perfectly good record, and every command that resolves the latest run — explain, diff, flaky, report, --rerun — enumerated by the uuid shape, so that record was invisible to all of them. --rerun was the sharp edge: it silently continued some older run instead of the one just produced.

    One predicate had been answering two questions whose risks point in opposite directions — may I delete this? is unsafe when broad, is this a run I can show you? is unsafe when narrow — so the deletion-safety choice had silently become a visibility choice. They are now separate: a directory is a record because it holds an events.jsonl, while rotation still deletes only uuid-named directories, so a custom-id run is discoverable and still never reclaimed by [run] keep-runs. Ordering stopped riding on the name too — uuid-v7 sorted chronologically until a custom-id directory joined the set and sorted by its first letter — and now takes the timestamp a uuid-v7 name carries (48 bits of unix milliseconds, which is why the lexical sort worked), falling back to directory mtime for a name that carries none.

  • A run with a failed optional: step no longer prints exactly like a spotless one. ConsoleMode::Failed’s own doc comment states the requirement and the Warned arm implementing it was unreachable: a scenario’s aggregate status was only ever Failed | Skipped | Passed, so a real optional failure was invisible under --console failed, showed a . rather than the documented w under --console dotted, and left the HTML report’s warned count and its filter-bar warned button permanently empty — four consumers and three docs describing something that could not occur. Steps carried Warned; scenarios never did. The aggregate now promotes, which changes what a run says and never whether it gates: Warned counts as passing in the exit code, the totals and JUnit.

  • explain prints a step’s authored name:, like its five siblings. step_label’s own doc enumerates the six surfaces that must render it — console, HTML, JUnit, TAP, the job summary, explain — and explain was the one that never called it, so the post-mortem tool showed one sentence repeated where the live console had told the steps apart.

  • The HTML “Slowest” section no longer counts [run] setup/teardown into “% of run time”. Every other aggregate on the page excludes phases (ADR-0014), including the tag table directly above it, so the share meant something different in that one section. A slow phase stays visible in the timeline and in its own block.

  • .cargo/audit.toml no longer suppresses advisories deny.toml deliberately un-suppressed. It carried the quick-xml pair (RUSTSEC-2026-0194/0195) with a comment claiming it mirrored deny.toml — which had removed them, precisely because the reason had expired (quick-junit 0.7 moved to the patched quick-xml 0.41, which the lockfile is on). So the nightly cargo audit job was suppressing for no reason, and would also have silenced any new advisory filed against that line.

  • A merged report covers the whole suite again, and its headline agrees with its page. Two independent failures in --rerun composition, against docs/CI.md’s promise that “one report stands for the composed result”. The overlay followed only the immediate rerun_of, so the ordinary fix → rerun → fix → rerun loop — the workflow the feature exists for — silently dropped everything from before the last link, with no banner saying so; the page just got smaller. It now walks the chain, newest verdict winning, with a cycle guard because rerun_of is a string read out of a record and records travel. And the headline took its numbers from the tail totals, which belong to the re-run, so one page read 2 passed · 0 failed above a tag table summing to eight and a sibling JUnit saying tests="8". The composed stream now declines those totals rather than inventing new ones, so the headline counts the scenarios actually rendered.

  • Run metadata reaches the two ADR-0020 §5 consumers that never received it. The GitHub job summary — named in the ADR, and the page a CI reader actually opens from the job — carried none, so the commit under test was in the record and the HTML report but not there. And diff --format json carried env but not metadata while diff’s human output printed metadata differences, leaving the machine surface a CI gate reads missing exactly the context the ADR was written to provide. Still handed over, never harvested.

  • The ADR-0002 grammar guard can now see the shapes the ADR names. The amendment claims the core’s hurl vocabulary is closed and pinned; the guard classified four shapes, and method lines — one of the four boundary recognisers the amendment’s own Measurement section names — was not among them. Teaching it surfaced one unenumerated token immediately: GET ${url:base}/PATH, sitting in bind.rs in the same literal as the already-pinned HTTP 200. The ADR’s table and the pinned set both now carry it, and the vocabulary is thirteen literals rather than twelve. The scan also stopped truncating at a file’s first #[cfg(test)] mod and now excises every test module: production code placed after one was silently unscanned, and two core files already carry a second test module.

  • --rerun on a truncated record no longer reports success over a suite that never ran. Record::scenarios is built from scenario_finished events alone, so a run killed mid-flight — SIGKILL, OOM, a full disk, a container eviction — leaves its unreached scenarios absent rather than recorded. The candidate list built from such a record named nothing, the “no failures” branch fired, and --rerun exited 0 having executed no scenario at all. explain saw the truncation the whole time; --rerun did not, and CI is exactly where truncation happens. The same class as the cancelled-run bug fixed in 0.14.0, which this code’s own comment describes.

    A truncated base inverts the question: not “what did the record say to re-run” but “what can the record prove finished” — everything else in the selected front runs, announced with a warning naming the truncation. That distinction now lives in a RerunFilter predicate rather than a list, because only the record reader knows which of the two questions applies.

  • diff no longer reports a scenario skipped in both runs as “now skipped … (was passing)”. Both halves were false — it did not become skipped, and it was not passing — and it fired for every @skip scenario on every diff, including two runs of an unchanged suite, handing --format json consumers the same wrong pair. The bucket exists for transitions (ADR-0019 §7); the guard makes that true of the code and not only of its name.

  • A tab in a bound value is refused where every other control character already was. The lower-time guard exempted \t, which hurl’s variable: grammar rejects like any other control character, so exactly one character kept taking the late path the guard exists to close — dying as emit::invalid_artifact against generated text the author never wrote, rather than as a refusal naming their own bind:.

  • --shard-weights no longer piles every zero-cost scenario into shard 0. Costs are whole milliseconds, so anything sub-millisecond stores as 0 — routine for a fast suite — and adding 0 never moved a shard’s load, so shard 0 stayed the minimum forever. An all-zero weights file put the entire suite in one shard and left the others selecting nothing: the flag doing the exact opposite of its purpose, silently, with the partition still exact so nothing complained. Assignments are now a tie-break alongside load, which also gives the right answer when weights genuinely cannot separate scenarios: equal cost, equal share.

  • A disk filling mid-run now reaches the exit code. A stdout that was already broken at start has failed loudly since the correctness series — but the human report’s own writes go through the console reporter, which swallows write errors (a reporter cannot report its own channel dying), so a disk filling during the run truncated the report while the run still exited by its verdict. The Tee under the reporter is the last place the failure is visible; it now latches the same stdout-failure flag outln! uses, and the exit funnel turns lost output into exit 3. Same closed-pipe exemption as ever — proef … | head is the reader ending the pipeline, not a failure — and a stderr console (machine mode) does not claim stdout failed. Pinned by a three-case test, mutation-checked.

  • The complexity guard added moments earlier was itself flaky, and now runs alone. It shipped in the ordinary suite on the reasoning that a ratio cancels out runner load. Measurement disagreed on its second full-suite run: 2.05× isolated, 3.09× under nextest’s full parallelism, against a bound of 3.0. The larger input has the larger working set, so memory-bandwidth contention penalises it more than the smaller one — the ratio drifts rather than cancelling, and interleaving the samples cannot fix a systematic effect.

    nextest’s test-groups bound concurrency within a group and do not isolate one from the rest of the suite, so the only mechanism that actually delivers isolation is #[ignore] plus a dedicated invocation: a CI step of its own and just perf. The samples are interleaved as well, which removes the one skew that ordering alone creates.

    TESTING-STRATEGY.md §7 previously asserted the opposite in as many words — that a ratio “survives a shared runner” — and is corrected with the numbers. The claim was reasoning, not measurement, which is the failure this whole section of the changelog exists to record.

Internal

  • A fixture that spells the record by hand can no longer drift off the schema. explain’s truncated-record test wrote its stream as three JSON string literals, and all three had drifted: a scenarios count on the head, a schema on the body events, a line on the close. Event carries none of them. Nothing failed and nothing could — the reader has no deny_unknown_fields, so a stale key parses cleanly and is dropped, and a fixture built to assert “a record holding one passed scenario” was three-quarters describing a format proef has never written. It is typed now, through the helpers its two neighbours already use.

    The class is closed by a sixth source_guards.rs rule: every string literal in the workspace that parses as a JSON object tagged event must deserialize as an Event, and every key in it must matter — a key is phantom when deleting it yields the same Event. Inertness rather than an inventory, so it stays correct through renames, #[serde(default)] and skip_serializing_if, none of which a key-set comparison survives. Substring assertions against records proef actually emitted ("event":"run_finished","passed":1) are skipped by construction — they are not objects, and they check the opposite direction.

  • Each doc check now lives in the half of the gate that its own rule names. tests/docs.rs holds the checks that need a built binary (they ask clap, rather than parsing help text into a model that could drift); xtask docs-check holds the ones that read files. The changelog-heading check added moments earlier read one file and parsed headings, so it sat in the wrong half — and the cost was concrete rather than tidy: it never ran in the fast doc-only CI step, only under a full nextest that had to build a binary it did not use.

  • A PR that changes source now has to record itself. RELEASING.md has always said that every landed change adds an [Unreleased] line in the commit series that lands it, and nothing checked it — this very entry is the one that was missed. Measured before being written: across the previous 21 source-touching merges the rule would have fired exactly once, on exactly the commit that broke it, so the check earns its place by count rather than by argument.

[0.16.0] - 2026-08-31 (the surfaces tell the truth: an eight-wave improvement programme, and the round that found what it missed)

Supersedes 0.15.0, which was cut (release: v0.15.0, 2026-08-25) but never tagged or published — its changes are all here, and crates.io goes 0.14.0 → 0.16.0 with nothing skipped.

Added

  • explain, diff and doctor speak --format json. They were the three commands with no machine output, and the three a consumer reaches for after a run. A run directory is artifacts/ + events.jsonl + run.log and carries no structured summary, so anything analysing a run it did not launch — a CI job reading another job’s artifact, a script, an agent — had to fold events.jsonl itself. That is the fold proef’s own two internal copies disagreed on three ways before report::suite_totals unified them; handing the canonical answer over is cheaper than inviting everyone to re-derive the one proef got wrong.

    Each object mirrors its prose field for field rather than modelling a richer view — the prose is the contract a reader already knows, and a machine surface that says something different is a second answer to one question. diff’s flaky/slower stay the rendered sentences for the same reason. The flag is the existing single-variant json enum the listing commands already use, renamed from ListFormat to JsonFormat now that it serves non-listing commands too. Machine mode owns stdout: notes whose content the object already carries are suppressed rather than repeated on stderr.

    doctor needed a real change to get there — it printed each check as it ran, so the verdict was the only thing a caller could see. Checks are collected before rendering now, which makes the JSON a second rendering rather than a second walk: the failure mode where one surface gains a check the other never learns about.

  • --console failed — the full BDD tree, but only for scenarios that failed or warned. A clean run prints the run line and the summary; a dirty one prints exactly what full would. The gap it fills is the CI one: full is a wall of green on a large suite, dotted drops the detail you need when something breaks, and quiet drops everything.

    Warned scenarios are shown, which the name does not say and the code explains: a warned scenario is one whose optional: step failed, RunSummary::passed counts it with the passes, and the summary line has no warned column — so a mode that showed only Failed would let a run in which something did fail print exactly what a spotless one prints. A fourth value on the existing flag rather than a new one.

  • proef flaky --by <key> splits flakiness by run context. --by env, or any [meta]/--meta key (--by runner), folds the history per context instead of pooling it. A scenario that flaps in one environment and is solid in another is not flaky but context-dependent — the fix is in the environment, not the test — and a merged history cannot reach that conclusion, because pooled failures and passes look exactly like one flapping test. The command names the scenarios whose verdict changes with where they ran, which is the finding the flag exists for. A run that never set the key becomes its own (unset) bucket rather than being folded in with runs that did; the context also rides in --format json. Reads the env/metadata provenance the record has carried since ADR-0020 — no new recorded field.

  • proef schema config publishes the proef.toml JSON Schema. TOML language servers (Taplo, tombi) validate against JSON Schema, so one file buys completion, hover documentation and typo detection in the config — before a run rather than after one. Generated from the same Rust model that parses the file, so it describes keys as they are written (runs-dir, not runs_dir) and inherits deny_unknown_fields, making an editor refuse exactly what proef refuses. proef schema keeps printing the pack schema, so one command answers “what may I write in this file?” for both authored formats rather than two verbs answering it once each.

  • An assertion that fails on values looking identical now says why. When the actual and expected values differ solely in whitespace, the failure carries a note repeating both with every whitespace character drawn — · for a space, \t/\r/\n for the usual escapes, \u{a0} for the exotic ones. hurl’s own message was already correct; the defect was simply invisible, so a trailing space, a CRLF fixture leaking \r, or a non-breaking space pasted out of a browser read as “the tool is wrong”. Taken from hurl’s structured actual/expected rather than parsed back out of its prose, and emitted per error so it sits beside the values it explains; silent whenever the difference is already visible.

  • proef flaky audits quarantine, which nothing else could. A @quarantine scenario’s failures gate nothing by design, so no exit code, no summary and no CI job reports them — which makes the tag’s own failure mode invisible: a quarantined scenario failing every run has been switched off and left in the suite. It now reads DISABLED rather than sharing the broken verdict with untagged always-failures, which wrongly implies someone is watching. The opposite case gets its own verdict too: green throughout the window is recovered, a tag that outlived its problem and is now suppressing the next real regression. Both print what to do, and --format json carries quarantined plus the verdict key so a scheduled job can gate on either.

    This needed the record reader to stop dropping data it was already given: scenario_finished has carried tags since 0.15.0, but ScenarioRun never parsed them, leaving every record consumer tag-blind.

  • Document symbols and hover. A feature outlines to its scenarios (with their tags), a pack to its macros (with the pattern each matches) — the vocabulary chosen by what discovery found in the file, never by its extension. Hover answers the question go-to-definition charges a round trip for: what a step binds, what a use: targets, what a ref: resolves to and which of its variables still need a bind:. Every fact is read from the same analysis the diagnostics come from, so a hover cannot contradict the squiggle on its own line. SuiteAnalysis gains a scenarios index, taken from the parse rather than from binding — an outline that hid exactly the scenarios you are debugging would be worse than no outline.

  • A panic no longer ends the editor session silently. Only the recompute was guarded, so a panic inside completion, definition or references escaped the message loop and killed the server — leaving an editor that shows nothing, which reads as “proef has no opinion here” rather than as a failure. Both entry points (a request, the debounced recompute) now wrap everything they do, the request is answered with InternalError rather than dropped, and the user is told once per suite state through window/showMessage — the channel an editor surfaces, unlike the stderr line that was the only report before. The next edit clears the report, because whether the new state also fails is news.

  • The editor can apply a “did you mean”, not just print it. Every misspelled-name diagnostic that already suggested a nearest spelling now carries the structured half of that suggestion — a span and a replacement — and proef lsp serves it as a quickfix code action: use: and ref: targets, with: and bind: keys, step kinds, Examples placeholders, and data-table columns. The suggestion is computed once and rendered twice (prose for a reader, an edit for an editor), so the message and the fix can never disagree.

    A fix is attached only when the edit is certain: the suggested name is near enough, and the misspelling occurs exactly once, as a whole token, in the diagnostic’s own file. Each of those failing means no fix rather than an approximate one — notably, a lowering error anchors on the feature step that invoked a macro while the typo lives in the pack, so it finds nothing to replace and offers nothing rather than editing the healthy file. The action is reachable from either the diagnostic or the token, because the two are regularly lines apart: a use: error carets the macro’s name key.

  • README answers the comparison a prospect actually runs: a “When something else fits better” section maps raw hurl (the exit stays open in both directions), Karate (choose it for embedded JS and whole-body fuzzy matching — the two mechanisms proef deliberately refuses; choose proef for one binary, deterministic reproduction, and files that run with no framework at all), and Postman/Bruno-class clients. The quick-start also points at proef init as the start that demonstrates the ref: body form — tests/features/ is deliberately fragment-free (the reference corpus is config-independent by design, and [run] fragments is a config key; the runnable ref: demo lives in the scaffold, pinned green against the fixture).

  • The docs site can get a visitor to a binary, and CI to a green workflow. New Installing page — install lived only in the repo README, outside the published site’s source, so the site’s first step sent visitors back to GitHub — and a new CI page with the paste-ready workflow the docs never had (zero runs-on blocks existed anywhere): install, secrets via PROEF_SECRET_*, --junit auto, a --shard matrix, --meta provenance, the diff --fail-on-regression baseline gate, --rerun continuation, and flaky over retained records. Nav reordered visitor-first (Installing → Getting started → Writing scenarios).

  • AUTHORING gains the three recipes every real suite needs: login-then-use-the-token (the docs’ most-asked absent question — zero “login” hits existed), waiting for an eventually-consistent result (finite retry: as the polling primitive, and why it must be finite), and test-data seeding/cleanup across its three scopes (Background:, [run] setup/teardown, saveAs: global).

  • Every release archive ships a .sha256 sidecar (basename inside, so sha256sum -c works from a download directory). Attestation covers the provenance story for gh users; the sidecar covers everyone who installs with curl — the half that was missing against the ripgrep/uv/starship baseline.

  • A broken proef.toml is a located diagnostic, not a bare sentence. The file is edited as often as any pack, and it was the one authored input whose errors carried no code, no source excerpt and no caret — while pack::yaml had all three for the structurally identical failure. New codes proef::config::toml (with toml’s own error span under the caret) and proef::config::unreadable, in the catalogue (73 → 75) and pinned by an integration test; proef lsp’s boot warning and doctor’s config row carry the same message.

  • Five help-less refusals gained their missing action. feature::parse (the shape of a feature file, and the most common way one stops parsing), bind::ambiguous_step (make one pattern more specific or retire the duplicate), bind::table_conflict (one source per param), pack::use_cycle (pull shared steps into a third macro), and the raw retry: -1 message now says why infinite retries are refused (hurl cannot be interrupted mid-call) and what to write instead — it used to cite “ADR-0007”, an internal document id with no in-band route to it.

  • A miss below the did-you-mean threshold names the valid set instead of going silent. All eleven suggestion sites ended closest(…).unwrap_or_default() — when nothing was near, the tail vanished, and unknown_step_kind said “not claimed by any registered engine” about a registry with exactly one member it never named. One matcher::suggest_or_enumerate now serves every site: the nearest spelling when one is near, else the set verbatim (small), else a count with the command that lists it ((9 known — proef macros lists them)). unknown_fake, unknown_variable and missing_config_var carry the same rendered tail through their typed errors.

  • unknown_placeholder fires once per authored defect, not once per Examples row — a 500-row outline with one typo’d <column> pushed 500 byte-identical diagnostics at one span (the console collapsed them; SARIF, one-result-per-site by design, did not). It also now names the header’s columns.

  • Every rendered error links the diagnostics catalogue. The stable codes were greppable and led nowhere — the catalogue was linked from every doc and reachable from no error. Rendered implements Diagnostic::url() and the LSP sets code_description, so editors show a clickable link on the code; on a terminal miette renders an OSC-8 hyperlink, and into a pipe or snapshot the URL prints as plain text beside the code (links ride the same TTY/NO_COLOR gate as color — an escape sequence a non-terminal sink must never see).

  • Pack diagnostics point at the defect, not the macro’s name. Every pattern-family and defaults: error anchored on the macro-name span — thirteen of the nineteen seeded pack snapshots underlined login: while the broken {rol} sat on a line outside the excerpt (one excerpted the previous macro). The match:-line span was computed since the pass was written and never reached a diagnostic; it does now, with the name span as fallback. locate::macro_span also stopped matching pack-root bind: entries (a macro sharing a name with a bind key anchored every diagnostic on the config line).

  • Parser errors speak hurl’s and gherkin’s prose, not Rust’s. A pack author was shown ResponseSectionName { name: "Wrong" } and Method { name: "" } — {:?} of internal enums from crates they never heard of. All three engine sites now render through hurl’s own DisplaySourceError (“the section is not valid. Valid values are Captures or Asserts”), and gherkin’s expectation-set tail is sort-normalized: it renders from a HashSet, so the same broken file printed two different messages across processes (observed live) — breaking snapshot determinism and the duplicate-collapse alike. Pinned.

  • A resolve::* error names the pack it lives in. The span is the feature step (the invocation), but ${nope} is written in a pack YAML the message never named — the reader was sent to a healthy .feature line while the sick file stayed anonymous. Every resolve error now carries (pack <file>).

  • The HTML report is triageable, linkable, and filterable. Every scenario block carries an id="s-<slug>" anchor (the same stem--name slug as its artifact, so the two cannot disagree) — a failure is now a URL a colleague can be handed. A “failed:” jump rail under the summary links straight to each failing block (blocks keep completion order — the rail is how a reader skips the green between failures), and a status-filter bar (all/failed/skipped/warned) toggles block visibility through a ~15-line inline script: progressive enhancement over classes the blocks already carry, no framework, still one self-contained file. Snapshot reviewed deliberately.

  • --watch reads like an inner loop. A visual rule with a rerun counter separates iterations (twenty edits used to stack twenty trees with nothing marking where the current one begins), and the post-run line says the verdict in words (“failures — details above”) instead of an exit number to decode.

  • Shell completions and a man page, generated by the binary itself: hidden proef completions <shell> (bash/zsh/fish/powershell/elvish) and proef man subcommands, and every release archive now carries completions/ plus proef.1 — generated during packaging by the exact artifact they ship beside, so they can never drift from it.

  • --env is global, like --config: proef --env staging test and proef test --env staging both work — five commands read the profile, and the position-sensitive spelling was a lesson nobody needed.

  • doctor examines the project, not just the engine: suite resolution (feature-file count, or the failure), hurl on PATH (a warning when absent — the engine is embedded, but ADR-0018’s stock-replay promise and the emitted # replay: hints need the binary), and runs-dir writability (probed with cleanup — the first-run create_dir_all failure was invisible to the one command whose job is diagnosis).

  • A typo’d --tags/--scenario names the nearest real spelling. The refusal held every scenario name and tag at the moment it printed “check –tags/–scenario” and used none of them; it now suggests the closest name and tag (glob atoms excepted — a glob selecting nothing is a fact, not a typo) and points at proef flows, the treatment [run] exclusive-tags always had.

  • proef fragments says why a listing is empty when no [run] fragments root is configured — previously indistinguishable from a configured-but-empty corpus, though the reader’s next move differs.

  • The console speaks in color, and every run ends on its identity. The status vocabulary (✓/✗/∅/⚠, the dotted glyphs, the summary’s verdict half) is ANSI-colored on a terminal — NO_COLOR, a dumb TERM, or a non-terminal stream turns it off, and the run.log mirror strips the paint either way (content verbatim, paint never). Color is paint on identical bytes: the record, the exit code and every text assertion see the same output. Each run’s final stderr line is now run <id> · <seconds>s — the run id is the reproduction key --shard, --shuffle and ${fake:…} all hang off, and it previously printed only at the top of the scrollback; a red run’s trailer adds the proef explain pointer. Wall-clock stays console-only, never entering the record.

Changed

  • Secret redaction no longer runs inside the reporter mutex. The sink masked each event while holding the lock that fans it out to the reporters, so every scenario thread queued behind work none of them share — and masking is the expensive half, a scan per text field per needle with roughly nine needles derived per secret. It reads the event and the needle set and writes neither, so it never needed the lock; the critical section now covers only the fan-out it exists for.

    Order is unaffected and the tests say why: a scenario is one thread, so its own events still reach the lock in the order it emitted them, and order across scenarios was never guaranteed. A new test emits from eight threads at once and asserts nothing is lost or doubled, everything arrives redacted, and each emitter’s own events keep their order. No timing assertion — the flake rule forbids one, and the change is justified structurally rather than by a stopwatch.

  • Pack validation is linear in the macro count, not quadratic. Every span locator scanned the whole pack file to find its macro’s block, so validating N macros scanned the file N times. A single indexing pass (locate::MacroIndex) records each macro’s name span and block region, and the locators became lookups into it. Measured on a release build over generated packs: 3200 macros went from 1.96 s to 0.03 s (~65×), and the curve changed shape — 4× per doubling before, ~2× after — so 6400 macros now cost 0.06 s where the old scaling predicts ~8 s.

    It also fixes an inconsistency the split readers hid: macro_span accepted a quoted "macro name": header while the region scan behind every other locator accepted only the bare form, so a quoted macro got a caret on its name and silently no span for its match:, use:, ref: or payload lines. One reader now gives one answer.

  • The editor stops re-analysing the suite on every keystroke. Completion, go-to-definition and find-references each ran the whole pipeline from scratch — read every pack and feature off the provider, parse, bind, lower — and threw the result away; between two keystrokes none of those inputs have changed, so the second run could only reproduce the first one’s answer. The server now holds the analysis and drops it exactly where an edit lands (the same notification path that already marks the suite dirty), so one recompute serves the debounced diagnostics publish and every request until the next edit. Measured on the two-file test suite: 10 provider reads per request before, none between edits after — pinned by a read-counting provider rather than by timing, per the flake rule.

  • The LSP’s type layer moved to the maintained generator: lsp-types 0.97 (unmaintained since; the crate that shipped its own fluent-uri Uri newtype) is replaced by gen-lsp-types 0.11 under the same lsp_types:: name — rust-analyzer’s own aliasing pattern, so every use path is unchanged. Its url feature aliases Uri to url::Url, which the embedded hurl engine already pulls in, so the swap adds no new crate and drops three (lsp-types, fluent-uri, serde_repr). Url::from_file_path/to_file_path are the native-path bridge documents.rs had to hand-roll under 0.97 — drive letters, segment joining, percent-encoding, ~90 lines — so the bridge is now a wrapper that only pins the source-name identity rule. Behaviour visible to an editor is unchanged; the one difference is what counts as a malformed URI (url percent-encodes a raw space where fluent-uri rejected it), and request dispatch now compares a method enum rather than strings, so an unknown method lands in Custom instead of matching nothing.

    Breaking (library): proef_core::report::percent_encode is private. It was public solely so proef-lsp could encode URI path segments against the identical unreserved set; that hand-rolled encoder is gone, and redaction needles — its only remaining caller — live in the same module.

Fixed

  • The record-size ceiling reached two of its four readers. 0.13.0 bounded the run-record read at 256 MiB because records travel — diff reads a downloaded baseline, flaky reads every retained run — and the read, the line split and the parsed Vec<Event> are resident at once, so a corrupt or hostile file was an OOM rather than an error. The bound lives in record::read_events, and explain and report each opened events.jsonl with a bare read_to_string instead, so neither had it. report even used the guarded reader for the base record two dozen lines below the raw read of the primary one.

    Both now go through read_events, which returns the parsed events — exactly the read-once/parse-once its own comment asked for. A source-scanning test makes the next reader go through the same door, the shape this project already uses for the raw-print and malformed-plural rules: a guard added in one place and left for the next call site to rediscover is how it went missing the first time.

  • cargo deny failed on a yanked transitive crate. rand 0.10.2 resolved chacha20 0.10.1, which was yanked from crates.io; the lock now takes 0.10.2. Not the secret store’s copy — chacha20poly1305 pins 0.9.1, which is unaffected — so nothing about encryption changed. Found by the gate, which is what it is for.

  • proef report -o wrote the author’s home directory into the file built to be shared. With the report inside the run dir the artifact links are a bare artifacts/…; with -o pointing anywhere else they were made absolute, which resolves only on the machine that produced them — and -o exists to put the report somewhere it will be published, which is exactly where that path is dead. 0.13.0 scrubbed machine identity out of the run record (R12-1); this put it back, twelve times over, in the HTML uploaded beside it. The href is now relative to the report, which resolves everywhere the absolute one did plus wherever report and artifacts travel together, and in the CI shape (-o public/report.html) names nothing outside the workspace. The href is built from path components joined with /, not from Path::display — Windows renders \, which is not a separator in a URL, so a Windows-generated report’s links would have been dead either way (the absolute path it replaces had the same flaw). A report written somewhere sharing no ancestor with the run dir still names the directories between them — that is what a correct relative path from there is, and it is no worse than what it replaces.

  • The report’s --skip colour failed WCAG AA, and every status pill failed it in dark mode. --skip was the one palette token the dark block did not redefine: a grey chosen against #0d1117 (5.48:1 there) left carrying white text on white at 3.45:1, against a 4.5:1 threshold — on the status a reader scans for after an interrupted run. It is now #59636e (6.11:1).

    Writing the guard rather than the fix found a second defect nobody had measured: .pill painted color:#fff on the status colour, and the dark palette’s colours are tuned as text on a dark ground, so all four dark pills sat between 2.52:1 and 3.45:1. The pill foreground is now a palette token — white on light, the page ground on dark — putting all four between 5.48:1 and 7.5:1. A test asserts the ratio rather than the hex, so a future palette change is free to move a colour and not free to move it below AA, and a second test pins that both palettes define the same token set (the absence that caused this).

  • The HTML report had one heading and no outline. The timeline carried an <h2>; the tag table and the scenario list — the body of the page — had none, so there was nothing to navigate by and no anchor to link a section with. Both gained one, sharing the class the timeline already used (renamed from .timeline-h to .section-h, since it now serves three). Pinned structurally, so a section added without a heading fails the test.

  • A step’s name: label reached the artifact and nothing else. A macro with more than one step turns one feature sentence into several engine steps, and they share a StepRef exactly — same file, same line, same text. The emitter has always written the authored name: into the artifact’s entry comment, which is why the .hurl could tell them apart; StepRef never carried it, so the console, the HTML report, JUnit, TAP, the job summary and explain all printed the same sentence once per step, with nothing but the status glyph to distinguish a warning from the failure beside it. The reference corpus demonstrated it: three step_finished events for the cookie session is exercised, byte-identical in the pinned snapshot, are now obtain the session cookie, optional probe (forces a split) and cookie survives the split.

    StepOutcome and step_finished now carry label, exactly as they carry fragment — the two answer neighbouring questions (which file did this request come from / which step of the sentence is this) and travel the same channels. One proef_core::report::step_label renders it for every sink, so the six cannot drift. Additive on the wire: absent when a step has no name:, so every pre-existing record still parses and re-renders unchanged, and the event schema stays 1.

    This retires two claims that were not true when written: AUTHORING.md’s “they anchor artifacts, events, and failure output” and LoweredStep::label’s own “(events/console)”. Same class as reproduce_hint in the R18 wave — computed all along, printed all along, dropped by the record.

  • A fragment’s text ran on into the comments introducing the entry below it. hurl attaches the blank and comment lines above a request to that request, which is exactly what makes the # @proef binding reliable — but it also means an entry has two different starts: where its lines begin and where its request begins. The scanner used one value for both, ending each fragment at the next entry’s request line, so every comment a corpus author wrote to introduce the next request was copied into the previous fragment and from there into the emitted .hurl. An artifact could carry # Destructive. Operators only. while containing no destructive request at all, and trim_end could not help — a comment is not whitespace. The same applied at the end of a file, where a trailing note became part of the last fragment. A fragment now runs from its annotation to the end of its own request and response; the gap between two entries documents the one below it and belongs to neither. Nothing executed differently, because hurl permits only comments and blanks between entries — which is why it survived: the only damage was to what the durable record says a request is.

    The property covering this asserted one request line per fragment, which is blind to comments; it now also asserts that no fragment holds any of the generator’s inter-entry filler.

  • explain and the HTML report disagreed about a truncated run’s totals. A record with no tail run_finished — a run killed mid-flight — is reconstructed by counting, and each surface carried its own version of that fallback. On the same bytes they differed three ways: the report dropped Warned scenarios from every column, counted [run] setup/teardown scenarios into a headline its own page labels “excluded from totals above”, and read a pre-0.6.0 record’s per-phase totals as the suite verdict where explain correctly declined to. One proef_core::report::suite_totals now holds the rule — prefer the tail event unless it cannot be trusted, else count suite scenarios with Warned riding along with Passed, exactly as the live path reports — and both surfaces call it.

    Also un-splices three doc comments in html.rs that an earlier change had merged into one, leaving render_tag_table and render_timeline undocumented and render_provenance_and_summary carrying all three.

  • A parse error pointing at a non-ASCII character produced a span that split the codepoint. gherkin reports a char-counted column, so the span’s start was correct; its end added one byte to that, landing inside a multi-byte character whenever the error pointed at one — a span that is not a valid slice of its own source. Nothing crashed, which is how it survived: miette tolerated it and drew the caret slightly to the left, and the LSP’s converter snaps to a boundary defensively, so every consumer defended itself instead of the producer being right. Found by the new fuzz_feature_parse target within a minute of first running.

  • A long --tags expression aborted the process instead of failing. and/or chains parse iteratively, and the module said so as though that settled it — but an iterative parse still builds a left-leaning tree as deep as the chain is long, and both eval and the derived Drop walk that tree recursively. A --tags expression of roughly twenty thousand and-joined atoms therefore overflowed the stack and died on SIGABRT: a signal, not one of the four exit codes ADR-0009 promises, and well within what a command line accepts. Expressions are now capped at 512 tokens, which bounds the tree and so bounds both walks, and past the cap you get a message naming the limit. (The test that was meant to cover this built 5 000 atoms and asserted success — one order of magnitude below the cliff.)

  • EDITORS.md no longer under-promises on built-in macros. It said the expect* family has “no jump target and no hover”; the first half is true and structural (their pack is compiled into the binary, so there is no file to open), the second is not — a built-in is in the analysis like any other macro, so hover answers with its pattern and params and names the pack as builtin:…, which is exactly why the jump is unavailable. Pinned by a test, since the page now claims it.

  • The tutorial’s ref: invitation no longer self-destructs. §3.6 showed a second [run] table that, pasted beside §3.5’s, was a TOML duplicate-table error naming a directory the tutorial’s layout doesn’t have; the fragments key now lives (commented) in §3.5’s one config block. “A suite is two things” undercounted its own mandatory proef.toml — it says three files now, and the tree shows all three. TROUBLESHOOTING stops listing hurl’s [Options] repeat: as if it were a proef step key.

  • The proef init scaffold goes green against the dev fixture. The advertised fastest path (init → fixture → test) ended 1 pass / 2 fail: the scaffold calls /search and /version, and the fixture served neither — a red first run that read as a broken tool. Both routes exist now, the whole path is pinned by an integration test, and the scaffold’s ref: fragment thereby executes against a live endpoint — the body form’s first runnable demonstration.

  • A failure no longer prints its detail twice. An engine fault quotes the failing step’s own detail, and the located step line just below printed the same ~200 characters again; when the fault message contains a failing step’s detail, the fault line now keeps the scenario identity and the step line carries the detail once.

  • JUnit failure and skip detail reaches every platform. The detail — assert diff, fragment provenance, @skip:reason, the quarantine notice — lived only in the message attribute; GitLab parses only the element text, and Azure maps the text to its stack-trace field, so half the platforms showed a bare failure (or a reasonless skip). Every non-success now carries both, and a failure’s text node additionally carries each failing step’s redacted reproduce hint — the content channel has the room the one-line attribute does not. Pinned alongside two library guarantees that were verified rather than assumed: quick-junit strips ANSI escapes and XML-1.0-illegal control characters on every setter (one binary response byte used to be the classic whole-report killer on Jenkins/GitLab), and time is plain three-decimal seconds; both now have tests so a dependency bump cannot shed them silently. A third pin: composed reports (suite + rerun-carried + teardown) yield each classname+name identity exactly once — GitLab silently drops duplicates.

  • The GitHub job summary can no longer vanish at the 1 MiB cap. The documented failure mode at GitHub’s limit is silent disappearance (and oversized writes have aborted jobs in shipped first-party actions); a failing rerun-overlay suite with per-tag tables crosses it more easily than it looks. The summary now truncates deterministically at a line boundary under a 900 KB budget, saying how many lines were cut and where the full detail lives.

  • ::error annotations budget for GitHub’s real limit. GitHub keeps ten error annotations per step and silently drops the rest — an uncapped emission made a forty-failure run look like exactly ten. The budget is now one annotation per failing scenario (its first failing step with detail, else its fault) capped at ten, with a closing ::notice naming what the ten are out of; title= is clipped under GitHub’s 255-character cap before encoding.

  • saveAs: global refuses a secret it can find, not just a secret it can equal. The gate lived in the hurl engine and matched whole-value equality against raw secret values — a capture merely containing one (Bearer <token>) or carrying an encoded reflection (base64/hex/percent/ JSON-escape) promoted to .proef-state.json in plaintext. The refusal now lives on the store’s owner (World::set_global), armed once per scenario by the runner with the same derived-needle set redaction uses (ADR-0005) — one needle list for both invariants, and every engine a scenario dispatches to is covered. The invariant is now genuinely property-tested (any composite carrying a guarded secret never enters the store), as CLAUDE.md had claimed of the single example test.

  • The SLA gate honors @quarantine. sla::check measured every scenario while the exit code excluded quarantined ones — so a quarantined, timing-marginal scenario (exactly what gets quarantined) could not fail the run on its assertions but still turned it red on latency. The latency population now applies the same non-gating list as the exit code.

  • A record that travels can no longer lie, crash, or steer. Reading a record predating scenario_finished.file (or any foreign baseline whose closes key under the serde default ""), the step buffer never attached: every scenario read as step-less, flaky could never see a retry or a duration, and diff --fail-on-regression certified green over empty step maps — the close now adopts its steps’ file when exactly one pending scenario matches by name (pinned by test). The head fold’s “first head wins” guard tested emptiness rather than position, so a second run_started in a concatenated or legacy record overwrote the run’s env/metadata/rerun_of wholesale (pinned by test). rerun_of — a string read out of the record — was joined onto the runs root unvalidated, so a crafted "../../elsewhere" spliced a foreign file’s events into the rendered report; it must now be a single path component, and --run-id gets the same rule at the CLI edge (a typed clap error on separators or .., on all four commands that accept one). Record reads gained a generous 256 MiB ceiling — the one input loaded with no bound — and every duration sum over record-supplied u64s (HTML report, tag table, flaky) is now saturating instead of a debug-build panic on a corrupt file.

  • A [tag-links] template can no longer be subverted by a tag’s spelling. The GitHub-summary sink substituted the tag into the URL raw, so @JIRA-1)[x](y closed the markdown link early and injected content into the job summary; the tag is now percent-encoded in the URL slot. Both sinks (HTML report and summary) also render non-http(s) templates as plain text rather than minting javascript:-class links.

  • An inverted Span degrades instead of exploding: Span::len and the SARIF byteLength are saturating — B1’s shipped class, closed in the type rather than at one construction site.

  • Eleven sites that swallowed an error and reported success now speak. The class the v0.6.0–v0.8.0 series was named for, still present at the edges: a poisoned store lock silently skipped persisting the World (every saveAs: global promotion of the run lost — now recovered, matching the runner’s own policy, which also stops failing an innocent scenario for another thread’s panic); an unreadable subdirectory silently shrank the suite to a confident “0 failed” (now warned, per entry too); doctor reported a clean “no packs” over a tree it could not read (now a Fail row) and fmt formatted nothing while reporting success (now warned); a non-UTF-8 environment value read as “not set” — the wrong cause — for ${env:…} (now named up front); a .map.json serialization failure was the one silent write in the run record (now warned); proef lsp booted with defaults over a proef.toml that exists but does not parse, silently diverging from the runner (now says so on stderr); one unreadable run aborted all of proef flaky (now skipped and counted, with the two-run floor re-applied over what was readable); xtask docs-check printed “aligned” when it could not read the directories it checks (now a failure); and a mis-severitied diagnostic pushed into the lowering error sink vanished entirely (any error-sink entry now fails the scenario).

  • fmt normalizes every literal-block spelling. The scan required the key line to end with |, so hurl: |-, |+, an indent indicator, or a trailing comment — all loadable — were silently skipped and --check certified them canonical. Folded scalars (>) stay out deliberately: YAML folding rewrites the line structure there is nothing line-preserved to normalize.

  • A Ctrl-C landing in --watch’s debounce window no longer launches one more full suite run. The ≥300 ms drain between “change detected” and the rerun never checked the interrupt, and the rerun then minted a fresh cancellation token — so the handler cancelled the finished run’s token, printed “leaving watch”, and a whole suite executed anyway. The interrupt is now checked inside the drain and again after the new token is stored, so a Ctrl-C from any point forward cancels the token the run actually carries.

  • --watch can no longer go silently deaf. A delivered watcher error and notify’s rescan signal (the kernel-queue-overflow event a git checkout burst produces) were both discarded by the event filter — the watch kept printing “watching … for changes” while missing every change. Both now retrigger a run, saying why. Two adjacent silent paths gained voices too: a runs dir whose path has no final component now warns that its writes cannot be excluded from the watch (the self-feeding-loop shape), and a failed Ctrl-C handler registration now says the two-stage interrupt is unavailable instead of silently dropping the contract.

  • A comment on a section header no longer blinds the scans that gate on it. hurl’s own section_name parser leaves the rest of the header line to the ordinary comment terminator, so [Options] # tuning is a real section — but proef’s scans required whole-line equality. Behind a commented header, validation pass 6 was off entirely: retry: -1 dry-ran clean (the abandoned-thread hole ADR-0007 exists to refuse), the delay cap and the double-declaration check with it, in inline blocks and fragments alike. The same equality bug made [Captures] # ids drop every capture under it from .map.json, and [Asserts] # note open a second section under an expect: merge. One is_section_header recogniser now serves every section scan.

  • delay: 5h is refused like delay: 90m always was. The duration table knew ms/s/m but not hurl’s h, so an hour-spelled delay five times over the 1-hour cap fell through the suffix parse and validated clean. The table now mirrors hurl_core’s DurationUnit in full.

  • A pack-scope bind: value resolves in the pack’s scope, not in whichever macro reached it first. The table resolved through the first ref-using macro’s argument scope and was then cached for the scenario — a bare ${param} silently took that macro’s value everywhere (or vanished, blaming an innocent macro). The pack table now resolves arg-free and default-free: namespaced references (${url:…}, ${vars:…}, ${secret:…}, ${fake:…}, ${env:…}) are its vocabulary, and a bare ${name} is a deterministic error attributed to the pack’s own bind: in every macro order.

  • A star-heavy tag atom can no longer hang selection or abort the process. The glob matcher was naive recursion: backtracking was exponential in the * count (a 19-character atom took seconds per tag per scenario) and recursion depth grew with pattern length (a long enough atom in --tags, [run] exclusive-tags or [tag-links] overflowed the stack — SIGABRT, outside the exit contract). Rewritten as the standard two-pointer match: linear-ish, iterative, oracle-property-tested against the old semantics.

  • A bound value carrying a lone \r is refused at lower time. lower::multiline_bind tested \n alone, so a carriage return (a value read off a CRLF file) sailed into the emitted [Options] variable: line and died one stage later as emit::invalid_artifact — blaming generated text the author never wrote. The guard now refuses any control character except tab.

  • An HTTP2-Settings: request header no longer mis-slots an expect: merge. The last-entry scan recognised a response line by the bare prefix HTTP, which the emitter’s own recogniser was already hardened against; both now share one is_response_line (HTTP / HTTP/).

Breaking

  • --output split by meaning: --format chooses a format, -o/--output names a path. test takes --format json|tap; the listing commands (flows, macros, fragments, flaky) take --format json — each through its own enum, so clap’s help can no longer advertise tap on four commands whose runtime rejected it (the old shared enum lied about a quarter of the surface, and -o changed category between siblings: format on five commands, directory on artifacts, file on report). --output json/--output tap no longer parse on those five commands — clean break, no alias; artifacts/report keep -o/--output for their paths, unchanged. The runtime json_only check is deleted: the type system does its job now.
  • Library: World::set_global returns bool (#[must_use]) — false is a refused promotion — and World gains guard_secrets; Redactions gains the taints probe. The hurl engine’s private equality-only gate is deleted in favor of the World’s.
  • Library: ConsoleReporter::new takes a fourth color: bool — the TTY/NO_COLOR probe stays at the CLI edge; the sans-IO core takes the answer as a plain value.

[0.15.0] - 2026-08-25 (the Robot Framework audit: visible skips, tag verdicts, explicit metadata)

Breaking

  • A quarantined test-failure reaches JUnit as <skipped> with a message, not <failure> — Jenkins marked builds UNSTABLE while proef exited 0; every dashboard now says what the exit code says (ADR-0019). Library: ScenarioSpec gains skip, ScenarioOutcome/ScenarioRun gain reason, Event::ScenarioFinished gains additive reason, write_junit takes the non-gating list.
  • --shard assignments re-deal: the hash gained a mixing finalizer. Raw FNV-1a’s low bit is the XOR-parity of the input bytes, so a scenario named after its feature file — the commonest Gherkin convention — collapsed to one shard at N=2 and left odd shards empty at N=4, silently (the empty shard exits 0). shard_bucket now finalizes with Murmur3’s fmix64; every scenario re-buckets, so all jobs of one matrix must run the same proef version (already true in practice). Round-18 finding, reproduced and mechanism-verified before fixing; the balance test gained the name-mirrors-file corpus it was structurally blind to.
  • Tag atoms glob. * and ? in a --tags / [run] exclusive-tags atom are now anchored wildcards (@FRD-* selects the family; ? is one character) — previously they were literal characters that silently matched nothing, the trap this closes. Metacharacter-free atoms are bit-identical to before, property-pinned. Case stays sensitive.
  • JUnit test identity is classname + name. classname carries the feature file, name the scenario alone; the old single name embedded file:line, so an edit above a scenario re-identified every test below it in Jenkins history and GitLab’s MR diff. Anything keyed on the old file:line name strings must re-key. The suite skipped count is now spelled skipped (was disabled, which no consumer reads).

Added

  • [tag-links] turns tag cells into tracker links (RF’s --tagstatlink, reduced to one mechanism): tag glob → URL template with {tag} substituted, honored by the HTML report’s by-tag table and the GitHub summary; the pattern language is the same anchored glob --tags uses. Library (Breaking): render_html takes the link map; tags::atom_matches_public exposes the one matcher.

  • --console dotted|quiet (RF wave 3): one glyph per scenario (. pass, F fail, s skip, w warn — lowercase is non-gating, the pytest/RF convention, flushed per glyph, wrapped at 80) or just the frame. Purely presentation: the record, every report, the post-pool failure details and the exit code are identical in every mode; run.log mirrors the console verbatim, dots included — events.jsonl is the full truth. Library (Breaking): ConsoleReporter::new takes a ConsoleMode.

  • A --rerun now produces the one JUnit and the one report that cover the whole suite (E2’s rerun half; Robot Framework’s rebot --merge shape, done as composition): the run head records rerun_of, the JUnit carries the base’s not-re-run scenarios as ordinary testcases, and proef report overlays the base into a merged page (banner named, base timestamps stripped so timelines never mix, rotated-away base degrades loudly). Exit code and totals stay the rerun’s own.

  • --meta key=value and [meta]/[env.<name>.meta] record explicit run metadata (ADR-0020, RF wave 2): commit, build URL, team — recorded in the run head, shown by the HTML report, GitHub summary, explain, diff (which now also warns on cross-env comparisons) and the --output json body (additive keys). The active --env profile name and the --shuffle marker ride the same head. proef never harvests: no git, no hostname, no CI env sniffing — the shell harvests, proef records. Everything passes the sink-boundary mask, keys and values both. Library (Breaking): RunRecord::open and exec::execute take the head inputs.

  • Per-tag verdicts in the HTML report and the GitHub summary (RF wave 2): tags now reach the record — additive tags on scenario_finished (finished-only: the cancel-skip path emits no start), additive exclusive on scenario_started (closes R11-6, the scheduler’s own bool) — and both reports roll them up per tag (suite-only, Warned counts with passed). Requirement-tagged suites (@FRD-3.1) get their traceability matrix for free. Tags are deduped at the one accumulation point (first occurrence wins); the quarantine list is now derived from the outcomes’ own tags — one owner, same behavior, pinned by the exit suite. Library (Breaking): ScenarioSpec/ScenarioOutcome gain tags.

  • @skip and @skip:<reason> park a scenario visibly (ADR-0019, RF wave 2): counted in every total, reasoned in the console, JUnit, TAP, the record, the HTML report, explain and flows --output json; the harness maps it to libtest’s ignored flag. All-selected-skipped exits 0; the empty-selection refusal stays exit 2. --tags "not @skip*" unselects both spellings; an authored skip is never re-queued by --rerun, and diff gives skip transitions their own bucket instead of reading them as fixed.

  • flows shows the feature description. The prose block under Feature: was parsed and then dropped — the one paragraph written for exactly the reader flows serves never reached them. Human output prints it under the feature header; --output json rows gain featureDescription: string|null (additive). Library: FeatureFile gains description.

  • --shuffle re-deals the execution order, seeded by the run id — one determinism knob for order and fakes alike, so --shuffle --run-id <id> reproduces an order-dependent failure exactly (Robot Framework’s --randomize, minus the parallel seed it threads separately). Applied after --shard, so membership never moves; under --watch every unpinned rerun re-deals, deliberately. The permutation is version-stable and pinned. Recording a shuffled marker in the run head is deferred to the planned RunStarted additions (env/metadata), one wire change instead of two.

  • The failing step’s reproduce: curl … reaches the record. The engine always computed the redacted curl and the live console always printed it — and the record dropped it, so explain and the HTML report knew less than the console did. StepFinished gains additive reproduce_hint (absent on passing steps and every pre-field stream); explain and the report print it; the sink-boundary mask covers it like detail.

  • README documents every flag the binary exposes, enforced. v0.14.0 shipped --shard and --max-fail with no README mention; the docs gate gains the flags direction (same vacuity guard as the command half), and the measured gap — those two plus schema --add-to — is closed.

  • JUnit carries what GitLab and Jenkins actually read (R3-6, specced from GitLab’s parser docs and Jenkins’ SuiteResult.java): file on each testcase (GitLab source linking), time on suite and root. timestamp and hostname stay absent deliberately — ignored or substituted by both consumers, and a hostname would undo R12-1’s provenance fix.

  • The docs corpus is a website: https://emrecdr.github.io/proef/. mdBook renders docs/ on every push to main that touches it; the nav is docs/SUMMARY.md, which the existing docs gates link-check like any other doc, and the pages workflow refuses a corpus doc that is not on the site. The crate homepage points there from the next release.

Fixed

  • A failure detail is bounded before it reaches any sink. hurl’s rendered assert error quotes the actual response, so a failed assert on a large body rode full-size into the record, JUnit, the HTML report and the GitHub summary at once. The engine now middle-cuts past 40 lines / 8 KiB with a marker naming the elision; the artifact pointer survives outside the cut, and the full output is one re-run away (Robot Framework’s 40-line rule, adopted at the boundary where all sinks are covered at once).

  • The machine-body contract closes its last two paths: an empty selection (--scenario/--tags matching nothing — loud exit 2 by design) and a corrupt global-state file both emitted zero stdout bytes under --output json.

  • Identical errors collapse like identical warnings — a broken macro usually fails to lower everywhere, so the error wall was the more common fifty-block wall; distinct errors still render separately, and SARIF keeps every site.

  • Injected [Options] lines respect every section-ending shape. The section-end move covered one shape of five: an unfenced JSON/XML body after an author [Options] swallowed the injected lines into invalid hurl (exit 2 on input that worked before), and an entry with an author section but no response line leaked its pending lines into the next entry, where hurl parsed retry: as an HTTP header and the artifact validated green. The section now ends at the first line that could not sit inside it.

  • A # inside a bind: value no longer hides the reads after it. The template probe parsed the value in an unquoted position where # opens a comment; it now probes the quoted variable: position bake actually injects into, so "{{a}} # {{b}}" reports both.

  • A setup that fails to load still emits the machine body — the last terminating path returning zero stdout bytes under --output json.

  • SARIF keeps one result per site again. The warning collapse shipped at the front-end aggregation, which also feeds SARIF — a code-scanning consumer lost every anchor but the first. The collapse now happens at console rendering only; SARIF carries all sites, the console one line with the count.

  • Every terminating path emits exactly one machine body (R17-2.3/2.4). An empty shard printed its prose note as the --output json body — jq failed on the very path a sharded matrix guarantees one job takes — and a setup abort printed nothing at all while JUnit carried the failure. The note now goes to stderr under machine output (and lost a stray-space run); never-ran paths report ADR-0014’s suite-only zeros with the exit code carrying the verdict.

  • A failed teardown reaches JUnit as its own suite (R17-2.5) — #78’s rule made symmetric: a phase appears in the reports when it fails. A gated pipeline used to read a fully-passing report on an exit-3 run.

  • A repeated warning is one warning with a count. One authored mistake in a macro shared by fifty scenarios rendered fifty times; identical warnings now collapse to their first occurrence plus “(N sites across the suite)”.

  • explain’s truncated-record fallback counts the suite only — a record that died mid-setup folded the phase scenario into the totals three lines above the label saying phases are excluded (ADR-0014).

  • bind: values are read by hurl’s parser, not a text scan (R17-2.2). A hurl function ({{newUuid}}, {{newDate}}) no longer counts as an unbound variable — proef refused input stock hurl runs — and a sibling literal bind whose name sorts earlier now counts as a supplier, since injected [Options] variable: lines are written and evaluated in name order. The seam answers the question once: FragmentSupport::template_reads.

  • A fragment’s own [Options] variable: lines now evaluate before the injected ones. Injection used to land at the section head, so the fragment-supplies-it route the unbound check accepts was assigned too late to be read at run time — accepted at dry-run, wrong at execution.

  • A {{x}} inside a bind: value is validated at --dry-run, not at run time. hurl templates the injected [Options] variable: line when the entry runs, so a name nothing supplies used to pass dry-run and die mid-run; proef::lower::unbound_placeholder now names the placeholder and the bind key at lower time, where the capture set is known. What legitimately supplies it: an earlier step’s capture, the fragment’s own [Options] variable:, a secret in scope, or a sibling literal bind whose name sorts earlier (injected lines are written and evaluated in name order).

  • A literal bind: that shadows an earlier capture is named, not silent. hurl’s variable: assigns into one shared set, so the bound value replaces the captured one from that entry on — sometimes intended, so it is a warning: proef::lower::bind_shadows_capture. A secret bind cannot shadow (it skips the [Options] path) and draws no warning.

  • A failed [run] setup reaches JUnit, the GitHub summary, and PR annotations. The abort used to return before the CI-report block, so a job gating on --junit saw no file at all — indistinguishable from proef never running. The reports now carry the setup scenario itself (suite named by the setup feature file); exit codes are untouched (ADR-0014), and nothing is fabricated for the pool that never ran.

Internal

  • The machine body has one exit. execute’s six terminating paths each pasted the same empty-body emission; they now return through a single funnel with the one emit_machine_body call after it, so a new path cannot forget the contract — and the empty-selection body takes its exit code from the refusal itself instead of restating it. Post-merge cleanup pass over the deep-audit cycle; behavior pinned by the existing path tests.
  • The probe and bake share the whole variable: line. template_reads re-spelled the injected [Options] variable: line by hand around the shared escaper; proef_core::lower::variable_option_line now builds it for both, and quote_option returns to being private (library-surface swap; unreleased either way). The section-end flush in bake_entry_options also drops its fence-branch duplicate — one check covers all shapes — and the console collapse builds its annotated message without cloning the diagnostic on the common single-site path.
  • The canary stopped trusting the index’s tail twice over: it skips prerelease versions (the sparse index is publish-ordered, so a 9.0.0-beta would have become “latest”), and refuses a backport older than the pin by semver ordering (an 8.0.2 published after 9.0.0 would have produced a green about a downgrade).
  • deny.toml’s advisory ignores were dead and are gone. The quick-xml pair was ignored under “the patched release is unreachable” — the quick-junit 0.7 bump made it reachable and the workspace has been on the patched line; the stale ignores would also have silenced any new advisory against it. quick-xml itself re-pinned to quick-junit 0.7’s in-tree copy (=0.41.0, one lock generation); lsp-server rides to 0.10, toml to 1.x.

Documentation

  • The corpus tells the truth again, audited claim-by-claim: CONFIG documents the [env.<name>.run] jobs-only rule a reader used to discover as a parse error; GETTING-STARTED can produce its own output (it now states the fixture token its §5 requires, and its reproduce command names the --secret the replay needs); EVENTS carries the provenance, totals, and field facts consumers implement against; TESTING-STRATEGY describes the CI that exists; RELEASING’s gate list predicts CI; TROUBLESHOOTING’s exit-1 row covers the --check family. The #N in an outline instance is documented as positional, with the column-placeholder naming that keeps identity stable across --shard and JUnit history.

[0.14.0] - 2026-08-18 (proef at CI scale)

Fixed

  • --rerun after a cancelled run continues it, instead of a false green. --max-fail (and Ctrl-C) stop a run early with the never-reached scenarios honestly recorded as skipped — but --rerun filtered to failures alone, so stop → fix → rerun ran only the old failures and reported exit 0 with most of the suite never executed in either run. Reproduced live before fixing (found by round-15 external review): stop at 2 of 6, fix, rerun → 2 passed · 0 failed, green, four scenarios untested. On a cancelled base record --rerun now runs failures plus the cancellation-skipped tail, and says so (note: the last run was cancelled before N scenario(s) ran…); scenario-level skips only exist under cancellation, so a completed base keeps the old semantics exactly. This also changes --rerun after Ctrl-C — continuing the unfinished work is what stop → fix → continue always meant. Mutation-tested: reverting the union fails the continuation test.

Added

  • proef test --shard I/N — stable hash-mode sharding (R3-3). A CI matrix runs --shard 1/N … N/N on separate machines; scenarios are assigned by a frozen FNV-1a hash of the run-wide (file, scenario) identity, so adding a scenario never re-buckets the others — the measured stability argument that rejected index-slicing at triage (inserting one scenario re-bucketed the whole shifted tail under slicing, nothing under hashing; the shard tests pin both directions, and the assignment itself is frozen by literals — the hash is a published contract, and changing it would be breaking). Sharding applies after every other selector (the pinned filter→shard order), so each matrix job partitions one agreed-on set. An empty shard of a non-empty selection is a note and exit 0 — a small suite over a big matrix is a fact, not a mistake — while an empty selection keeps the loud typo’d-filter refusal, sharded or not.

  • proef flaky — flakiness verdicts over the retained run history (R3-2). The 2026 discipline is detect → quarantine → resolve, and proef already owned the middle step: @quarantine runs a scenario without gating the exit code. This is the missing detect, a fold over the records runs-dir already retains — the window is [run] keep-runs, and no new state is written. Three signals from fields the record already carries (ADR-0008): flapping (verdict changed between consecutive observed runs more than once — transition-counting, not fail-rate, which is what separates flaky from broken: a scenario failing every run is consistently broken, a different problem), passes only on retry (green, but some step needed more than one attempt — the latent flake pass/fail-history tools structurally miss; the record keeps per-step attempts), and always failing. A cancellation-skipped row is not evidence and does not count toward a scenario’s history; phases are excluded (ADR-0014). --output json emits one object per scenario with the counts behind each verdict. Fewer than two runs is refused (exit 2), the same answer diff gives.

  • proef test --max-fail N stops the run after N suite-scenario failures (1 = fail fast) — the convention Playwright (--max-failures), pytest (--maxfail) and cargo-nextest (--max-fail) share, with the shared honest semantics: in-flight scenarios finish, the never-run rest record as skipped (not absent, never passed), and teardown still runs on its own token. The stop rides the graceful-cancel path Ctrl-C already exercises, so the record is a complete cancelled run — which diff --fail-on-regression already refuses to certify, exactly right for a deliberately-partial one. [run] setup/teardown failures never count toward the threshold (a broken fixture is not a failing test, ADR-0014).

Documentation

  • The R3 enhancement registry is triaged (OPEN-FINDINGS): --max-fail built; a flakiness verdict over the run history and hash-mode sharding validated as build-next (the 2026 flaky pipeline is detect → quarantine → resolve, and the @quarantine tag already owns the middle step); CTRF, pack doc and the pre-M6 seam refactors deferred with named triggers; OTel and Cucumber-Messages exporters declined under ADR-0008’s one-record rule; items defined only in the absent v1 research document held for a spec.

[0.13.0] - 2026-08-17 (a record that travels, and a secret that stays one)

Added

  • proef diff takes a path. Each side is now a run id, a record directory, or an events .jsonl file under any name — the stream is the record (ADR-0008), so all three must mean the same thing. The file form is the CI baseline flow an adopting suite asked for: download the base branch’s events.jsonl artifact and proef diff baseline.jsonl <new> --fail-on-regression gates the PR, with no shared record store. Previously every argument was joined onto runs-dir, so a path produced .proef-runs/<your path>/events.jsonl: No such file — the argument mangled into the complaint. A path that does not exist now names itself; a --baseline flag was considered and declined as a second spelling of the same positional.

  • [run] keep-runs bounds how many past run records runs-dir retains. The policy already existed as a hard-coded 200; it just could not be expressed, so a suite re-run on every save accumulated records for a day with nothing signalling a ceiling. 0 keeps none but the run in flight. Rotation still only ever deletes directories named by a generated run id — runs-dir may be . — so a --run-id <name> record sits outside the budget and is never rotated, now stated in CONFIG.md rather than left to be discovered. Filed as R12-2.

Fixed

  • A run record no longer names the machine that produced it. [run] suite resolves against the config directory (0.12.0), so a path-less proef test handed the front end an absolute path — and every emitter printed it: the .hurl # source: header, .map.json’s feature.file, every step_finished event, the console, and pack diagnostics. Two checkouts of one suite stopped producing equal artifacts, which is the property ADR-0010 exists to guarantee; an adopting suite hit it as /Users/… in 133 artifact lines and 64% of its event stream by bytes.

    The resolution rule was right and stands. What was missing is its naming dual: resolve against the project, then name against the project again. front::SourceNaming is now the one boundary that answers “how is this path spelled”, for features, packs and fragments alike — replacing the fragment corpus’s separate cwd-relative strip, which was a second anchor for the same question. The four ways to name one suite — derived from [run] suite, typed, typed absolutely, or reached from a subdirectory — now emit one artifact, byte for byte.

    A path that arrives relative is recorded exactly as it arrived; a suite or corpus genuinely outside the project keeps its absolute name, there being no project-relative spelling of it. Filed as R12-1, and it closes R9-6, which had described the same defect as safe from the project root — it no longer was.

    Breaking, by the rule in docs/RELEASING.md: it changes emitted artifact bytes, which is inherently breaking and takes a MINOR bump. Migration: nothing to do for a suite invoked with a typed relative path — those bytes are unchanged. A tool reading step.file or feature.file out of a record produced by a path-less run now sees a project-relative path where it saw an absolute one; join it onto the directory holding proef.toml. Records written by earlier versions are not rewritten.

Security

  • An encoded reflection of a secret is redacted (S1). Redaction was exact-match on the raw secret bytes, and a server that reflects a bearer token encoded — an OAuth introspection endpoint, a debug echo, a JWT claim — defeated it: a failing assert quoted the base64 form in its detail, and a string trivially base64 -d-able back to the live credential reached the console and events.jsonl, the retained record CI uploads. Demonstrated live against 0.12.0 by an external research pass and reproduced here before fixing. Redactions::new now derives each secret’s common encoded forms as additional needles — base64 (standard and URL-safe alphabets, with and without padding), hex (both cases), RFC 3986 percent-encoding, and the JSON-string escape — so every construction site (the CLI sink, the engine’s internal renderer, TAP) is covered by construction. This is the remedy GitHub’s own log-masking documents for the same limitation: register each transformed value too. The needle set covers the reversible transforms that occur at HTTP boundaries and does not claim completeness — a secret reflected hashed or re-encrypted matches no needle list. Over-redaction is the accepted failure direction. Property-tested over every derived form, pinned end-to-end by a fixture route that echoes the bearer base64-encoded, and recorded as an ADR-0005 amendment.

  • The fragment corpus read is bounded. [run] fragments names a directory proef did not write and does not control, and it was read with no per-file or total cap: a 279 MB file cost 601 MB of resident memory on proef flows — a command that never looks at a fragment — because the text is read whole and then copied into an Arc<str>. A file over 8 MiB is now skipped (proef::pack::oversized_fragment_file) and the reader stops past 64 MiB total (proef::pack::fragment_corpus_too_large). The size comes from the directory entry, so an oversized file is never allocated at all; the same bound applies in proef lsp, where the corpus is held between requests rather than for one command. Skipped, never fatal — a corpus is foreign by design, so one bad file must not sink the ones beside it. An unreferenced corpus still costs nothing: the scan stays lazy, so nothing is reported unless a pack actually names a fragment. Filed as R9-3.

Internal

  • A hung test is now a five-minute failure, not a five-day zombie. The nextest config had slow-timeout with no terminate-after, which only labels a test SLOW and never kills it — an lsp_stdio test wedged on an unbounded child.wait() ran for five days with its proef lsp child alive. Both layers fixed: the two bare child.wait() sites got the file’s own bounded-watchdog pattern (a server that fails to exit now fails the test in 10s, naming what did not exit), and the runner gained terminate-after = 2 (120s), sized from a cold-cache census of the whole suite (slowest ordinary test: 5.1s). The harness_ trio — which shells cargo test inside the test and measured 216s on a fully cold cache — gets a per-test override to 600s, the nextest docs’ own tight-global-plus-overrides pattern. The process-group kill (a spawned server dies with its test) was verified empirically with a deliberately hung test holding a live child.

  • Cleanup pass over this cycle’s four PRs (reuse/simplification/efficiency/ altitude review). The corpus-bound decision moved into core as pack::CorpusBudget — it was abstracted in the CLI and hand-copied in the LSP, agreeing by copy rather than by construction; both readers now share it and only measurement stays reader-local. Redactions stopped allocating on the miss path (nearly every call: per string field per event under the reporter-stack mutex, with the needle list ~9× larger since the encoded forms) — clean fields now hand back their original Arc. A relative source path is left exactly as it arrived, per its documented contract — it had been falling through to a per-file canonicalize that could rewrite a ../-typed spelling. The LSP’s percent-encoder folded onto core’s (byte-identical copies, one character set to drift). The fixture’s hand-rolled base64 became the crate call — its dependency-surface rationale died when this same cycle made base64 a workspace-wide compile. diff’s path-or-id resolution moved beside its sibling in record. A deny.toml home for the curl floor was tried and reverted by mutation test: cargo-deny 0.19.8 mismatches build-metadata versions (curl-sys@<0.4.90 banned the good 0.4.90+curl-8.21.0); the floor stays a unit test, now scanning every lockfile entry rather than the first.

  • The bundled libcurl cannot silently regress under the June-2026 CVE batch. curl-sys 0.4.90+curl-8.21.0 in the lockfile is past the batch — but only as a transitive accident of resolution, and the usual gates are structurally blind here: RUSTSEC carries no advisories for CVEs in a *-sys-bundled C library, so cargo audit/deny stay green however stale the bundled curl is. A test now asserts the lockfile floor, and each release build prints the libcurl actually linked into that artifact (proef doctor already reported it; the release log now carries it per target). The hurl-8.1 watch items — variables-file:’s missing sandbox first among them — are recorded as a pin-bump checklist in the thin-fork runbook.

  • Fuzzing reaches the fragment rules. fuzz_pack_load ran against an empty corpus, so ref: resolution, bind: keys nothing reads, a bind: colliding with a variable the fragment supplies itself, and unbound placeholders were covered on paper and unreachable in fact. The new fuzz_fragment_binding target is structure-aware: it builds a well-formed pack and corpus and spends its budget on the name space where those rules live. That shape was chosen from measurement, not taste — a byte-oriented version never once resolved a ref: in 1.45 million runs, because reaching the rules meant discovering valid YAML and a matching corpus at the same time. The corpus is read by a synthetic scanner rather than hurl’s, which is what keeps the fuzz workspace free of native libraries: cargo dependencies are package-level, so one engine-dependent target would compile hurl for all of them.

  • Hurl’s own annotation scanner is property-tested, in proef-engine-hurl where the native libraries already are. The properties pin what the entry-boundary arithmetic is for: every reported line lies inside the file, every entry is accounted for exactly once, the starts are ordered and distinct, and — the one that matters — no fragment’s text runs into the entry after it. That last assertion exists because a first draft without it passed while the boundary was deliberately broken.

  • The fuzz target list comes from cargo fuzz list. It had been spelled out in ci.yml and nightly.yml, so a new target ran nowhere until both were edited, and nothing failed to say so.

[0.12.0] - 2026-08-14 (one path rule, and a watcher that stops lying)

Fixed

  • A runs-dir edited mid---watch no longer feeds the loop its own output. Reruns re-read the config (the fix below), so records went to the new directory while the watcher’s exclusion still named the one it had frozen at startup — and every rerun’s artifacts/*.hurl, now under an unexcluded directory, requeued the next run. One edit produced 39 runs in 12 seconds, firing real traffic. This was the third outing for the watch-feedback class, so the fix removes the second answer rather than resynchronising it: each rerun registers where it is about to write, and the exclusion is derived from the same config the run is. A directory a previous run wrote stays excluded too, since its events can still be in flight. Filed as R11-8.

  • --watch --config <relative path> retriggers on config edits. The watcher compared the config by exact path while notify reports events under the spelling the OS resolved them to, so --config proef.toml never matched and config edits produced nothing — silently, because feature edits kept working and the loop looked alive. Symlinked and /tmp-style aliased paths failed the same way and are also fixed: the flag is made absolute when it is stored, and identity is settled by comparing canonical paths, which is a stricter question than being absolute. The same relative-path flaw silently cost proef lsp --config <relative> go-to-definition across the whole fragment corpus, since documents::name_to_url refuses a relative name. Filed as R11-9.

  • doctor reports a proef.toml that will not parse. The discovery arm had become a silent unwrap_or_default, so a malformed config left doctor reporting on invented defaults and printing “all checks passed”, exit 0 — with the parse error, which the previous code printed, discarded. It is a project: row now, so it reaches worst and the exit code a CI script actually reads. Being absent is still not a finding: doctor must run outside a project. Filed as R11-10.

  • proef fragments exits non-zero when a [run] setup/teardown phase fails to load. It printed error: setup feature failed to validate: and exited 0, because the phase half flattened its failure to “not measured” while the suite half kept its code. Withholding the counts was right; reporting success while printing errors was not.

  • proef.toml has one path rule. A path written in the config now resolves against the directory holding the config; a path typed on the command line still resolves against the working directory. [run] fragments already worked this way and everything else did not, so two keys in one table meant two different roots: from a subdirectory fragments = "hurl" resolved while suite = "features" reported “neither a feature file nor a directory”. With --config the split was worse than inconsistent — pointing at a config in another tree ran dry-run OK over whatever suite happened to sit beside the shell, and never looked at the configured one.

    The rule now covers suite, setup, teardown, runs-dir and the tests/ convention probe, plus two files nothing had inventoried: .proef-state.json (the persistent World) and .proef-secrets.json (the secret store), which were anchored on the working directory — so two shells in one project were two Worlds and two secret stores. It is the convention Cargo, tsconfig.json and pytest’s rootdir all follow. Absolute values are taken as written, and with no proef.toml in scope written paths stay relative to the working directory, so the config-independent reference corpus is unaffected. Filed as R11-1.

  • --watch rereads the config it retriggers on. Editing proef.toml retriggered a run that still used the snapshot loaded at startup: changing [url] base produced a rerun that dutifully called the old host, and the same went stale for jobs, [env.*] and exclusive-tags. Watching a file whose contents you then ignore is worse than not watching it, because the rerun reports that the edit was taken. Each rerun now re-reads the file and re-resolves the suite from it; a config that no longer parses fails that rerun and leaves the loop watching, since half-typed TOML is the normal state of a file being edited. Which directories the loop watches is still fixed at startup, so changing [run] fragments or [run] suite needs a restart to be watched. Filed as R11-2.

  • --config is honoured or refused by every subcommand. doctor printed the error for a missing named file and then reported on defaults, exit 0 — the “fall back to defaults” CONFIG.md forbids — while fmt, init, schema and secret accepted a nonexistent path silently, against the “global to every subcommand” claim in CONFIG.md, README.md and this file. A named file that is not there is now exit 2 everywhere, including where nothing reads it; doctor stays lenient about discovery, which is a different claim. secret additionally uses the flag, since the store is the project’s. Filed as R11-3.

Breaking: the secret store, the persistent World and the run records move with the config rather than with the shell. What decides whether this reaches you is where you invoked proef, not where proef.toml sits: runs started from the project root are unchanged, but a run started from a subdirectory used to write .proef-state.json, .proef-secrets.json and .proef-runs/ beside the shell, and now writes all three beside the config.

Nothing is migrated, and none of it announces itself. A World written from a subdirectory reads as empty, so saveAs: global values start over on the first run after upgrading; stored secrets read as absent; and the old run records are simply invisible to explain, report and diff, which say “no run records” rather than erroring. To carry them over, move .proef-state.json, .proef-secrets.json and .proef-runs/ from the directory you used to run from into the one holding proef.toml. Otherwise re-run proef secret set and take a fresh baseline.

Breaking (library): proef_cli is not a published library surface, but for the record front::run takes the state-file path, ProjectConfig::runs_dir returns a PathBuf, setup/teardown return Option<PathBuf>, suite is gone (fold into default_suite_path), and the secretstore entry points take the store path. proef_core gains one item: pack::FragmentCorpus::unreadable_file.

  • [run] exclusive-tags validates itself. --dry-run did not parse the expression at all, so a malformed one exited 2 from proef test and passed dry-run OK … 0 warning(s) from the gate CI runs. And a well-formed expression matching no scenario was silent: @soloz against a @solo suite put every scenario back in the shared pool, exit 0, nothing said — the exact silent degradation the key was designed as a config expression to prevent, and one that reads as flakiness rather than as a typo. Both paths now parse it, and a zero-match expression warns, naming it and pointing at proef flows. Judged over every scenario the suite loaded rather than the ones selected, so a --tags filter that removes the matches from one run is not reported as a broken setting. Filed as R11-4 and R11-5.

Changed

  • proef fragments says which half it could not measure. --check reported “needs a suite that binds” when the suite had bound perfectly well and a [run] setup/teardown feature was the thing that failed to load, sending the reader to inspect the half that was fine. The degraded listing also now carries the note macros prints, so withheld counts read as “not measured” rather than as a corpus nothing uses.

  • proef fragments --check refuses to pass with no corpus configured. With [run] fragments unset it printed 0 entries and exited 0, indistinguishable from a fully-used corpus — so a CI gate disarmed silently the day the key left the config. The listing still works; only the gate is now a user error.

  • proef fragments --output json carries annotated on both row shapes. The annotated and unannotated rows differ in eight fields, and consumers had to probe for the absence of one to tell them apart.

Documentation

  • CONFIG.md’s “everything else keeps running at jobs width” was false: queueing is strict FIFO, so nothing new starts while an exclusive scenario waits at the head. The cost is bounded, not absent, and is now described.
  • The one caveat [run] exclusive-tags carries is written down in CONFIG.md and ADR-0007: exclusivity is enforced against the dispatcher’s active set, which a watchdog-abandoned scenario leaves while its detached thread is still issuing requests (hurl cannot be cancelled mid-entry).
  • TECH-SPEC §10 gained proef fragments and the global --config; §11’s [run] inventory listed three of seven keys.
  • DIAGNOSTICS.md carried a pack::load row nothing emits — a reader who grepped it found a plausible cause that could never be one — and filed lower::multiline_bind under proef::pack::*. Both fixed, and the two-way agreement between the file and the emitted codes is now a test, since this drifted twice.
  • OPEN-FINDINGS R9-2 still said fuzz_tag_expr “sits in neither fuzz loop” three sections after recording that it is in both.

[0.11.1] - 2026-08-12 (the gaps 0.11.0 shipped with)

Fixed

  • An output path creates the directories it names. --junit, --sarif and report -o failed when the parent directory did not exist, while artifacts -o and the run directory created theirs — no rule, four sites deciding separately, with the two used most in CI on the failing side. Every adopter paid the same mkdir -p. pytest --junitxml, jest-junit, cargo-nextest’s JUnit store and the hurl proef embeds all create them. This does not weaken the “side effects should be explicit” principle: that is about writing files the user did not name, and here they named exactly this path.

  • proef fragments counts [run] setup/teardown usage. A fragment only a phase feature reached was reported UNREACHABLE — no macro refs it, which was false, and failed --check — a false CI failure in the workflow --check exists for. The verdict also depended on where the phase file sat: inside the suite directory it was discovered as an ordinary feature and counted. The listing’s universe now matches the runner’s, and a phase that fails to load withholds every count rather than guessing. Filed as R10-2.

  • One predicate answers “is this a fragment file?” (FragmentSupport::claims). Three answered it before — CLI discovery via Path::extension, the core scan via rsplit('.'), and the LSP’s corpus invalidation case-insensitively — so they disagreed about api.HURL (the editor rebuilt its corpus for a file nothing would scan) and about a dotfile named .hurl. Filed as R10-3.

  • --config reaches proef lsp and --watch. The flag bypasses the upward search so a proef.toml beside the suite becomes usable — but the editor re-discovered its own config and --watch watched whatever a fresh search found. So in exactly the layout the flag exists for, proef test --config … ran green while the editor reported every ref: as unknown, and editing the config driving the run never retriggered it. ProjectConfig now keeps the file it was read from (with root derived from it rather than stored beside it), and both consumers use the config actually in force. For proef lsp the flag also outranks the client-announced workspace root. Filed as R10-1.

[0.11.0] - 2026-08-12 (the adoption response)

Added

  • [run] exclusive-tags — a tag expression selecting scenarios that run with the pool to themselves. Real suites contain scenarios that cannot run beside anything: one asserting absolute positions (items[0]) needs a store no concurrent scenario writes to, and the only workaround was several CLI invocations driven by tag discipline in a Makefile, each producing its own run record, JUnit file and exit code to aggregate in shell.

    A matching scenario waits for the pool to drain, runs alone, and the pool refills after it, with discovery order unchanged so an exclusive scenario never loses its place. Queueing is strict FIFO, so nothing new starts while one waits at the head — the throughput dip around each exclusive scenario is the price, and it is bounded. A config expression rather than a reserved tag name, because with a bare convention a scenario added months later lands untagged in the parallel pool and breaks isolation intermittently — which reads as flakiness rather than as a missing declaration. A malformed expression is a user error, never a silently-ignored key.

    This is exclusion, not ordering: a scenario that must run before the rest belongs in [run] setup, which already runs once before the pool exists. Deliberately one axis of the two cargo-nextest settled on — per-group concurrency limits (rate-limiting a shared dependency) are a real future need that nobody has asked for, and a group table can be added later without breaking this key.

  • proef fragments — the corpus listing, symmetric with macros. Until now no proef output stated how many fragments there were, so neither way a fragment can die had a denominator to be noticed against: one no macro references was unobservable, and one reached only through a macro no scenario binds looked covered because the macro was flagged. Both are now named apart, unannotated entries are listed by line (they have no name to list by), and --check exits 1 when something never runs. --require-annotated extends that to unannotated entries and is deliberately opt-in: an unannotated entry is inert by design (ADR-0018), so “not done yet” is a porting team’s meaning, not every adopter’s. Reachability is read off the lowered scenarios, so a fragment reached through a chain of use: counts as reached.

  • --config <path>, global to every subcommand, naming the proef.toml to read instead of searching up from the working directory. Discovery only goes up, so a config beside the suite is unreachable from the repository root — a layout an adopting team planned and abandoned after it failed. A named file that does not exist is a user error rather than a fall back to defaults: discovery finding nothing means “no project here”, but a named path that is not there is a typo, and a silently unconfigured run is what that used to buy.

  • proef doctor sees the fragment corpus — a row reporting how many fragments loaded from [run] fragments, warning when the configured root is not a directory. A misconfigured path used to surface much later as pack::unknown_ref: an error about a name when the cause is a path.

  • proef init scaffolds both body forms — a one-entry .hurl file with a # @proef annotation, [run] fragments, and a pack macro of each kind. The newcomer with most to gain from ref: is the one who already owns a hurl corpus, and a scaffold teaching only hurl: | reads as “proef wants your files transcribed into YAML”.

Fixed

  • A bind: key nothing reads is refused (proef::pack::unread_bind_key), with did-you-mean over the names actually in scope. bind_without_ref only caught a table with no ref: at all, so bind: { token: …, toekn: … } validated clean — the one authoring mistake in the fragment path that produced no signal whatsoever. Checked as a union over the scope, never against one fragment: a pack-scope table is the plumbing every macro in the file needs, so a key serving one macro and not its siblings stays correct.

  • duplicate_fragment no longer says “in both x and x” for two entries in one file, and stops offering file.hurl#name as the remedy there — that qualifies by file and cannot separate two entries inside one. Annotating a corpus adds many names to few files, which makes same-file the likely collision.

  • unbound_placeholder names all three supply routes. The omitted one was the fragment’s own [Options] variable: — the route that makes a corpus file runnable standalone, which is the property ADR-0018 exists to preserve.

  • A fragment’s [Options] escaped the ADR-0007 value caps. retry: -1, repeat: -1 and an unbounded delay: were rejected in an inline hurl: block and accepted in a ref: fragment — byte-identical text, exit 2 one way and “dry-run OK, 0 warning(s)” the other, then written verbatim into the executed input. The scan lived inside the inline-only linter; only the twinned-option half of pass 6 had crossed to fragments. It reads the text alone, so it now runs against a fragment’s too, anchored on the ref: line and naming the fragment file and line. This is the case the caps exist for: hurl has no cancellation, so an infinite retry makes the batch budget unestimatable and leaves the watchdog abandoning a thread it cannot stop.

  • A step declaring both ref: and a payload was told, falsely, that its pack had no ref: at all. The conflicted step is reported and dropped, so the loaded bodies stop showing every ref: the author wrote — and the pack-scope bind_without_ref check then drew a conclusion from the gap. It now infers nothing from a pack whose steps did not all normalize.

  • A pack-scope bind: with no ref: anywhere was silently dropped. AUTHORING.md said bind_without_ref applies “at every scope” while only the macro and step scopes were checked — and a setting ignored in silence is the bug those two exist to refuse. The check was the better half of the disagreement, so the pack scope now has it too.

  • A multi-line bind: value blamed the artifact. A hurl [Options] variable: value is a single-line scalar, so a newline could never reach the entry — but it surfaced one stage later as emit::invalid_artifact, pointing at generated text the author never wrote. Refused by name at lower time as lower::multiline_bind, naming the inline hurl: | form that is what splices a multi-line body (ADR-0018’s splicing-versus-binding boundary, enforced where it can be explained).

Changed

  • Breaking (library): AnalyzeCtx takes the fragment corpus instead of building one. Building it internally meant a fresh scan memo per call, so the LSP re-read and re-hurl-parsed the whole corpus on every request — each completion popup, each go-to-definition, each debounce tick. The server now holds one and rebuilds it only when a fragment file changes; editing a pack or a feature, which is nearly every keystroke, leaves it alone. It is also what core purity already required: the caller does the IO.

  • Breaking (library): StepKindSpec gained options, an engine-contributed recogniser mapping a raw option key to what ADR-0007’s budget rules should make of it. The fragment half of that rule already crossed the seam while the inline half matched "retry-interval:" as a literal inside proef-core — one rule at two altitudes, and a second engine would have had its fragments linted and its inline blocks not. A kind contributing no recogniser is not linted, since the core has no way to know what its option keys mean.

  • Breaking (library): proef_core::engine::FragmentScanner returns ScannedFile { fragments, unannotated } rather than Vec<ScannedFragment>. An engine’s scanner now also reports the 1-based lines of entries carrying no annotation — lines only, never built-then-discarded fragments, so a foreign corpus still costs a push per unannotated entry. Without it “which entries did I forget to annotate?” is unanswerable: a missing annotation produces a green run and a silently absent test, and the entry that would prove it was never built. FragmentCorpus gains fragments(), unannotated() and diagnostics(), because the scan is gated on some pack naming a fragment — so PackSet::fragments is empty for exactly the suite a listing has most to say about.

Documentation

  • Config discovery is a requirement, not a convention. proef.toml is found by searching up from the working directory, so a config beside the suite (tests/proef/proef.toml) is never found from the repository root — an adopting team planned that layout and discovered it by failure. CONFIG.md now says so, and notes that keeping the file at the root collapses the one place [run] fragments (config-relative) and suite/setup/teardown/runs-dir (cwd-relative) differ.

  • The release runbook could not work as written. main is a protected branch, and step 4’s git push origin main --follow-tags fails in the dangerous direction: --follow-tags is not atomic, so the branch is rejected while the tag still lands — and the tag is what release.yml triggers on, starting a release build from a commit that is not on main. It happened cutting 0.10.0. The runbook now routes the release commit through a PR and tags the merged commit, and the cargo publish section carries the dry-run, tag-check and --locked sequence plus why only four crates go ([workspace.package] publish = false is the default). Also drops step 1’s reference to changelog “bottom links”, which do not exist.

[0.10.0] - 2026-08-12 (named hurl fragments)

Breaking (library): proef_core::pack::load takes a &proef_core::pack::FragmentCorpus between the packs and the step kinds (&FragmentCorpus::empty() for the previous behaviour, or FragmentCorpus::new(sources, kinds) to supply fragment files), and PackSet::fragments is an Arc<BTreeMap<…>> so one scan can be shared by every load; LoweredScenario::secrets is a BTreeMap<String, String> of engine-variable → secret name rather than a BTreeSet<String>; Prepared and ScenarioCtx each gain a secret_bindings field carrying that map to the engine; and SourceProvider::discover_fragments is a required method (return Ok(Vec::new()) to serve none) — it was briefly defaulted, and the default silently disabled fragments for a provider that forwarded the other two; and ScannedFragment::name is a String rather than Option<String>, because a scanner now reports only the entries it found an annotation on; and ScannedFragment and pack::Fragment each gain a supplied_variables: Vec<String> (Vec::new() for none), which an engine’s scanner must fill from the entry’s [Options] variable: lines — leaving it empty reinstates the silent last-wins it exists to refuse; and both LoweredStep, StepOutcome and Event::StepFinished gain a fragment: Option<String> field and analyze::FragmentDef gains placeholders: Vec<String>, so a literal construction of any of them needs one more line (None / Vec::new() reproduces the previous behaviour). The wire schema is unaffected — the event field is skipped when absent, which is what keeps existing records byte-equal.

Added

  • The docs are checked mechanically, not only read. xtask docs-check gained two passes — every relative link resolves, and every fenced toml/yaml example parses with the product’s own parsers, so the check means “proef would accept this example” rather than “some parser would”. A third pass, whether a documented command or long flag actually exists, needs a built binary and so lives in crates/proef-cli/tests/docs.rs.

    All three were written against defects already in the tree: ADR-0018’s first example could not load (an unquoted ${…} inside a YAML flow mapping, where { opens a nested mapping), and a row marked shipped documented proef report --html, a flag that never existed. Both had correct prose around wrong code — the failure mode review does not catch.

  • Packs can name fragments: ref: and bind: (ADR-0018). A macro step’s body may be ref: <fragment> instead of an inline hurl: block, and bind: supplies the fragment’s {{…}} variables at pack, macro and step scope, most specific winning. Fragment names are global, and file.hurl#name qualifies one — the same two spellings, resolved the same way, that use: already accepts.

    Refused at load, each with its own code: a ref: naming no loaded fragment (unknown_ref, suggesting the closest, and saying so plainly when no fragment file was loaded rather than implying a typo); two files declaring one name (duplicate_fragment); a file the engine cannot read (bad_annotation — its siblings still load); a step that is both ref: and a payload (body_form_conflict); and bind: on a step with no ref: (bind_without_ref — an inline block takes ${…}, so that binding would feed nothing, and a setting silently ignored is the bug this refuses to ship).

    A fragment declaring its own retry alongside a step’s retry: is the same option_declared_twice an inline block gets, so the two body forms behave identically rather than differing by where the hurl text happens to live.

    A fragment may also supply a variable to itself with an ordinary [Options] variable: line — that is how a corpus file stays runnable on its own, so it counts as an answer to that fragment’s own {{…}} and needs no bind:. Supplying and binding the same name is refused (option_declared_twice): both reach the entry as variable: k=, hurl takes the last, and the fragment’s own line is last — so the bound value would silently never be sent, and would stay unsent for every later entry, since hurl’s variable: assigns into the run-level set rather than scoping.

    Discovery arrives below, so a ref: resolves end to end.

  • [run] fragments — the hurl files a pack may ref:. Names one root, scanned recursively for the extensions the registered engines claim, so discovery never learns a file type of its own. Unset means no fragments: there is no convention fallback, because unknown_ref saying “no fragment files were loaded” beats guessing at a directory.

    Relative paths resolve against proef.toml’s own directory, not the working directory. The config is found by walking up from the cwd, so a path in a config three levels above must mean “relative to the project” — otherwise proef flows from a subdirectory reads the right config and then cannot find anything it names. [run] suite predates this and stays cwd-relative; it is only consulted when no path was given, so the difference is not observable there.

    The LSP resolves fragments through the same root, so ref: does not read as unknown in an editor while the suite runs green. --watch retriggers on .hurl edits and watches the fragment root separately, since a corpus may live outside the suite. proef fmt still refuses .hurl in both discovery branches — it locates hurl blocks inside YAML, and a corpus proef did not write is not proef’s to rewrite — now pinned by a test.

  • Fragments lower, bind, and execute. A ref: step emits the fragment’s own text with its non-secret bindings baked in as per-entry [Options] variable: lines, so the artifact stays the executed input and replays identically under the stock CLI (ADR-0010). Values are always quoted: variable_value tries null/bool/number before string, so an unquoted records, 2 and true would become three different types by accident.

    Two refusals guard the parts that could otherwise pass silently:

    • lower::unbound_placeholder — a fragment reading a {{variable}} that no bind: in scope supplies and no earlier step captures, anchored on the .hurl line the variable is on rather than on the pack. hurl’s [Options] variable: assigns into one shared set rather than scoping, so an unbound name would inherit whatever a previous entry happened to leave and run green against the wrong value.
    • lower::secret_in_composite_bind — a bind: value mixing ${secret:…} into a larger string. To inject that, the composite would have to be materialized into the artifact, which ADR-0005 forbids; bind the secret alone and let the fragment spell the surrounding text.

    Secrets keep their own path: recorded as engine-variable → secret name and injected via insert_secret at run time, never as an [Options] line. That indirection is what lets bind: { auth_token: "${secret:apiToken}" } give a secret the variable name a corpus proef did not write already uses.

    Bindings resolve once per scope instantiation — pack scope once per scenario, macro scope once per invocation, step scope per step — so one binding is one value and two bindings are two. A macro with no ref: step resolves nothing, so an unused table never advances the ${fake:…} counter.

  • The engine seam can describe fragment files (ADR-0018, groundwork). StepKindSpec gains fragments: Option<FragmentSupport>, and proef-core gains ScannedFragment / FragmentScanError / FragmentScanner. The hurl engine implements the scanner over hurl’s own AST: the # @proef <name> annotation is read from the entry’s line_terminators, so the annotation↔entry binding is exactly as reliable as hurl’s parser and no text is scanned for structure. An entry’s required inputs and produced captures are read from the same AST, which is what will let an unbound placeholder be an error rather than a runtime surprise.

    Additive only — nothing was removed from proef-core’s surface, and no hurl type appears anywhere in it. Discovery asks the registry for the extension instead of naming .hurl itself, so this stays ADR-0002’s “adding an engine leaves proef-core diff-empty” rather than an exception to it. Nothing observable ships yet: no pack can reference a fragment until the schema lands.

    StepKindSpec::fragments is one Option<FragmentSupport> rather than a separate extension and scanner, so a kind that claims a format it cannot read is not expressible; a file no kind claims is skipped rather than handed to whichever engine happens to be registered first. ScannedFragment::declared_options lists option families rather than flagging retry alone, so the core applies its double-declaration rule to delay: too — through the same bake_entry_options path, so leaving it out reproduced the very last-wins bug the rule exists to refuse. supplied_variables is separate from it because the two clash on different keys: an option family family-to-family, a variable name-to-name.

    A note for whoever extends the scanner: hurl’s Visitor treats templates as leaves, and visit_template, visit_url and visit_filename are three separate no-op defaults that do not forward to one another. Overriding only visit_template silently under-reports an entry’s inputs — and a missing input reads as “needs no binding”.

  • A run record says which fragment a step ran, and explain prints it. step_finished gains a fragment field carrying file.hurl#name (additive per ADR-0008: absent for an inline hurl: block, so no pre-existing record changes a byte — the reference event-stream snapshot is unmoved), and proef explain renders it under a failure as via tests/hurl/admin.hurl#admin.search. A step that never ran reports it too: “not run” is exactly when someone is reconstructing what the suite was about to do.

    This closes a promise ADR-0018 made rather than adding a new one — three files per test was accepted on the condition that explain and go-to-definition earn it back, and only go-to-definition had. The name is qualified at lowering rather than by the reader, because a record has to stand alone: by the time it is read, the pack that named the fragment may say something else.

    JUnit, the GitHub job summary and the ::error annotations name it too, as a trailing (via file.hurl#name) on the failure message, and the HTML report renders it under the reason. CI is where a reader is least able to go looking for themselves, so it is the last place provenance should drop out — and all three sinks share one helper rather than a format string each, because three copies is how one of them quietly stops agreeing with the run record.

  • bind: completes against what the fragment actually reads. With the cursor in a bind: table — flow or block style — the editor offers the {{variables}} of the fragments that pack ref:s, nearest ref: ranked first, each labelled with the fragment that wants it. The names come off the engine’s own AST at scan time (analyze::FragmentDef::placeholders), so this is the file’s real interface rather than a second description that could disagree with it.

    Until now the only route to a foreign corpus’s variable names was to run the suite and read proef::lower::unbound_placeholder — a lower-time error, so the names arrived only after a failure. bind: exists at three scopes and only the step one names a single fragment unambiguously, so the list is a union rather than a guess; the owning fragment rides in each item’s detail.

  • The fragment corpus is scanned once per command, not once per pack load. A proef test loads packs up to four times — the suite, then [run] setup and [run] teardown, each validated and then run — against different feature paths but always the same corpus, and each load re-read and re-parsed every .hurl file. Measured on a 200-file / 15k-line corpus: 140 ms → 40 ms warm, with pack loading falling from ~28% of the run to a single pass. The win scales with the corpus, which is the direction adoption goes.

    The corpus is now read once per invocation (front::fragment_corpus) into a FragmentCorpus that scans itself lazily, at most once. Laziness is the part worth guarding: load_collecting still scans only when some pack actually has a ref:, which is what makes CONFIG.md’s “pointing at a corpus you did not write costs nothing” true. Hoisting the scan to the caller to share it would have bought the speed by breaking that promise, so the memo lives with the corpus instead — and a test proves the eager version fails, by pointing an unreferenced corpus at a file that cannot parse and asserting no diagnostic appears.

    Built per invocation rather than in a static: --watch re-enters the same process after each edit, and a corpus outliving one run would serve pre-edit fragments to the next.

  • Go-to-definition on a ref: worked again, then briefly did not. Shortening the [run] fragments root to a cwd-relative spelling — done so a run record would not carry an absolute, machine-specific path — also shortened the root proef lsp hands to its source provider. The LSP keys document identity on absolute names (name_to_url yields None for anything relative), so every ref: go-to-definition returned null and .hurl-positioned diagnostics stopped publishing, while the suite still ran green. That is the capability restored two commits earlier.

    Resolution and spelling are now separate concerns: ProjectConfig::fragments() returns a resolvable path, and the shortening happens at the naming boundary in front::fragment_sources, which only CLI runs pass through. Both properties hold at once — the editor resolves, the record stays portable.

    Covered by an end-to-end proef lsp stdio test with a real proef.toml, the seam the unit tests could not reach: they inject absolute names through a fake provider, so they never exercise config → provider → URI. The test canonicalizes its temp root deliberately — on macOS a tempdir is /var/… whose real path is /private/var/…, and without that the cwd comparison silently no-ops and the test passes vacuously.

  • Every failure sink names the fragment, not just the CI ones. via() moved from ci_reports to render, and the console failure list and TAP diagnostic now carry it too. A helper scoped to one delivery channel was how proef test printed no provenance on stderr while report.junit.xml from that same run printed it — the drift the helper’s own comment says it exists to prevent.

Internal

  • The secret-name join has one home. proef_core::engine::secret_variables pairs a scenario’s secret_bindings (variable → secret name) with its secrets (name → value) and is the only place that join is written. Doing it engine-side invited injecting under the secret name, which makes a renamed binding (ADR-0018) resolve to nothing — the request then leaves with an unresolved {{…}} and fails far from the cause. It yields borrows on purpose: an owned variable → value map would put a second copy of every secret value in memory per scenario, and ADR-0005 keeps values in one place.

  • engine::OPTION_FAMILIES names the vocabulary the double-declaration check compares against, and MacroStep::declared_options derives the other half of that comparison once for both body forms. The two sides were previously hardcoded lists that met by string equality with no test spanning the crates — a spelling only the engine knew would have matched nothing and quietly disabled option_declared_twice, reinstating the hurl last-wins it exists to refuse. A proef-engine-hurl test now asserts every family the real scanner emits is one the pack can declare; delay was untested there entirely.

  • Lowering’s two diagnostic sinks are one Sinks value. They were adjacent parameters of the same type threaded through seven functions and a closure: transposing them at any of a dozen call sites compiled cleanly and routed every error into warnings, so a scenario that should have failed lowered “successfully” and the run exited 0. No &mut Vec<Diag> parameter remains in lower.rs, which makes the mistake unspellable rather than merely unmade.

Documentation

  • AUTHORING says which body form to reach for, and why. A table contrasting splicing against binding — what each can substitute, whether it can be reused, whether stock hurl can run it, and when an unknown variable is caught — plus the rule that decides it: inline when you need to splice something hurl cannot template (${docstring} as a body has no binding equivalent), ref: when the request is shared, foreign, or must stand alone. CONFIG.md gains [run] fragments with a worked three-file example.

  • The hurl non-goal is about generation, not direction (PRD §3 amendment). It read “importing/round-tripping hand-written hurl files into Gherkin (artifacts flow outward only)” — a clause and a parenthetical saying two different things, the parenthetical forbidding hurl text from being an input at all. What the non-goal protects is that proef never authors a test for you, and that reasoning is untouched (ADR-0016 stays declined on it). It does not extend to hurl being an input source, which §1’s own framing — “there is no tool that joins the two” — describes as the product’s purpose. Recorded honestly: OPEN-FINDINGS M3 asked for this re-examination to arrive with a measured port cost, and it has not.

  • ADR-0018 — named hurl fragments. A macro step’s body may be ref: <fragment> naming one entry in a real .hurl file, annotated # @proef <name>, with proef values supplied by an explicit bind: map instead of ${…} splicing. The file stays valid hurl, so the same file runs under proef test and under stock hurl. Inline hurl: | is unchanged and stays: the two are splicing versus binding, with different capability envelopes, and the 844-line corpus port is recorded in the ADR as evidence the inline path is sufficient for real work. No behaviour ships with this entry — the ADR and the charter amendment land first, deliberately.

Fixed

  • --watch reran itself forever. ADR-0018 added the engines’ fragment extensions to the retrigger allowlist — .hurl among them — while every run writes .proef-runs/<id>/artifacts/*.hurl. A watched tree containing its own runs dir fed itself: 49 runs in 15 seconds, firing real traffic in a tight loop and churning record rotation. The filter now excludes generated trees by directory name, reusing discovery’s own skipped_dir so there is one rule with two consumers, and takes [run] runs-dir for the case where it is not a dot-directory. OPEN-FINDINGS P5 had closed this “by inspection”, naming .hurl as a file that could never match; the note is corrected in place.

  • One unreadable file sank the whole corpus. A fragment root is foreign by design, but a single binary or latin-1 file in it exited 3 from every command — flows included, which never looks at a fragment. Read failures are now per-file diagnostics (pack::unreadable_fragment_file) that never sink their siblings and stay silent until something ref:s the corpus, matching what pack loading and the annotation scan already did.

  • schema --add-to rewrote fragment files. It prepended a yaml-language-server modeline to a .hurl corpus file and dropped the pack schema beside it — violating ADR-0018’s “fragment files are inputs proef never writes”. It now refuses anything that is not a pack, reusing the is_pack_file predicate fmt already had.

  • A # in an annotation name was accepted but unreachable. # separates a file from a fragment in ref: file.hurl#name, so such a name could be declared and never referenced — and the failure suggested the exact spelling that had just failed. Refused at scan time.

  • proef lsp answered every URI-keyed request with null on Windows. A source name is an identity compared as a string, and the two sides spelled it differently: Path::join appends without rewriting what is already there, so a proef.toml saying suite = "tests/features" — the portable spelling the docs use — produced C:\proj\tests/features\packs\api.yaml from discovery while the client’s document URI produced C:\proj\tests\features\packs\api.yaml. The two never matched, so go-to-definition, find-references and completion all found nothing while the suite itself ran green. Discovered names are now rebuilt in native form. Unix has one separator and was never affected, which is why every gate stayed green.

  • A fragment’s path was absolute everywhere it was named. [run] fragments resolves against the config file’s directory, so fragments = "tests/hurl" became /home/you/project/tests/hurl — and that spelling then named the file in every diagnostic and, once steps recorded their provenance, in the run record too. Feature and pack names are project-relative because the path the author typed was; a path the author never typed had no such luck. Records went machine-specific: the same suite on two checkouts stopped comparing equal, and a temp-dir path could reach a durable artifact. The root is now shortened back to a cwd-relative spelling when it is under the working directory — resolution is untouched, so which file gets read never changes.

  • Every ref: was an error in the editor while the same suite ran green. SourceProvider::discover_fragments shipped with a default Ok(Vec::new()), and the LSP’s overlay provider — which forwards feature and pack discovery to disk — never overrode it. So the analyzer saw no fragments at all: go-to-definition on a ref: did nothing, ref: completion returned nothing, and every ref: rendered as proef::pack::unknown_ref. Exactly the diagnostics-you-cannot-trust drift the fragment-aware analysis was added to prevent.

    The default is gone; discover_fragments is a required method. Every implementation lives in this workspace, so the default bought no compatibility — it only let a forwarding provider inherit “no fragments” silently instead of failing to compile. An integration test now drives the real provider chain and asserts a ref: jump lands on the annotation in the .hurl file.

  • A fragment file saved with a BOM failed at line 1, blaming the request. Every other text entry point (feature::parse, the inline-payload probe) strips a leading U+FEFF; the fragment scanner did not, so the mark reached hurl’s parser as the first character of the first request. The file is now normalized by the same rule, and the mark cannot travel into an artifact that has to be valid hurl.

  • A macro-scope bind: with no ref: step was silently dropped. The step-scope version of this mistake has been a hard error since bind: landed; one scope up it vanished at lower time. That is the half authors actually hit, because factoring plumbing upward is the habit — and the tempting reading, that a use: target will pick the table up, is wrong: the child resolves its own scopes. Now proef::pack::bind_without_ref at both scopes, with a message that says so.

  • A ref: step’s name: reported a ${fake:…} value it never sent. A label is a replay of what the request was built from, not a fresh use of it: the inline path rewinds the ${fake:…} occurrence counter, resolves the label, then restores it to the high-water mark. The ref: path reproduced that tail without the rewind, so a step binding ${fake:email} and naming ${fake:email} minted two identities — the console and the event stream announced one address while the request sent another, and every later step’s fake values shifted by one. Both body forms now end in one shared finish_step, so the rule is stated and enforced in a single place rather than copied.

  • An escaped $${secret:…} in a bind: value was refused as a composite. $${ is the escape (ADR-0005), so $${secret:token} is the literal text ${secret:token} and names no secret — but the composite check searched for the substring "${secret:", matched at offset 1, and rejected the binding with secret_in_composite_bind. Both the whole-value and composite tests now read the value through the resolver’s own reference scanner, so there is one thing that knows what a ${…} is and $${ stays an escape everywhere.

  • A step that set retry: twice ran the value it did not name. A pack could declare retry: (or delay:) as a step key and again inside the block’s own [Options]. Lowering extends an author’s existing section rather than opening a second one, so proef’s baked line landed above the author’s; hurl resolves a duplicated option last-wins, and the raw value therefore won every time. The pack said retry: 10, the run did retry: 3, and nothing anywhere said so — the finite-retry lint only ever looked for -1 and over-cap counts, so a plausible finite value passed untouched. Declaring an option in both places is now proef::pack::option_declared_twice, refused at load with the span on the raw line that used to take effect.

    The scan is deliberately scoped to [Options] sections rather than matching any retry:-shaped line: retry is a legal request-header name, and a header is name: value like an option is, so a line-shaped match would have turned an ordinary header into a hard error. Pinned by a test that a header named retry on a step carrying a typed retry: still loads.

[0.9.0] - 2026-08-11 (tool-surface integrity & authoring guidance)

Breaking: proef secret set --value was removed in favour of --stdin, and proef macros --output json’s pattern field changed from a boolean to string|null.

Added

  • The run record says which scenarios were lifecycle phases. phase ("setup"/"teardown") is now on scenario_started/scenario_finished — additive and optional (ADR-0008), so older records read as “no phases”, which is what they had. Without it a teardown scenario was indistinguishable from a suite one except by feature path, so every consumer re-derived phase membership from proef.toml and three of them got it wrong in different ways. Fixing them off one signal is what the three entries below have in common.

  • proef doctor reports a missing pack schema. init installs it automatically, but noticing when it is absent never shipped — so a suite whose editor completion had been silently off had nothing telling it so. Reported as a warning, never a failure: it costs autocomplete and load-time validation in the editor, not a run, and doctor’s exit is the environment verdict. Uses the same predicate init uses, so the two cannot disagree about what “installed” means. Runs outside a project too — no config or no suite is reported, not failed.

  • bind::unbound_step names proef macros again, from the CLI. The pointer was removed from the diagnostic in #25 for a correct reason — that text also renders in an editor’s diagnostics pane through the LSP, where the affordance is completion, not a command — but nothing put it back on the terminal side, so a terminal reader saw it zero times. It is now added by the CLI’s own renderer, which legitimately knows it is the CLI. The core diagnostic still names no tool.

  • proef macros answers when the suite does not bind. Listing the vocabulary previously required every scenario to bind — so the command refused in exactly the situation that sends an author looking for it: a step that matched no macro. It now prints the diagnostics, then the vocabulary the packs offer, and keeps its exit code unchanged (2), so scripts see no difference. Pack loading precedes binding and does not depend on it, so the listed vocabulary is complete. Every count-derived verdict is withheld in that mode — calls/unused render as —/null rather than 0/false, because a feature that failed to bind contributes no calls and would otherwise make its own macros look dead. proef flows deliberately still refuses: its contract is to list every scenario, and a partial list that silently omits the unparsed feature is the wrong answer, not a degraded one.

  • A failed run says when the suite is still the untouched scaffold. A freshly scaffolded project cannot pass — its target and its routes are both placeholders — and init says so once, two commands earlier, in a parenthetical the failure never referred back to. The run now names the situation and the remedy. It fires only on the conjunction ([url] base still byte-identical to what init wrote and no PROEF_BASE_URL): an operator who set the override did name a target, so their failure is about their API and is not second-guessed. Exit codes are untouched — whether an unreachable target is a user or a system fault is a taxonomy question decided in the engine (ADR-0009), and the reader’s actual problem is vocabulary.

Changed

  • proef secret set --value is gone; use --stdin. Breaking. A secret in argv is visible to anyone who can run ps, and the failure path steered people to it — the hidden prompt’s error said “pass --value in scripts”, which fires exactly in the non-TTY/CI case where the exposure matters. There is now no flag that takes a value: --stdin reads it from a pipe (same shape as docker login --password-stdin), stripping the trailing newline the pipe added, and the prompt stays the default. Scripts using --value must pipe instead: printf %s "$TOKEN" | proef secret set NAME --stdin.

  • proef macros prints the sentence, not just the identifier. A test author writes prose that binds to a vocabulary somebody else maintains — and the one command that lists that vocabulary showed health where the author needs the service is healthy. The match: pattern was already loaded and already linted; both renderers discarded it on the way out. It now appears in the text listing, and --output json’s pattern field carries the string itself (null when a macro is use:-only) instead of a bare boolean.

Fixed

  • proef fmt refuses a file that is not a pack. It took an explicit path on trust, so it rewrote whatever it was pointed at: proef fmt src/main.rs stripped trailing whitespace from Rust source, printed formatted:, and exited 0. A mistyped path was a silent edit. Formatters parse before they write and refuse what they cannot parse; this one locates blocks textually, so the extension is the check available — and it is now the same predicate discovery already used, rather than a second opinion about what a pack is. Only the explicit-file path was affected: a directory was always filtered.

  • proef fmt leaves the YAML skeleton alone, as it always said it did. Its documented scope is hurl blocks — the module doc promises the skeleton, comments included, is never touched, and the code claimed the trailing newline was the only normalization applied outside a block. Both were wrong: every line was trimmed. A pack whose blocks were already canonical failed fmt --check on nothing but a trailing space in a comment, which is a CI red an author cannot explain from the documented scope. This is the same over-reach the line-ending fix removed in 0.8.0, in the same function, one line above where that fix landed.

  • A truncated record no longer drops a warned scenario from its totals. With no run_finished to read, explain recounts the scenarios present — and counted Passed/Failed/Skipped but not Warned, so a scenario whose optional: step warned vanished from every column. The live path counts Passed | Warned together (RunSummary::passed is “passed, warnings allowed”), so the reconstruction silently disagreed with the run it was reconstructing — and optional: exists precisely so a scenario can warn and still pass.

  • A failing run says when the scaffold’s routes are still placeholders. The scaffold has two halves to fill in, and a reader can have done either. Someone who follows init’s instruction — point ${url:base} at your API — then hits the other half: /health and /search 404, and the target-side note deliberately cannot fire, because they did configure a target. They had been told about the routes once, parenthetically, two commands earlier. Now they are told at the failure. Decided from the pack’s bytes, never from what the server answered: a 404 proves a route is missing, not that it is a placeholder, and inferring the second from the first is the class of claim removed in 0.8.0. The two notes are mutually exclusive — a reader with one unfinished half is told about that half, not handed a list.

  • --dry-run’s “next” command is the run that was validated. After --dry-run --env prod --tags smoke it printed a bare proef test, which is a different run — another [url] base from the profile, and every scenario rather than the tagged subset. The operator could not tell: the command works and simply tests something else. Every selector that chose what ran is echoed now (--env, --tags, --scenario, --scenario-file, and the path), quoted so a tag expression or a scenario name with spaces survives a paste. Deliberately selectors only — a general “reprint the invocation” is how secret-bearing arguments reach stdout.

  • --sarif emits startLine. GitHub keys inline annotations on it, so a log carrying only byteOffset/byteLength uploaded cleanly and annotated nothing — the flag looked wired up and delivered none of what it advertises. Sources are read once each at the IO edge and only to count newlines; Diag keeps carrying byte spans, and no column arithmetic is introduced.

  • --watch retriggers on proef.toml. It watched the suite path recursively, and the config lives above it — so editing a [url]/[vars]/ [env.*] value that every scenario resolves through changed nothing, which reads as the watcher being broken. Matched by exact path rather than by a .toml extension, so an unrelated manifest in the tree still does not requeue.

  • Three places interpolated a value into a format without escaping it. Same shape each time, so they are fixed together:

    • LSP completion snippets. $, } and \ are LSP snippet syntax, and a match: pattern is prose — prose carries $. the price is $5 made the client read $5 as tabstop 5 and drop the text, so accepting the completion inserted something the author never wrote. Literal characters are escaped now; the tabstops the generator writes stay syntax.
    • GitHub annotations. file= was passed raw while title= and the message beside it in the same writeln! were encoded. A path carrying , or : — every Windows path carries a : — broke the key=value,key=value parse.
    • The GitHub job-summary table. The scenario name and file went into Markdown cells unescaped; a | in either ends the cell and shifts every column after it, and the row still renders, which is why it goes unnoticed.
  • A templated retry:/delay:/repeat:/max-time: no longer under-counts the batch budget. The estimator matched literal values only, so a {{var}}-driven option fell through and read as no retries — the budget was then computed for a single attempt, and the watchdog abandoned a scenario that was retrying exactly as authored, reporting it as an environment fault (exit 3). A placeholder resolves inside hurl at run time and cannot be estimated, so the engine now says so: batch_budget returns None, whose contract already routes the batch to the orchestrator’s default budget. An infinite count is treated the same way, since it is unbounded by definition. TROUBLESHOOTING described the old behaviour as if the budget could see these values; it now says what actually happens.

  • --output json’s exit_code is the code the process exits with. A failed JUnit write escalates the run to 3, and that escalation was applied by a return after the body had been printed — so a machine consumer read a verdict the program then exited past, with nothing to signal the disagreement. The escalation is now folded in before anything serializes it.

  • proef fmt keeps each line’s own ending. Its scope is hurl blocks, not line endings, but it split the whole file with str::lines() — which throws the terminator away — and rejoined with a single one. A file mixing CRLF and LF was therefore homogenized, and fmt --check came back red on a pack whose blocks were already canonical. The earlier fix moved from “always LF” to “the dominant ending”, which still rewrote the minority lines. Terminators now travel with their line, so an untouched line is written back byte-for-byte; the only ending fmt still supplies is a trailing newline on a file that lacked one.

  • proef --help describes macros as it now behaves. It still said “with its call count” after the command started printing the sentence each macro binds — the README table was updated and the clap text that actually produces --help was not.

  • proef lsp adopts the workspace root the client announces. The root was resolved at the process edge, before the handshake, from the working directory — so an editor launched anywhere but the project analysed the wrong tree, and nvim ~/proj/x.feature from $HOME rooted the analyser at $HOME. The initialize params were bound and discarded. The server now reads workspaceFolders, falling back to rootUri (deprecated since LSP 3.16, and the spec is explicit that folders win when both are present) and then to the previous config-then-cwd resolution. proef-lsp still knows nothing about proef.toml: it calls back into the CLI, which owns config (ADR-0012).

  • A mixed suite+phase failure kept the phase label. explain chose the label from the whole report (failed == 0), so it appeared only while every failure was a phase failure — and vanished the moment a suite failure joined one, leaving 1 failed above two indistinguishable blocks. The disambiguation disappeared exactly where it was needed. Labelled per block now, from the record.

  • --rerun after a phase-only failure says there is nothing to rerun. It returned the failed teardown, which build_specs cannot match because the phase is excluded from the pool — producing a run that matched nothing and reported “no scenarios matched the filters (check –tags/–scenario)”, naming flags the operator never passed. Phases are invisible to --rerun (ADR-0014); it now exits 0 saying so.

  • diff no longer counts a failing teardown as a test regression. A cleanup fault makes test exit 3, not 1, so blending phases into the regression buckets made diff --fail-on-regression contradict the run it was diffing. Phase scenarios are excluded from the verdict and the exclusion is reported.

  • Records written before 0.6.0 no longer report the wrong verdict with confidence. They carry one run_finished per phase and their totals counted every phase; read under today’s suite-only meaning, a genuine suite failure was reported as 1 passed · 0 failed and labelled setup/teardown. The schema field cannot distinguish them — that change was semantic and never bumped it — but the structure can. explain now detects the multiple pairs, recomputes the totals from the scenarios present, and says the record predates 0.6.0. A reader must be able to consume a record or detect that it cannot; quietly doing neither was the one unacceptable option.

  • proef init no longer destroys a proef-pack.schema.json you wrote. The never-overwrite loop walks a fixed four-entry array; the schema is not in it, and is written afterwards by the shared installer. So the one unguarded path was pack-absent + schema-present: init scaffolded the pack, then the installer replaced an authored file — reported as created 5 file(s), skipped 0, while the README promised the opposite in as many words. init now asks the installer to preserve what is already there and reports it as skipped; proef schema --add-to still refreshes, since that is an explicit install and how the schema is updated after upgrading proef.

  • The first-run note no longer fires on real suites. It keyed on [url] base still equalling the value proef init writes — which looks init-specific and is not: GETTING-STARTED teaches that exact line to people building a suite by hand, and proef’s own proef.toml uses it. So a hand-built suite whose server was up and whose assertion genuinely failed was told “this suite is still the proef init scaffold — its target and its routes are placeholders, so it cannot pass yet”: every clause false, moments after the suite reached a real verdict. The deciding evidence is now the run itself — the note appears only when nothing was reachable (no scenario passed and every outcome is a system fault). A suite that got an HTTP response, even a 404, has a target; whether its routes are placeholders was a guess, and the note stated it as fact. Wording softened accordingly.

  • Suite discovery no longer walks build output, and one unreadable directory no longer empties the suite. The walk had no exclusions, no depth bound, and a canonicalize() per directory — and it re-runs on every language-server request, so entering target/ cost that price over and over for a subtree that cannot contain a suite. It now skips target/, node_modules/, vendor/ and dot-directories (tested on children only: a suite may legitimately be rooted at such a name), and refuses beyond 32 levels rather than recursing until the stack runs out. A Permission denied on one descendant used to abort the entire walk, and proef lsp swallowed that error into an empty analysis — so a single unreadable subdirectory silently emptied the suite. Unreadable descendants are now skipped, the way find and ripgrep do; an unreadable root is still a loud error, because that path is the caller’s own.

  • Ctrl-C no longer skips cleanup in silence. Teardown shared the run’s cancellation token, so an interrupt left every teardown scenario Skipped — and because a skipped phase carries no fault, the worst-wins fold passed it without a word. Whatever setup created stayed created and nothing said so, against this ADR’s own premise that suite cleanup is reliable. Teardown now runs on its own, independent token (not child_token(), which cancels with its parent and would have re-implemented the bug): the pool stops at its batch boundary, the operator is told cleanup is running, and it completes. A second Ctrl-C still hard-exits (130) — the escape hatch ADR-0007 relies on — and the announcement says so. Amends ADR-0014.

  • A phase that only skipped is now a failure, not a pass. That silence was the shape that hid cancelled cleanup. A setup completing no scenario aborts the run rather than letting the suite execute against state setup never created — which is also what keeps teardown gated on setup-success, since the abort is the gate; a teardown completing no scenario is reported and fails.

  • --dry-run validates [run] setup and [run] teardown — which ADR-0014 always claimed (“validated like any other feature but never executed”) and nothing did: --dry-run never read the keys. A broken teardown therefore surfaced only after a full suite had run — real requests, a run directory, artifacts — while the identical mistake in setup failed in milliseconds. Both are now validated by one loader shared with proef test, which also pre-flights teardown before the pool. A bad phase path is a user error (exit 2) rather than a blanket system fault (exit 3), and creates no run record.

  • proef schema --add-to and proef init now announce the schema file they write. Both wrote proef-pack.schema.json silently, so init listed four files and then reported “created 5 file(s)” — the first output a new user reads, not reconciling, with the unannounced file being the one that powers editor completion.

  • proef init no longer sends you to install editor completion that is already installed. A re-run named proef schema --add-to unconditionally, even with the schema sitting beside the pack. It now says which of the two situations you are in.

  • The nextest harness no longer reports green having listed no tests. A PROEF_HARNESS_SUITE set to bytes that are not valid UTF-8 read as unset, which the harness treats as “expose nothing” on purpose — so cargo test passed having run zero scenarios. A PROEF_BIN it could not read fell back to proef on PATH, silently invoking a different binary than the one named. Both now surface as a failing proef::config trial, the same loud shape the harness already used for flows-contract drift, whose comment states the invariant this violated: never run zero tests green.

Documentation

  • AUTHORING shows how to write a validation-error catalogue. Two patterns that were reachable but not signposted, and that compose into one. A validation suite’s cases differ structurally — one omits a key, one empties it, one adds a key the caller may not set — so a single parameterised macro cannot express them and an Examples cell cannot practically hold JSON; the answer is one named macro per malformation, whose sentence says what is wrong in business terms. The expectation side then does not grow with the catalogue: because an expect: merges into the previous request entry, one parameterised the error code is {code} covers every case in the set, typically the largest de-duplicator in a validation pack. That merging was documented as a mechanism in two sentences and never shown as the pattern it is. The cost is stated rather than hidden — the pack grows with the catalogue, which is what buys feature files a non-engineer can review.

  • An outline’s <column> placeholders substitute into the docstring, and AUTHORING now says so. They always have — TECH-SPEC §4.4 specifies it and the code has done it since — but the author-facing guide named only step text and table cells, and StepDefn’s own doc comment named the substitution on text and table while describing docstring as just “raw request bodies”. Naming it twice and omitting it once reads as a deliberate exception, so a reader concludes the opposite of the truth: this is exactly the capability an author reaches for to data-drive a request body without leaving the feature file. AUTHORING gains a worked example. Pinned by tests for the first time — every other outline test asserts on step text, so a regression would have emitted a literal <label> into an artifact with the suite green.

  • The docs-drift backlog is closed. EDITORS.md said go-to-definition cannot land on a match: line — it has since 0.5.1, and definition_on_a_step_lands_on_the_match_line proves it; the bullet now names the gap that is real (built-in macros live in a pack compiled into the binary, so there is nothing to open). TECH-SPEC §10’s command surface gained --run-id/--rerun/--sarif. GETTING-STARTED no longer shows a scaffold comment with a word the scaffold does not write. ADR-0015 described a worker on ScenarioFinished that is always None, because that event is emitted from the dispatcher thread rather than the worker — an errata records what shipped, which EVENTS.md had right all along.

    Two entries did not reproduce and are recorded as such rather than dropped: CONFIG.md carries no claim that [env.<name>.run] overrides any section, and the 0.5.2 changelog does mention the directory-valued-phase error.

  • WRITING-SCENARIOS’s two sample outputs match the binary again. The macros sample showed two builtins with no ellipsis and omitted the (builtin, unused here) marker and the trailing count; the missing_config_var sample dropped the (or in the active [env.<name>.url]) clause. Both read as verbatim transcripts, so a reader comparing them against a real run found differences that were the document’s, not theirs.

  • One worklist instead of four documents to cross-read. Four files read like backlogs and only one was: OPEN-FINDINGS now carries every open item, including the residue of both UX reviews (R1–R3) and the decisions taken against them, each entry self-contained. The two review documents were removed once their open items landed there — their transcripts and citations remain in git history, and a retired review left on disk is exactly the thing that reads as a backlog. IMPROVEMENT-PLAN stays a separate file — five ADRs cite it by section number — but its master table gained a Status column, because its ✅/⚠️ glyphs mean “fits the architecture”, never “done”, and 13 of its 16 items had already shipped while the table gave no way to tell. Item 14’s cited mechanism (Refs::default() resetting per lower() call) was corrected: 0.6.0 replaced it, and only the cross-scenario half of that caveat still holds.

  • A page for the persona the product is named after. PRD §4’s first persona writes prose against a vocabulary somebody else maintains — and every document labelled “test authors” taught pack authoring, so that reader had no route through the tool. docs/WRITING-SCENARIOS.md covers only their loop: what a sentence is, how to list the ones available, the dry-run cycle, and the two diagnostics they will actually hit. The index now labels each author-facing page with the persona it serves instead of calling six P2 documents “test authors”.

  • bind::unbound_step leads with the action its reader can take. The help opened on “add a macro to a pack” — the pack maintainer’s move, which a scenario author cannot make — and buried theirs in a parenthetical. It now opens with matching a sentence the suite’s packs already bind. It names no tool: Diag.help reaches an editor’s diagnostics pane verbatim through the LSP as well as the terminal, and each front end already has its own way to show the vocabulary (completion in the editor, proef macros in a shell) — proef-core does not know which one is reading. The YAML stub is unchanged: it is load-bearing for the maintainer and stays verbatim.

  • ADR-0014 now records the question it was silent on. It is specific about a failing setup and a failing teardown, so a reader reasonably infers the cancellation case was considered — it was not. What teardown does on Ctrl-C is unspecified, and today it silently skips: the phase runs with the already-cancelled token, every scenario resolves Skipped, and phase_failed ignores a phase that only skipped, so cleanup never runs and nothing says so. The ADR now states the gap and the two defensible answers, since an implementer working on teardown reads the ADR, not the findings list.

  • The open-findings list is now in the repo, not on one machine. A v0.5.3 review was validated claim-by-claim (40 claims, 38 confirmed) and the record lived only in a gitignored scratch directory, so ~26 still-open defects — the Ctrl-C teardown gap, LSP rooting, --sarif line numbers, several docs drifts — existed nowhere durable. docs/OPEN-FINDINGS.md carries them, plus what shipped against them, so a fixed finding is not re-reported and an open one is not lost.

  • proef init is now in the command tables it was missing from. It shipped in 0.6.0 and was documented in GETTING-STARTED.md and in the README’s prose, but not in the README’s CLI table or TECH-SPEC.md’s command surface — so the two places a reader scans for “what can this tool do” both omitted the command that starts a first run.

  • CLAUDE.md’s status list now records the v0.6.0–v0.8.0 correctness series rather than ending at post-M5, so the three releases that closed the reports-success-on-wrong-output bug class are visible to anyone picking the project up.

[0.8.0] - 2026-08-09 (CLI output & exit integrity)

Changed

  • A set-but-unreadable environment variable is now a loud user error, never silence — breaking for a pipeline that relied on the old silent fallback. std::env::var collapses “unset” and “set to bytes that are not valid UTF-8” into the same Err; .ok() erased that distinction at five call sites, so a value proef could not read was indistinguishable from one the user never set. A non-UTF-8 PROEF_KEY fell through to the key file and decrypted with the wrong key, reporting tampering instead of the real cause (and doctor reported the key source as the file instead of the override); a non-UTF-8 PROEF_SECRET_<NAME> fell through to the store and reported a missing secret; a non-UTF-8 PROEF_ENV ran silently against the wrong environment, including in proef lsp, where it meant analysing against the wrong config profile. Four of the five sites now exit 2 (user error) naming the variable; doctor instead reports it as a failed check alongside its other unready-environment findings and exits 3, the same as an unreadable key file. A pipeline that today tolerates a mis-set PROEF_ENV, or a non-UTF-8 key/secret, will start failing after this upgrade.
  • A failed stdout write now reaches the exit code — breaking for a pipeline that tolerated truncated output. Writing to a full disk or other failed stdout exited 0 with truncated output; it now exits 3. A closed pipe (proef … | head) still exits cleanly. A pipeline that captures proef’s stdout somewhere that can fail mid-write (a full disk, a device error) previously reported success over truncated output; it now gets a nonzero exit it can act on instead of trusting truncated bytes. Per docs/RELEASING.md, any breaking change is MINOR — together with the environment-variable change above, this forces the next release to be 0.8.0, not 0.7.1.

Fixed

  • proef fmt rewrites line endings wholesale, violating its hurl-blocks-only promise. fmt split pack files with text.lines() (which strips both \n and \r\n) and rejoined with hardcoded "\n", so CRLF files became LF. On an autocrlf checkout (a supported way to clone this repo), fmt --check was permanently failing through no fault of the author. fmt now detects the file’s dominant line ending and preserves it when rewriting.

  • run.log could gain duplicated fragments when the console accepted a short write, because the tee re-wrote the full slice on every retry. It now mirrors only the accepted bytes.

  • proef report -o outside the run dir wrote artifact links relative to the run dir, so every link 404’d from the report’s own location while the command reported success. The href is now absolute when the report is written elsewhere.

  • proef diff reported a brand-new retried step as newly flaky, because a step absent from the base run was assumed to have run once. Steps with no baseline are now skipped, and the ordinal-shift caveat inherent to positional step keying is documented in TROUBLESHOOTING.

[0.7.0] - 2026-08-07 (record & artifact integrity)

Changed

  • ${fake:…} values no longer repeat across a scenario’s steps. The occurrence counter restarted on every step, so two steps each asking for a fresh ${fake:email} received the same address. Every independent ${fake:…} reference within a scenario — across steps, and within one step’s payload/when:/label — now gets its own value and never collides with another, however many a single step ends up resolving. A step’s name: label (shown in artifact comments and events) is the deliberate exception: it is not independent of its own payload, so it replays from the start of the step’s own occurrence window instead of minting new ones, matched by position (the label’s Nth ${fake:…} reference reuses the payload/when:’s Nth occurrence, regardless of generator kind) — so it reproduces the payload’s own value when the label’s references mirror the payload’s in kind and order, and shows a different generator’s output when they don’t. Even a label with more ${fake:…} references than its payload still reserves each extra one, so a later step can never be handed a value the label already displayed. Values remain deterministic for a given --run-id, but suites using ${fake:…} will see their emitted artifacts change. Known limitation, not fixed here: the counter resets at the start of every scenario, not the run, so two different scenarios that each resolve ${fake:email} at the same position in their own step order still collide — that is a separate bug with its own snapshot-moving fix.
  • proef_core::resolve::resolve changed signature (public API break for downstream proef-core consumers): it now takes an additional &mut usize occurrence counter supplied by the caller, and Resolution::fakes was removed — resolve() no longer owns the counter itself.

Fixed

  • run_finished is once again the last line of a run record. A scenario the watchdog abandons keeps running on a detached thread and only notices its cancellation token at the next batch boundary, so it went on appending events after the sweep had recorded its outcome — and after the run itself was finalized. docs/EVENTS.md has always said the last line is run_finished; it was not, so anything reading a record as a stream (the JSONL consumer, report, explain) could see events arrive after the terminal one. Late events from a finalized scenario are now dropped at a single gate rather than by asking every emitter to check. Abandonment itself is unchanged and stays cooperative (ADR-0007) — only the record’s tail is affected.

  • .map.json no longer loses a request’s captures when the pack comments one of them. A comment inside an open [Captures] run is the author’s note about a capture, not the start of the next entry, so it no longer closes the scan — previously it dropped every capture after the comment. The entry that follows opens with a method or response line, and that closes the run on its own.

  • .map.json no longer lists captures that were never made. The sidecar’s capture scan was fence-unaware — a literal [Captures] line inside a fenced (…) body re-armed it — and it recognised only the stock HTTP methods, so an entry opened by a custom method (PROPFIND, …) never ended the previous scan. Both let capture names that don’t exist in the emitted entry land in .map.json, a normative artifact (ADR-0010). The scan is now fence-aware and shares the lowering pass’s method recogniser (is_method_line) instead of carrying a second, weaker copy.

  • pack::empty_expect now also catches a whitespace-only hurl: fragment. The diagnostic already existed for an expect: item with neither status: nor hurl: at all; a hurl: key present but carrying no non-blank assert line slipped past it, lowered to an empty asserts block. It also gains a remediation hint and the seeded corpus case it was missing. Scope: this check reads the unresolved pack text, so a fragment that is non-blank as authored but resolves to nothing at lower time (e.g. ${vars:key} naming a proef.toml value that is "" in the active environment, or an unset ${global:key} under --dry-run) still lowers to an empty asserts block — see the sidecar-emitter entry below for how that residual case is handled.

  • The sidecar emitter can no longer produce an inverted .map.json span. A Then step whose asserts all resolved to nothing — reachable even after the pack::empty_expect widening above, since pack validation cannot see what a fragment resolves to, only what it says — lowered to a zero-line merged-asserts step, and the emitter’s line-span arithmetic underflowed: the start offset exceeded the end. Such a step now gets no sidecar row at all instead of an inverted one — nothing was appended to the artifact, so there is nothing to report a span for.

[0.6.0] - 2026-08-07 (first-run UX & run-record correctness)

Added

  • proef init scaffolds a working suite. It writes the files GETTING-STARTED.md teaches — proef.toml, one .feature, one matching pack — installs the pack JSON Schema for editor completion, and prints the next command. Nothing is ever overwritten, so a second run is a no-op and no --force flag exists to destroy authored work. A test asserts the scaffold passes --dry-run unchanged.
  • The README now shows a parameterized macro and states the load-bearing non-goals, including the supported path for teams that already have a hurl corpus.

Changed

  • A passing --dry-run now names the next command. Every failure path already named a remedy; the success path stopped talking at the moment a new user decides whether to continue.
  • A scenario with no steps is now an error, not a silent pass — breaking. A Scenario: with a commented-out or never-written body previously bound to nothing, ran nothing, and exited 0; it now exits 2, through proef test, proef flows, the libtest-mimic harness, and proef-lsp (which re-analyzes on didChange, so a half-typed Scenario: now shows a live error while you’re still typing it). Per docs/RELEASING.md, any breaking change is MINOR — this forces the next release to be 0.6.0, not 0.5.4.

Fixed

  • resolve::missing_config_var now suggests the closest key defined in the same namespace, matching resolve::unknown_variable and resolve::fake_unknown. Candidates are namespace-scoped, so a ${url:…} typo can never suggest a [vars] key. The code also gains the seeded corpus case it was missing.
  • proef init no longer rewrites a pack it declined to create. Installing the editor modeline ran unconditionally, so a hand-authored suite/packs/api.yaml reported as “already exists” was still modified; the schema install is now gated on the file having been created, and an existing pack gets a hint naming proef schema --add-to instead.
  • Setup and teardown no longer corrupt the run record. Each phase bracketed its own run_started/run_finished, so one record held up to three pairs and proef explain reported the last phase’s totals — printing “1 passed · 0 failed” above a failure it had just listed. The record now carries one pair, and its run_finished totals are the main suite’s own verdict — [run] setup/teardown scenarios still appear as their own events in the record, but are never folded into passed/failed/skipped, so those numbers agree with the console summary: line, JUnit, --output json, TAP, the SLA gate, and the exit code. The console run header also prints once per run instead of once per phase.
  • report and explain flag a truncated run. Both rendered an incomplete record as if it were whole; explain also derived its headline solely from the missing tail event, reporting all zeros for a record that held completed scenarios. Both now read through the same record reader diff uses.
  • explain’s step/attempt totals count a still-in-flight scenario. A step only attached to the record once its ScenarioFinished landed, so a scenario still running when a truncated record’s stream ended had its step evidence silently dropped from the headline — the one place a post-mortem tool most needs it. Totals now fold the raw events directly instead.
  • explain’s failure detail is keyed (file, scenario), not scenario name alone. Two same-named scenarios in different files previously bled each other’s failure output together.
  • worker is the slot a scenario occupied, not a per-scenario counter. The timeline drew one lane per scenario regardless of --jobs.
  • Run rotation only treats hyphenated UUID directories as run records. The parser also accepted bare 32-hex, urn:uuid: and braced spellings, which rotation could then delete when the runs directory points somewhere shared.
  • The nightly canary can fail again: its step piped through tee without pipefail, so a red canary exited 0 and the open-an-issue step was unreachable.
  • The raw-print-macro guard now covers proef-lsp, where stdout is the JSON-RPC channel and a stray print corrupts protocol framing.

Documentation

  • The stdout/stderr macro rule is now written down where contributors look: docs/CONTRIBUTING.md (“Rules that are easy to trip over”) and CLAUDE.md. 0.5.3 began enforcing it with a source-scanning test, so a raw println! or eprintln! in proef-cli failed the suite with nothing explaining the rule or naming render::outln!/errln! as the sanctioned spellings.

[0.5.3] - 2026-08-06 (closed-pipe safety)

Fixed

  • The CLI no longer panics when stderr is a closed pipe. Every remaining raw eprintln! in proef-cli now routes through the EPIPE-safe errln! guard added in 0.5.2, so proef test … |& head ends the pipeline with the contracted exit code instead of aborting with 101 — a code outside the typed 0/1/2/3 taxonomy (ADR-0009). The execution failure summary, which writes several lines per failing scenario, was the largest remaining exposure. A source-scanning test now keeps raw eprintln! out of the crate.
  • The language server no longer dies while recovering from a panic. proef-lsp reports a caught analysis panic on stderr; that report used a raw eprintln!, which panics when its write fails — so a closed stderr (EPIPE) took down the very server the surrounding catch_unwind exists to keep alive. The write is now explicitly unchecked. Ships without a test: reaching the line needs a real analysis panic and a closed stderr, and the panic is not injectable without a test-only hook in shipping code; the mechanism itself is already covered by the CLI’s closed-pipe tests.

Changed

  • proef report derives its output directory through the shared fsutil::parent_dir helper instead of an open-coded empty-parent fallback, so there is one spelling of that derivation. Internal consistency only — the emitted artifact links are unchanged.

[0.5.2] - 2026-08-05 (CLI correctness)

Fixed

  • A directory-valued [run] setup/teardown is now a loud user error. ADR-0014 defines setup/teardown as a single feature file; a directory ran every feature under it as the phase and again in the pool (a silent double-run) — that path is closed.
  • Diagnostics no longer panic when stderr is a closed pipe: print_all and report_front_error’s trailing "{errors} error(s)" summary line are now routed through an EPIPE-safe errln! guard (mirroring outln!’s stdout guard), so proef test --dry-run <broken suite> |& head exits cleanly instead of panicking (exit 101).
  • diff step records are now keyed by (text, occurrence ordinal) instead of text alone — macro-expanded steps that share text no longer collide in the last-write-wins map and silently drop out of the diff.
  • diff --fail-on-regression now fails when the new run is incomplete or cancelled (was a silent pass), and banners any incomplete/cancelled record in the diff output either way. Its slower-step duration math is hardened against overflow (saturating arithmetic).
  • A bare-filename [run] setup/teardown (or suite path) now resolves its packs and assets from the current directory. A path with no directory component (e.g. setup = "setup.feature" at the project root) has an empty Path::parent(), which produced a cannot read directory failure; it now normalizes to . (the current directory) via a shared fsutil::parent_dir helper at the pack/asset base-derivation sites.

Documentation

  • The second-interrupt hard-exit code 130 (128+SIGINT) is now documented for test and watch (TECH-SPEC §10, ADR-0009) — a deliberate escape hatch outside the typed 0/1/2/3 ExitCode taxonomy.

[0.5.1] - 2026-08-05 (LSP go-to-definition + correctness)

Added

  • LSP go-to-definition: use: references and match: landing (ADR-0017). Go-to-definition now jumps from a use: reference in a pack to the macro it targets, and lands on the macro’s match: line rather than its name key (falling back to the name key for use-only macros with no match:).

Fixed

  • LSP: the stdio server now exits cleanly. proef lsp dropped the connection after joining the transport threads, so the writer thread (holding the sole channel Sender) never ended and the process leaked. It now drops the connection before joining. Covered by a real stdio subprocess lifecycle test.
  • LSP: a malformed request no longer crashes the server. A bad document URI or out-of-range position propagated a deserialization error out of the event loop and exited the process; the request now gets an InvalidParams (-32602) reply and the server keeps serving.
  • LSP: one broken pack no longer blanks the whole suite. analyze_suite now keeps the packs that loaded (and reports the broken one’s diagnostic) instead of zeroing all bindings, completion, and go-to-definition on any pack error.
  • LSP: analysis is scoped to the configured suite. The server roots at [run] suite (else the tests/ convention) under its launch directory rather than walking the entire working tree, sharing the CLI’s suite resolution.
  • LSP: unsaved edits are honored for paths with special characters. The open-buffer overlay is keyed by source name instead of the raw file URI, so a path segment containing sub-delimiters ((, +, ', …) no longer misses.

Documentation

  • Documented proef-lsp and the lsp/macros/diff/report subcommands across the README, TECH-SPEC CLI/dependency references, and the RELEASING publish order.

[0.5.0] - 2026-08-04 (LSP language server)

Added

  • proef lsp language server (ADR-0017). A server-only, generic-LSP stdio binary — a second front-end over the sans-IO core — giving feature/pack authors live editor support: diagnostics (the whole --dry-run validation set, republished across the suite as you type), go-to-definition (Gherkin step → the macro that binds it), completion (macro-pattern step completions, prefix-ranked by relevance to the typed prose), and find-references (every step a macro binds). Wired into Neovim/Helix/Emacs via generic LSP config — see docs/EDITORS.md. No VS Code extension in v1. proef.toml config is a startup snapshot (restart the server after editing it). Works on Linux, macOS, and Windows. Pinned lsp-server 0.7.9 / lsp-types 0.97.0.
  • New proef-core public surface enabling the language server: the injectable SourceProvider seam (proef_core::provider), the collect-all analyze_suite analysis (proef_core::analyze) — the same headless analysis the CLI runs, driven over an overlay-then-disk provider so the LSP re-validates the whole suite on every edit — and matcher::prefix_rank for prose-prefix completion ranking. All keep the core sans-IO (the IO is injected).

[0.4.0] - 2026-08-03 (external config & environments; competitive-review breadth)

Added

  • Suite setup & teardown (proef.toml [run] setup/teardown, ADR-0014). Each names a feature run once around the whole suite (the Playwright/Jest globalSetup model). setup runs before the parallel pool and merges its saveAs: global promotions into the shared store before any scenario lowers, so it seeds fixtures/shared state every scenario reads via ${global:…}; teardown runs once after for cleanup. A setup failure aborts the run as a user/system fault (never a test failure, exit 1); teardown runs only if setup succeeded and its failure is a distinct exit 3 (never a silently green suite). Both are excluded from the pool, so a setup/teardown feature inside the suite never also runs as an ordinary scenario.

  • proef test --output tap — a TAP version 13 stream to stdout, one test point per scenario, derived from the run’s own outcomes (not from hurl), for prove/tappy and TAP-native CI. The human report moves to stderr (as with --output json). @quarantine scenarios map to the # TODO directive (their failure does not gate); skipped scenarios to # SKIP; failure detail rides in a redacted YAML block. --output tap is rejected on flows/macros (a user error, not a silent human fall-back).

  • proef macros now flags near-duplicate pattern macros — two that differ only in their {capture} names (identical literal skeleton), which are confusable to authors. Advisory only (never gates the exit code); --output json gains a nearDuplicateOf field beside unused for a CI hygiene check. The heuristic is deliberately tight (skeleton equality), so a legitimately similar family with distinct literals is left alone.

  • Localized Gherkin (# language:) is now verified and test-covered — a localized feature parses, its dialect keywords are stripped, and a localized scenario outline with Examples expands like any other. Outline detection now keys primarily on Examples presence (dialect-independent) with the English keyword as a fallback, so this no longer relies on an English-only heuristic. (A localized outline that omits its Examples still degrades to an unbound-step error, since gherkin 0.16 does not expose its dialect keywords.)

  • Built-in expect: shape-macro library. The embedded Core pack gains a curated, product-neutral set of response-shape assertions — the value at {path} is a string / … a number / … a boolean / … a uuid / … an ISO date / … present / … a non-empty list — each merging one hurl type predicate (isString/isUuid/isList + count, …) into the previous request. It is a convenience layer over the existing expect: mechanism (no new engine capability, no marker DSL); the raw-hurl assert vocabulary still covers anything the macros don’t.

  • Run-level SLA gate (proef.toml [sla]). An opt-in latency budget: after a run, per-step wall-clock durations fold into p95-ms (95th-percentile ceiling) and max-ms (slowest-step ceiling); a breach prints the offending metrics + the slowest steps and maps to exit 1 (a test failure). It is off by default (no [sla] table = no gate, run byte-identical to before), env-overridable via [env.<name>.sla], introduces no new exit code, and never downgrades a User/System fault. Distinct from hurl’s per-request duration < assert — the gate is an aggregate budget over the whole run. Skipped steps are excluded from the population.

  • External config & environments (proef.toml, ADR-0012). New [url] and [vars] tables hold non-secret suite variables, referenced in packs as ${url:<key>} / ${vars:<key>}; [env.<name>.<section>] profiles deep-merge per-environment overrides over the base tables (url/vars/http/run). proef test --env <name> (or PROEF_ENV) selects the active environment. proef.toml is discovered by searching up from the working directory (like cargo/git), so it is found from any subdirectory. Adds the proef::resolve::missing_config_var diagnostic.

  • Default suite path. [run] suite sets the path proef test/flows/ artifacts use when given none (falling back to the tests/ convention), so proef test runs with no argument. An explicit path still wins.

  • Documentation set completing the corpus: docs/DIAGNOSTICS.md (all 57 diagnostic codes, corpus coverage marked), docs/CONFIG.md (proef.toml reference), docs/EVENTS.md (the events.jsonl wire schema for CI), docs/TROUBLESHOOTING.md (exit codes, glyph legend, frequent failures), docs/CONTRIBUTING.md and docs/SECURITY.md (threat model, private vulnerability reporting), and an IDE-integration section in AUTHORING.

  • proef test --scenario-file <file>: scope a --scenario name filter to one feature file (duplicate scenario names across files stay disjoint; the libtest-mimic harness uses it to keep the Trial↔scenario bijection).

  • scenario_finished events now carry a file field — the run-wide scenario identity alongside scenario (additive, ADR-0008; absent in older records).

  • Diagnostics pack::pattern_duplicate_capture (a {capture} written twice) and lower::kind_unrouted (internal registry-drift safety net).

  • proef macros lists every loaded macro with its call count and flags user-pack pattern macros that no scenario binds (dead prose bindings); use:-only helpers and unused builtins are listed but never flagged. --output json for CI dead-code gates.

  • proef test --run-id <id> pins the injected run id (like artifacts --run-id), so a run’s ${fake:…} data — which keys on the run id — is reproducible; the JSON summary echoes the id.

  • proef test --dry-run --sarif <path> serializes validation diagnostics (unbound steps, pack lint, non-finite retries) to a SARIF 2.1.0 log — a shift-left gate that renders findings as inline PR annotations. The export is additive: the dry-run’s exit code is unchanged.

  • proef test --rerun re-runs only the scenarios that failed in the last run (read from its JSONL record, keyed on the run-wide (file, name) identity); it composes with --tags/--scenario, and reports “nothing to rerun” (exit 0) when the prior run was clean.

  • @quarantine tag: a scenario so tagged runs and reports normally, but its test-failure no longer gates the exit code (a System/User fault still does — quarantine is for flaky tests, not broken input or infra). A note prints when a quarantined scenario fails, so it is never silently swallowed.

  • proef diff [base] [new] compares two run records (defaulting to the previous and latest runs) and reports scenario status transitions — regressed, fixed, still-failing, new, removed — keyed on the run-wide (file, scenario) identity, plus per-step flakiness (rising retry counts) and perf deltas (steps diffed on text, never the volatile authored line). It is a derived view over events.jsonl, never a second record (ADR-0008); --fail-on-regression exits 1 when a scenario regressed, for CI gating.

  • Flaky-failure detail: a step that passes only after a retry now records the messages from its earlier, failed attempts as attempt_details on the step_finished event (additive, ADR-0008); JUnit surfaces them as <flakyFailure> under the passing test case, so a green-on-retry run is honest instead of indistinguishable from a clean pass. The engine already collected the earlier-attempt errors — they were being discarded on success.

  • proef report [run-id] writes a self-contained HTML report for a run — scenario tree with pass/fail pills, per-step attempts and timing, a per-scenario timing waterfall (each step’s bar offset by the steps before it and as wide as its own duration — the sequential cascade within a scenario, derived purely from step durations), a cross-worker timeline (a lane per worker, each scenario a bar on a shared run-relative axis, so concurrency is visible at a glance), failure detail, and deep-links to the executed .hurl artifacts (bodies are not inlined).

  • Injected run timing (ADR-0015). scenario_started/scenario_finished events gain optional timestamp_ms (run-relative) and worker (0-based index) fields, stamped at the CLI sink on the worker thread so the sans-IO core stays clock-free. Additive (absent on records without timing); they power the HTML timeline. Records without them degrade to the waterfalls alone. A pure proef_core::html::render_html derives it from the event stream (ADR-0008, snapshot-locked); the events are already redacted at the sink, so the page is too. Defaults to report.html inside the run dir; -o redirects it.

Changed

  • --tags is now a boolean expression, not a comma-separated list. It takes a single expression over and/or/not and parentheses (the @ stays optional), e.g. --tags "@api and not @slow"; a bare tag still works. The grammar and evaluator live in the sans-IO core (proef_core::tags, deterministic and fuzzed); a malformed expression is a user error (exit 2), as is a selection that matches nothing. This replaces the old CSV OR-list — there is one selection mechanism, not two.
  • --output is a typed value: an unknown format (e.g. a jsonl typo) is a user error (exit 2) instead of silently degrading to the human report.
  • --watch reruns only on .feature/.yaml/.yml changes — the watched tree can now contain proef’s own run output without a self-trigger loop.
  • The example corpus (tests/features/) and the dev fixture use a neutral workspace / activity-board domain (record · note · event · attachment · session · channel) — no product-specific vocabulary.
  • CHANGELOG.md, CONTRIBUTING.md, and SECURITY.md moved under docs/ (root keeps only README.md and CLAUDE.md).
  • Pack root key renamed templates: → macros: (ADR-0004 amendment): one canonical spelling for the prose→engine binding layer (the entry is a macro, the file a pack). No templates: alias — packs using the old key fail to load.
  • The dev-loop fixture (cargo run -p xtask -- fixture) binds the advertised default port 8787 — falling back to an ephemeral port (and printing a PROEF_BASE_URL line) only if 8787 is busy; ... -- fixture <port> overrides. So proef.toml’s default base reaches it with no PROEF_BASE_URL export (ADR-0011 amendment). Its GET /health now returns a versioned identity — name, a numeric version (1.0), and the RFC 3339 time it answered.
  • The unbound-step diagnostic (bind::unbound_step) now prints a paste-ready pack-macro stub — quoted tokens in the sentence become {argN} captures — alongside the existing did-you-mean suggestion, so an author can add the missing macro without hand-writing the match:/hurl: scaffold.
  • CI reporting surfaces failures and flakiness more honestly. Under GitHub Actions the run emits a ::error file=,line=,title= annotation per failure (rendered in the PR “Files changed” gutter; gated off when --output json owns stdout). The job summary gains a flaky passes section and per-failure attempt counts, and the JUnit report records “passed on attempt N” for a scenario that only went green after retries — a silent green-on-attempt-2 is no longer invisible.
  • docs/AUTHORING.md gains an “Asserting responses” cookbook surfacing the hurl 8.0 predicate/filter/RFC-9535-JSONPath vocabulary that raw hurl: blocks already accept — documenting existing capability, not new engine work.
  • A failed step now prints a curl: reproduce line — the redacted curl for the failing request, surfaced from the embedded engine via a new engine-agnostic StepOutcome.reproduce_hint — so a failure can be replayed request-by-request without leaving the terminal. Secrets are masked.

Removed

  • The # key: value feature-file directive mechanism (e.g. # baseURL:, ADR-0012 amendment). Variables now have exactly one home — proef.toml ([url]/[vars]) — so a .feature file can no longer define a variable (one-way-to-do-one-thing). # comment lines stay valid gherkin comments; they are simply no longer parsed. The env-override the directive provided is preserved by embedding ${env:NAME:-default} in a config value (resolved recursively). ${…} plain-name resolution is now args > defaults only.

Fixed

  • Optional-batch error path no longer double-reports later batches into the JSONL run record (ADR-0008); saveAs: global promotions are no longer dropped when the store lock is poisoned; the event sink recovers from a poisoned lock instead of truncating the record.
  • expect: merge scopes to the last entry (fence-aware); [Options] injection can no longer duplicate a section; the use: graph walk is node-linear instead of exponential on multi-edge chains.
  • The embedded-hurl version lockstep is now asserted by a test; the encrypted secret store maps user vs. environment faults to exit 2 vs. 3 (ADR-0009); run.log / artifact-write / malformed-proef.toml failures surface instead of being swallowed.

[0.3.1] - 2026-07-29 (secret-management hardening)

Added

  • proef secret rm NAME removes a stored secret (locked atomic rewrite; removing an absent name exits 2).
  • PROEF_KEY env override supplies the project key directly (base64) — a committed ciphertext store now decrypts in CI without shipping the key file; a set-but-invalid key errors instead of silently falling through.
  • proef doctor reports secret store/key health (readable, parseable, private permissions); a corrupt .proef-secrets.json no longer bricks secret set — it is moved aside to .corrupt and a fresh store begins.

Fixed

  • Secret-valued captures never reach .proef-state.json: a saveAs: global capture whose value equals a known secret is refused — the owning step warns with the reason — closing the one sink the redaction invariant (ADR-0005) did not cover.
  • Secret resolution reads the store and key once per run instead of once per secret (no torn view against a concurrent secret set).
  • Warned steps now print their reason on the console (↳ …) — a bare ⚠ glyph explained nothing, for optional: failures too.

[0.3.0] - 2026-07-29 (data-safety blockers, Then visibility, taxonomy)

Fixed (v0.2.1 review — every finding reproduced before fixing)

  • Asset copy destroyed user files: proef artifacts -o pointing at the suite truncated referenced assets to 0 bytes, and .. references escaped the output directory. Copies now refuse absolute/.. references (exit 2), never copy a file onto itself, and surface IO errors (exit 3).
  • Run rotation deleted arbitrary directories: with runs-dir shared with user content, rotation could recursively delete user directories — and its own in-flight run. Only uuid-named run records rotate now, never the live run, and rotation happens before the new run dir exists.
  • Zero-entry payloads passed silently: a comment-only hurl: block ran nothing while the scenario reported green. Load-time lint rejects it; the engine backstop emits Skipped outcomes for anything that slips through.
  • proef flows … | head (and every other command) tolerates a closed pipe; a non-UTF-8 environment variable no longer aborts any command.
  • Raw [Options] retry:/repeat: values are parsed and capped (10000), and delay: is capped at 1 hour in both typed and raw forms; repeat: now counts toward the batch budget so long repeats aren’t blamed on the environment.
  • Concurrent proef secret set calls no longer lose keys (advisory-locked, atomic 0600 temp+rename store; the key-creation race resolves to the winner’s key). proef fmt and schema --add-to write atomically.
  • proef fmt keeps fenced body bytes verbatim (blank lines and trailing whitespace inside ``` fences are the bytes the test sends).
  • Nested suites now load their packs: pack discovery recurses like feature discovery (packs/ directories at any depth); proef fmt shares the rule.
  • Duplicate/empty Examples header columns are a named error instead of a silent last-value-wins; an empty .feature gets a plain-language error; a UTF-8 BOM is stripped instead of shifting every diagnostic span.

Changed

  • Then steps are visible everywhere: expect: macros now surface as their own step rows in console, events, JUnit, and explain, with assert failures attributed to the authored Then line — the host request no longer inherits its followers’ assert failures. Artifact bytes are unchanged; sidecars gain one row per Then (schema-compatible).
  • Error taxonomy: mistakes in the test’s own text (undefined {{var}}, bad JSONPath/regex/URL/options, unreadable body file) exit 2 instead of 3, anchored on hurl’s own assert-context flag.
  • when: guards skip on a literal false/0 as well as empty — an author writing when: ${flag} with flag=false means skip.
  • proef.toml is no longer gitignored (it is documented, committed project config).
  • proef-core public API: removed dead surface (NormalizeReporter, the never-populated config resolution tier, StepOutcome.artifact_span, LoweredStep.retry, StepKeyword, and friends); added EngineErrorClass::UserInput, StepPayload::MergedAsserts, ScenarioOutcome.artifact_slug, Guard::skips.

[0.2.1] - 2026-07-29 (review P0 + failure UX)

Fixed

  • [Options] header detection follows hurl’s token grammar — the injection can never land inside XML/JSON/prose bodies (class closed, unit-tested).
  • proef artifacts survives a closed pipe (exit 0, best-effort writes).
  • --dry-run honors --scenario/--tags with the same zero-match exit 2.
  • Duplicate scenario names dedup feature-wide (#N): unique artifacts, console buffers, and events — no silent overwrite.
  • .proef-secrets.json is created 0600, gitignored, and documented.

Changed

  • Failure details surface hurl’s computed expected/actual (fixme) anchored on the error’s own artifact line, not the entry’s first line.
  • GETTING-STARTED uses PROEF_BASE_URL, points the reader at a runnable target, and frames sample output honestly.

[0.2.0] - 2026-07-29 (correctness, output contract, author docs)

Fixed (v0.1.0 deep-review follow-up — all three blockers reproduced first)

  • [Options] injection is body-fence-aware: a retry:/delay: step whose body contains method-looking lines no longer gets options spliced into the body it sends.
  • Step↔entry correlation is a partition anchored on each entry’s request line: a comment-only step can no longer cause the next request to be sent twice (one authored POST is one POST, asserted via the event stream).
  • delay: joins the watchdog budget (with saturating duration math throughout), so delayed steps are no longer killed as system errors; retry.count is capped at 10000 by the pack lint.
  • A panicking scenario thread is contained (catch_unwind), reported as a System fault under its real identity immediately — never a budget timeout; abandoned scenarios keep their real file/name/line; steps in batches never reached report Skipped instead of vanishing from every report.

Changed

  • Output contract: --output json owns stdout exclusively (human report on stderr — pipeable into jq); StepFinished events carry a detail failure field (additive); optional: failures report Warned everywhere consistently; engine failure details use hurl’s own error descriptions instead of Rust Debug; diagnostics drop ANSI when stderr is not a terminal; a filter selection matching nothing exits 2; failure output prints a ready-to-run reproduce: hurl … line; the artifact replay header names required --secret placeholders; the undocumented .env autoload was removed.

Added

  • Author-facing documentation: docs/GETTING-STARTED.md (first suite in ten minutes) and docs/AUTHORING.md (the full pack/feature reference).
  • Mechanical alignment gates: xtask docs-check (crates and ADRs must appear in their indexes) runs in PR CI; xtask public-api snapshots proef-core’s public API surface (1.4k items) and fails CI on unreviewed changes — the mechanical form of the zero-core-diff invariant.

[0.1.0] - 2026-07-29

Initial release.

Added

  • Authoring: Gherkin .feature files in plain business prose; YAML macro packs bind prose to executable steps via match: patterns, typed params, defaults, use: composition (cycle-checked), expect: assert-only macros, optional:, finite retry:, delay:, when: guards, and saveAs: global promotions.
  • Validation: proef test --dry-run binds, lowers, emits, and parse-validates every scenario without touching the network; stable diagnostic codes with source-span rendering; a seeded error corpus pins every code; pack payloads are validated at load by the engine that claims them.
  • Execution: the hurl engine runs artifacts in-process (exact-pinned hurl 8.0.1); contiguous same-engine steps batch maximally; variables and cookies chain across batch splits; per-entry [Options] override batch defaults; finite budgets with a watchdog bound every scenario; Ctrl-C cancels gracefully (twice = hard exit); parallel scenarios share a typed World with write-set-only merge-back and a persistent global store.
  • Artifacts as the contract: every scenario emits canonical .hurl text that is byte-identical to what the engine executes, plus a sidecar map (entry ↔ feature anchors, explicit batch/step indices), .vars, and any referenced file assets — replayable with stock hurl --test.
  • Record & reporting: a versioned JSONL event stream is the run record (live per-entry progress included); console BDD tree; JUnit XML; GitHub job summaries; proef explain replays the record; secrets are encrypted at rest, injected via hurl’s redaction, and value-redacted once at the event sink — never present in artifacts, events, logs, or reports.
  • Tooling: proef flows, artifacts, schema (merged JSON Schema with editor modelines), secret set|list, fmt (canonical hurl blocks), doctor, --watch; a libtest-mimic harness exposes one test per scenario to nextest/IDEs; ${fake:*} deterministic synthetic data seeded from the run id.
  • Quality gates: unit + property tests, fuzz targets, insta snapshot corpus (artifacts, diagnostics, events), fixture-server integration suite, assert_cmd CLI/exit-code suite (0/1/2/3 contract), cargo deny/machete/ zizmor in CI, cargo audit nightly, a scheduled canary against the next hurl release, and CI on Linux, macOS, and Windows.
  • Distribution: tagged releases build five targets (macOS arm64/x86_64, Linux arm64/x86_64-gnu, Windows x86_64-msvc) with cargo auditable, ship a Homebrew tap formula and a cargo binstall-compatible layout, and attest SLSA provenance once the repository is public.