proef — documentation index
proef (Dutch: test, trial — and tasting) is a declarative, modular, multi-engine
end-to-end test runner, mostly Rust. Tests are Gherkin .feature files in business prose;
macro packs bind the prose to executable steps; a pluggable engine runs each step batch —
embedded Hurl for API testing (the seam admits future engines; none are scheduled).
This folder is the project corpus: the product requirements, the decision log, the
normative technical spec, the milestone plan, and the testing strategy. Written
2026-07-28 from a validated research round (a working spike ran 5/5 scenarios green
under both a prototype native runner and stock hurl 8.0.1 on identical generated
artifacts); implementation has since delivered milestones M0–M5 and everything after
them — through v0.19.0 (the 0.18 survey waves, then the gates that could not see
what they covered). Only M6 (a second
engine) is unscheduled. The repo-root CLAUDE.md carries the live status. This
corpus is also published as a website: https://emrecdr.github.io/proef/.
Reading order
| # | Document | What it answers | Audience |
|---|---|---|---|
| 0 | WRITING-SCENARIOS.md | Write prose against a vocabulary somebody else maintains: see the sentences, the dry-run loop, the two errors you will hit | P1 test authors |
| 0 | GETTING-STARTED.md | Your first suite in ten minutes — including the packs behind it | P2 pack maintainers |
| 0 | AUTHORING.md | The pack/feature reference from the author’s seat | P2 pack maintainers |
| 0 | EDITORS.md | Wiring proef lsp into Neovim/Helix/Emacs for live diagnostics, jump-to-macro, completion | P1/P2, whoever sets up the editor |
| 0 | TROUBLESHOOTING.md | Exit codes, glyphs, frequent failures, digging into runs | everyone |
| 0 | CONFIG.md | Every proef.toml key with defaults | P2 pack maintainers |
| 0 | DIAGNOSTICS.md | The greppable index of every diagnostic code | P2 pack maintainers |
| 0 | EVENTS.md | The events.jsonl wire schema for CI consumers | CI engineers |
| 1 | PRD.md | What are we building, for whom, and how do we know it works? | everyone |
| 2 | adr/ — ADR-0001 onward | Why is it built this way? Each decision, alternatives, consequences | engineers |
| 3 | TECH-SPEC.md | How exactly is it built? Types, pipeline, schemas, verified seam facts | implementers |
| 4 | IMPLEMENTATION-PLAN.md | In what order, with what acceptance criteria? M0–M6 task breakdown, risks, runbooks | implementers |
| 5 | TESTING-STRATEGY.md | How is every layer verified? | implementers |
| — | IMPROVEMENT-PLAN.md | Post-M5 competitive review: the feature roadmap, each item carrying a Status column (14 of 16 shipped) | maintainers |
| — | OPEN-FINDINGS.md | The worklist. Every open defect and gap, whichever review found it, plus what shipped against each | maintainers |
| — | RELEASING.md | Versioning policy and the release runbook | maintainers |
| — | CONTRIBUTING.md | Setup, gates, and the rules that are easy to trip over | contributors |
| — | SECURITY.md | Threat model and vulnerability reporting | everyone |
| — | CHANGELOG.md | Per-release change log (SemVer) | everyone |
| — | ../CLAUDE.md | Repo-root guidance for Claude Code: constraints, seam facts, commands, status | coding agents |
Decision log (ADR index)
| ADR | Decision | Status |
|---|---|---|
| 0001 | Embed hurl’s crates in-process as the API engine | Accepted |
| 0002 | Multi-engine core: EngineFactory/EngineSession seam, step-kind routing, batching | Accepted |
| 0003 | Exact pins + thin zero-diff fork as patch vehicle + upgrade canary | Accepted |
| 0004 | Packs = YAML skeleton + embedded raw Hurl blocks | Accepted |
| 0005 | ${…} author-time / {{…}} run-time variables; World; secrets | Accepted |
| 0006 | Engine traits are sync + dyn; no async machinery in v1 | Accepted |
| 0007 | Cooperative cancellation at batch boundaries + budgets (hurl has none) | Accepted |
| 0008 | Serde event enum = run record; decorator reporter stack; libtest-mimic mode | Accepted |
| 0009 | User/TestFailure/System → exit 2/1/3; miette at the CLI edge | Accepted |
| 0010 | Emitted .hurl artifacts are the executed input (same bytes) + sidecars | Accepted |
| 0011 | Fixture server is synchronous tiny_http, not axum (tokio-runtime ban) | Accepted |
| 0012 | Project config & environments in proef.toml ([url]/[vars]/[env.*], --env, deep-merge) | Accepted |
| 0013 | Typed macro parameters (params name→type map; best-effort literal-args lint) | Proposed (defer) |
| 0014 | Suite-level setup/teardown ([run] setup/teardown features, CLI-edge orchestration) | Accepted |
| 0015 | Injected run-relative timestamps + worker id (sink-stamped, sans-IO core) for the HTML timeline | Accepted |
| 0016 | OpenAPI → suite generator: one-shot seed allowed under a bright line; oracle/drift mode permanently rejected | Proposed (defer) |
| 0017 | proef lsp language server: sync lsp-server, whole-suite wholesale recompute, injectable-provider + collect-all front-end refactor | Accepted |
| 0018 | Named hurl fragments: ref: as a second macro body form, # @proef <name> in real .hurl files, explicit bind: scopes | Accepted |
| 0019 | Reserved tags and the authored skip: @skip[:reason] at the CLI edge, reasons in every sink, authored-vs-mechanical split for --rerun, quarantine aligned in JUnit | Accepted |
| 0020 | Run metadata is explicit-injection-only: --meta/[meta]/[env.<name>.meta], run_started gains env/metadata/shuffled, harvested-vs-handed-over is the boundary | Accepted |
| 0021 | Run discovery and rotation are separate questions: a directory is a record because it holds one, rotation still deletes only uuid-named dirs, ordering follows the uuid-v7 timestamp, then mtime | Accepted |
Naming & identifiers
Project/binary proef · crates proef-core, proef-engine-hurl, proef-cli,
proef-fixture, proef-harness, proef-lsp
· run records .proef-runs/ · persistent
World .proef-state.json · config proef.toml. The crates.io names
proef, proef-core, proef-engine-hurl, and proef-lsp are published and owned.
Provenance
Grounded in two verified sources: (1) hurl master (Orange-OpenSource) — source-level verification of every library seam used, with file:line citations in TECH-SPEC §5; (2) a working spike proving the front end and artifact contract end to end (the spike predates this repo). Ecosystem practices are drawn from cargo-nextest, cucumber-rs, sqlx, rustls, probe-rs, and current (2026) Rust guidance.
Installing proef
Pick whichever fits your machine:
# Homebrew (macOS or Linuxbrew, arm64 or x86_64):
brew install emrecdr/proef/proef
# Prebuilt binary via cargo-binstall (any Rust dev environment):
cargo binstall proef
# From source via crates.io:
cargo install proef --locked
Or grab a prebuilt archive from a
GitHub Release — five
targets ship per tag (macOS arm64/x86_64, Linux arm64/x86_64-gnu, Windows
x86_64-msvc), each with SLSA build provenance
(gh attestation verify <archive> --owner emrecdr) and a .sha256 sidecar
(sha256sum -c proef-<tag>-<target>.tar.gz.sha256 from the download
directory). The Windows zip bundles the libcurl/libxml2 DLLs the binary
needs; Linux binaries expect the distro’s libcurl4 and libxml2 (present
on virtually every system, or apt install libcurl4 libxml2).
Prebuilt binaries need no Rust toolchain and no build prerequisites.
Building from source (the cargo install/cargo binstall-fallback path)
needs the native libraries the embedded hurl engine links:
apt install build-essential pkg-config libssl-dev libcurl4-openssl-dev libxml2-dev libclang-dev on Linux; the Xcode command-line tools suffice on
macOS.
Completions and the man page
Shell completions and a man page travel in every release archive
(completions/, proef.1) — or generate them from any installed binary:
proef completions zsh > "${fpath[1]}/_proef" # also bash, fish, powershell, elvish
proef man > /usr/local/share/man/man1/proef.1
In CI
cargo binstall proef resolves the release archives, so any workflow with a
Rust toolchain installs in seconds; without one, download an archive and its
.sha256 from the Release directly. A complete GitHub Actions workflow —
secrets, JUnit, sharding, the regression gate — is in CI.
First run
proef doctor verifies the environment (embedded engine, native libraries,
project layout, runs-dir writability). Then Getting started
takes you from proef init to a green run.
A note on the dev fixture
The tutorial’s zero-network first run uses the dev fixture API
(cargo run -p xtask -- fixture), which lives in this repository and is not
part of the installed binary — it needs a checkout and a Rust toolchain. With
an installed binary alone, point [url] base at any API you own instead; the
proef init scaffold passes as-is against the fixture, and its routes
(/health, /search, /version) are one file away from targeting yours.
Getting started — your first suite in ten minutes
proef runs end-to-end API tests written as plain Gherkin prose. The prose stays readable by anyone; a YAML pack binds each sentence to real HTTP work. This walkthrough builds a two-file suite from nothing and runs it.
Install first (Installing), then verify:
$ proef doctor
engine `hurl`:
[ok ] embedded hurl hurl 8.0.1 (exact pin, ADR-0003)
...
0. Or scaffold it: proef init
$ proef init
created ./proef.toml
created ./suite/case.feature
created ./suite/packs/api.yaml
created ./hurl/api.hurl
created ./.gitignore
created ./suite/packs/proef-pack.schema.json
ok ./suite/packs/api.yaml (modeline added)
created 6 file(s), skipped 0
next: proef test --dry-run (then point ${url:base} at your API — the scaffold's routes are placeholders)
proef init writes the three files this walkthrough builds by hand below —
proef.toml, suite/case.feature, suite/packs/api.yaml — plus a one-entry
hurl/api.hurl, the pack JSON Schema and a .gitignore. The .hurl file is
there because a step has two body forms: the scaffold’s pack shows an inline
hurl: block and a ref: naming that file’s # @proef entry, so the form an
adopter with an existing hurl corpus wants is visible from the first command
rather than only in the docs. It is a deliberately smaller suite than the
one this tutorial builds: no Then the first hit is record "r-1" step, no
firstHit expect: macro, no ${secret:apiToken}, and search targets a
plain /search route instead of the tutorial’s /api/v1/admin/search/records —
while adding one step the tutorial does not, And the corpus reports its version, whose macro is the scaffold’s ref: example.
Pasting the tutorial’s Then line into the scaffold’s feature file fails with
bind::unbound_step — read on to build the fuller suite by hand, or extend
the scaffold yourself once you understand the pieces.
1. A suite is three files
suite/
case.feature # the prose — what the test says
packs/
api.yaml # the vocabulary — what each sentence does
proef.toml # the configuration — where URLs and variables live (§3.5)
Layout is convention, not configuration. proef test suite takes one path and
discovers everything under it: every *.feature file is a test file (at any
depth), and every *.yaml/*.yml file directly inside a directory named
packs is a macro pack — the packs directory itself may sit at any depth.
All packs merge into one vocabulary shared by all feature files — grow the
suite by adding files; nothing needs registering.
2. Write the prose
suite/case.feature:
Feature: Directory search
Scenario: A known record is found
Given the service is healthy
When the operator searches for "Acme"
Then the first hit is record "r-1"
The feature is pure prose — no URLs, no environment data. The target host and
any variables live in proef.toml (§3.5) and reach the packs as
${url:…} / ${vars:…}.
3. Bind the prose
suite/packs/api.yaml:
macros:
health:
match: the service is healthy
steps:
- hurl: |
GET ${url:base}/health
HTTP 200
search:
params: [term]
match: the operator searches for {term}
steps:
- name: search records for ${term}
hurl: |
GET ${url:base}/api/v1/admin/search/records
Authorization: Bearer ${secret:apiToken}
[Query]
q: ${term}
HTTP 200
[Captures]
recordId: jsonpath "$[0].id"
firstHit:
params: [id]
match: the first hit is record {id}
expect:
- hurl: |
jsonpath "$[0].id" == "${id}"
Three macro shapes are on display: a fixed sentence (health), a
parameterized one ({term} captures the quoted word — quotes are shed), and
an assert-only expect: macro whose lines merge into the previous request’s
asserts. The hurl: blocks are raw Hurl — validated with
the real parser the moment the pack loads.
3.5 Where URLs and variables live (proef.toml)
Variables are declared in proef.toml — never in the .feature files — and
referenced as ${url:…} / ${vars:…}. The pack’s ${url:base} above resolves
from here; proef finds the nearest proef.toml, searching up from the working
directory (like cargo/git):
# proef.toml
[run]
suite = "suite" # `proef test` needs no path argument
# fragments = "hurl" # uncomment to `ref:` entries of real .hurl files (§3.6)
[url]
base = "${env:PROEF_BASE_URL:-http://127.0.0.1:8787}" # → ${url:base} (env override wins)
[vars]
apiVersion = "v1" # → ${vars:apiVersion}
[env.staging.url] # per-environment overrides (mirror the base tables)
base = "https://staging.example.com"
[env.prod.url]
base = "https://api.example.com"
[env.prod.http] # an env may override a runner setting too
timeout-ms = 60000
A macro then reads GET ${url:base}/api/${vars:apiVersion}/… with nothing
declared in the feature. Pick an environment at run time:
$ proef test # base [url]/[vars]; runs `[run] suite` — "suite/" in the file above (`tests/` is only the fallback when the key is unset)
$ proef test --env prod # [env.prod.*] deep-merged over the base (or set PROEF_ENV=prod)
The rule is uniform: under [env.<name>], url.* / vars.* override variables and
http.* / run.* override runner settings — anything unlisted inherits the base,
so an environment names only what changes (the Cloudflare-Wrangler / Cargo-profile model).
Secrets never live here — they stay in the encrypted store (${secret:…}, §5).
3.6 A step’s body has two forms
Everything above uses an inline hurl: block, which is complete and permanent. The other
form is ref: <name>, which points at one entry of a real .hurl file marked
# @proef <name>, with values supplied by a bind: table:
Uncomment fragments = "hurl" in §3.5’s proef.toml (a second [run]
table would be a TOML error — one table, both keys), put the annotated file
under hurl/, and point a step at its entry:
steps:
- ref: admin.search # one entry of hurl/admin.hurl
bind: { q: "${term}" }
That file stays valid hurl, so the same bytes run under stock hurl and under proef —
useful when a corpus already exists and you would rather annotate it once than transcribe
it. Choose by capability, not taste: inline splices text (so ${docstring} can carry a
multi-line body, which no binding can express), while ref: buys a name, reuse, standalone
runnability, and a static check that every variable is supplied. AUTHORING.md §“hurl: or
ref:” has the full comparison.
4. Validate without a network
$ proef test suite --dry-run
This binds every sentence, resolves every ${…}, emits the artifacts, and
parse-validates them — no request is sent. Typos in prose, packs, or hurl
blocks all fail here, with the file, line, and a “did you mean” where one
exists. Wire your editor too: proef schema --add-to suite/packs/api.yaml
gives packs autocomplete via the JSON Schema.
5. Provide the secret and a target, then run
Point PROEF_BASE_URL at any HTTP API you can reach — or start proef’s own
dev fixture in a second terminal: from a checkout, cargo run -p xtask -- fixture binds the default base port (8787), so no PROEF_BASE_URL is needed
(it prints a PROEF_BASE_URL line to export only when it ends up somewhere
other than 8787 — the port was busy, or you passed a different
... -- fixture <port>). It also prints the fixture’s own token line:
export PROEF_SECRET_APITOKEN=fixture-token — the value the next step needs,
because every /api/v1/ route rejects anything else with a 401. Then:
$ proef secret set apiToken # paste fixture-token; or: export PROEF_SECRET_APITOKEN=fixture-token
$ proef test suite
running 1 scenario(s) with 8 job(s) — run 019f…
Scenario: A known record is found (suite/case.feature)
✓ suite/case.feature:3 — the service is healthy (2ms)
✓ suite/case.feature:4 — the operator searches for "Acme" (5ms)
✓ suite/case.feature:5 — the first hit is record "r-1" (0ms)
✓ scenario A known record is found
summary: 1 passed · 0 failed · 0 skipped
Secret values never appear anywhere — not in artifacts, events, logs, or reports.
6. When it fails
A failing assert names the feature line, the artifact line, and hands you a
reproduce command (absolute paths; captures ride along in a .vars file):
✗ suite/case.feature:5 — assert failure (artifact case--a-known-record-is-found.hurl:12)
reproduce: hurl --test /abs/path/.proef-runs/<run-id>/artifacts/case--a-known-record-is-found.hurl --variables-file /abs/path/….vars
Because this suite binds ${secret:apiToken}, the replay also needs the
secret — proef never writes its value anywhere, so add it yourself:
--secret apiToken=fixture-token (the artifact’s # replay: header says so).
Every run leaves a record under .proef-runs/<run-id>/: events.jsonl (the
machine-readable event stream), run.log (the console mirror), and
artifacts/ — the exact .hurl files that were executed, byte for byte.
proef explain summarizes the latest run from the record, proef diff
compares two runs — surfacing what regressed, what got fixed, and which steps
turned flaky or slower — and proef report writes a self-contained HTML page
of a run (scenario tree, timings, and deep-links to the executed artifacts).
Where next
AUTHORING.md— the full pack reference: composition withuse:, retries, optional steps, guards, captures across scenarios, fakes.TROUBLESHOOTING.md— exit codes, the glyph legend, and the frequent failures;DIAGNOSTICS.mdindexes every error code you might hit.proef flows suitelists every scenario with tags;--format jsonfeeds the nextest harness (one IDE test per scenario).proef test suite --watchreruns on every file change.- Tag a scenario
@skip/@skip:reasonto park it — still counted, its reason in every report — or@quarantineso a flaky one stops gating CI while you fix it (AUTHORING.md · Reserved tags). - Exit codes are a contract:
0pass ·1test failure ·2your input ·3environment — safe to wire straight into CI, with--junitfor reports.
Writing scenarios
For the person who writes tests, not the person who wires them up.
You describe what should happen, in sentences. Somebody on your team maintains the vocabulary — the list of sentences proef understands — in files called packs. You never have to open one. This page is your whole loop.
If you also maintain the vocabulary, you want AUTHORING.md instead; this page deliberately stops at the boundary.
1. A test is sentences in a file
A test lives in a .feature file and looks like this:
Feature: Directory search
Scenario: A known record is found
Given the service is healthy
When the operator searches for "Acme"
- Feature — what area you are testing. One per file.
- Scenario — one situation worth checking. A file can hold many.
- Given / When / Then — the steps, in order. proef strips the opening word and matches only the rest of the sentence, so all of them behave identically; pick whichever reads best. Given for setup, When for the action, Then for the check. And and But work too, and read better than repeating yourself:
When the operator searches for "Acme"
Then the response status is 200
And the value at "$.results" is a non-empty list
Each step is one sentence, and every sentence must be one the vocabulary knows. That is the only rule you have to satisfy.
2. See the sentences you may write
$ proef macros
builtin:core.yaml
expectPresent 0× the value at {path} is present (builtin, unused here)
expectStatus 0× the response status is {status} (builtin, unused here)
…
suite/packs/api.yaml
health 1× the service is healthy
search 1× the operator searches for {term}
10 macro(s) · 0 unused
The right-hand column is what you can say. The left is the internal name — you
do not need it. The 1× counts how many steps across the suite use that
sentence.
{term} is a blank you fill in. the operator searches for {term} means
you write:
When the operator searches for "Acme"
Quotes are the house style and make the value easy to see.
This command works even when your file has a mistake in it — which is exactly
when you need it. If some step does not bind, you still get the list; only the
1× counts are withheld (they would be misleading), and they show as —.
Editor tip. If someone has set up
proef lspfor your editor, these sentences appear as autocomplete while you type, and mistakes underline as you go. It is worth asking for — but nothing on this page needs it.
3. Check before you run
$ proef test --dry-run
This reads your file and tells you whether every sentence binds. It sends no requests and touches nothing, so it is always safe and takes well under a second. Run it constantly.
ok suite/case.feature — 1 scenario(s), 2 step(s), 1 batch(es)
dry-run OK: 1 feature(s), 1 scenario(s), 2 step(s) …
next: proef test
When it says next: proef test, your sentences are good.
4. Run it
$ proef test
Scenario: A known record is found (suite/case.feature)
✓ suite/case.feature:3 — the service is healthy (6ms)
✗ suite/case.feature:4 — the operator searches for "Acme" (0ms)
✓ passed · ✗ failed · ∅ skipped, because an earlier step in the same
scenario failed.
A ✗ here is a real result — proef reached the system and the answer was
wrong. That is the test doing its job. Take it to whoever owns the API, with
the line it names.
5. The two mistakes you will actually make
Everything else is someone else’s problem. These two are yours.
“no macro matches …”
proef::bind::unbound_step
× no macro matches `the operator serches for "Acme"` — did you mean
`the operator searches for {term}`?
╭─[suite/case.feature:4:5]
4 │ When the operator serches for "Acme"
· ────────────────────────────────────
╰────
You wrote a sentence the vocabulary does not know — usually a typo, a plural, or a word order that drifted. proef points at the exact line and, when it can, guesses what you meant.
Your fix: say a sentence that exists. proef macros lists them.
The message also shows a macros: block. That is instructions for the person
who maintains the vocabulary — hand it to them if the sentence you need
genuinely does not exist yet. It is not something you need to write.
“url variable … is not set”
proef::resolve::missing_config_var
× in macro `health`: url variable `bse` is not set — define `[url]` `bse`
in proef.toml (or in the active `[env.<name>.url]`) — did you mean `base`?
A sentence you used needs an address that nobody has filled in. This one lives
in proef.toml, which is a short settings file, not a pack:
[url]
base = "https://api.your-company.com"
Your fix: if you recognise the name, correct the typo it suggests. Otherwise this is a setup question for whoever configured the project.
One more you may meet once
If a brand-new project fails on its very first run and says it is still the
proef init scaffold, nothing is broken — the starter files point at a
placeholder address and placeholder routes on purpose. Somebody needs to point
[url] base at the real API and replace the example routes. After that, the
loop above is all yours.
6. The loop
read the sentences → proef macros
write a scenario → your .feature file
check it binds → proef test --dry-run
fix what did not → the two errors above
run it → proef test
That is the whole job. When you need something the vocabulary cannot say, that is a conversation with whoever maintains the packs — and AUTHORING.md is the page for them.
Where to go next
| You want to | Read |
|---|---|
| Understand a symbol, exit code, or failure | TROUBLESHOOTING.md |
| Get autocomplete and live error underlining | EDITORS.md |
| Maintain the vocabulary yourself | AUTHORING.md |
| Set up a project from scratch | GETTING-STARTED.md |
Authoring reference — packs and features from the author’s seat
Everything here is validated at load or at --dry-run time; nothing fails
only at execution that could have failed earlier. Start with
GETTING-STARTED if this is your first suite.
Feature files
Standard Gherkin: Feature:, Scenario:, Background: (prepended to every
scenario), Rule:, Scenario Outline: + Examples: (expanded, #N-deduped —
note the #N is positional: inserting an Examples row above renames every
instance below it, which re-keys their JUnit history and re-buckets them
across --shard. A column placeholder in the outline’s name
(Scenario Outline: search finds <q>) keeps each instance’s identity tied to
its data instead of its row number),
data tables (rows become step arguments), and docstrings (delivered to the
macro as the docstring param). Keywords (Given/When/Then/And) don’t affect
binding — only the sentence text does. Prose between the Feature: line and
the first scenario is the feature’s description: never executed, but
proef flows prints it (JSON: featureDescription) so a file’s intent
travels with its inventory.
An outline’s <column> placeholders substitute into the docstring too, not
just the scenario name, step text and table cells. That is how a request body
gets data-driven without leaving the feature file:
Scenario Outline: Posting <label>
When a record is posted
"""
{"label": "<label>", "priority": "<priority>"}
"""
Then the response status is 201
Examples:
| label | priority |
| alpha | high |
| beta | low |
A <name> that is not an Examples column is a parse-time error wherever it
appears, docstrings included.
Variables are declared in proef.toml ([url] / [vars]), never in the
feature file, and referenced from packs as ${url:key} / ${vars:key} — see
CONFIG.md. Feature files stay free of URLs and environment data.
Tags (@smoke) accumulate feature→scenario; proef test --tags <expr>
selects scenarios by a boolean expression over them — and, or, not, and
parentheses, with the @ optional (e.g. --tags "@api and not @slow" or
--tags "(smoke or nightly) and not wip"). A bare tag is a valid expression; a
selection matching nothing is an error, not a silent green run. Atoms may
glob: * matches any run of characters, ? exactly one — anchored to the
whole tag and case-sensitive — so --tags "JIRA-*" selects every
ticket-tagged scenario, while a metachar-free atom stays plain equality.
Reserved tags
Two tag names carry behavior; every other tag is yours (selection via
--tags, grouping, traceability):
-
@quarantine— the scenario runs and reports, but a test-failure does not gate the exit code, reaches JUnit as<skipped message="quarantined failure (non-gating): …">, and maps to# TODOin TAP. For flaky tests while they are being fixed — a User/System fault still fails the run.Quarantine is a holding pen, not a destination, and
proef flakyis what keeps it one: over the retained records it separates a quarantined scenario that fails every run (DISABLED— switched off, and nobody is watching those failures because by design nothing reports them) from one that has been green throughout (recovered— the tag outlived the problem and is now suppressing the next real regression). Neither is visible any other way. -
@skip/@skip:<reason-token>— the scenario is parked: never prepared or run, counted as skipped, and the reason (the tag spelling itself) appears in the console, JUnit, TAP, the record, the report,explainandflows.--tags "not @skip*"unselects both spellings when you want them gone from the totals too. All-skipped exits 0;--dry-runstill validates a skipped scenario (skip is not a validation waiver). In[run] setup/teardownfeatures, reserved tags have no effect.
A tag that is almost reserved — @quarantined, @skipped, @Skip — is an
ordinary tag with no effect, and since 0.18 it warns (tags::reserved_tag_typo)
with the spelling it likely meant, because a scenario its author believed
quarantined would otherwise gate the build in silence.
Conditional, data-dependent skipping is a step concern and stays in packs:
when: guards a step, optional: soft-fails one, retry: bounds one.
That split is deliberate and permanent — scenario prose stays declarative;
there is no scenario-level IF/WHILE/TRY and none is planned (ADR-0019).
Macros (macros: in a pack)
macros:
name:
match: the record {name} is resolved # sentence pattern (optional)
params: [name] # declared parameters
defaults: { index: records } # defaults for optional params
description: One line for humans.
tags: [Admin]
steps: [...] # OR expect: [...] — never both
match:binds prose. Patterns need at least one literal word (no capture-only patterns), captures are{name}, quoted arguments shed their quotes, and matching is leftmost with ambiguity rejected — two macros that could claim the same sentence fail pack load.- Every
{capture}must be a declared param;defaults:keys must be declared params; adjacent captures ({a} {b}with nothing between) are rejected. - A macro without
match:is composition-only (reachable viause:). proef macroslists every macro with itsmatch:prose — the sentence a feature file may say — and its call count across the corpus, flagging pattern macros no scenario binds (dead prose bindings);use:-only helpers and unused builtins are listed but not flagged. It also flags near-duplicate pattern macros — two that differ only in their{capture}names (the same literal skeleton), which are confusable to authors. Both are advisory only: they never change the exit code, and--format jsoncarriespattern,unusedandnearDuplicateOffields for a CI hygiene gate.- When a step does not bind,
macrosstill lists the vocabulary (that is when you most need it) and keeps exit 2. Counts are withheld, not zeroed:calls/unusedrender as—/null, because an unbound feature contributes no calls and would make its own macros look dead.proef flowsstill refuses — it promises every scenario, and a silently partial list is a wrong answer.
Steps
steps:
- name: human label (${…} resolves here too)
optional: true # failure warns instead of failing
retry: { count: 10, interval_ms: 300 } # finite; 1..=10000
delay: 250 # ms before the request (capped at 1 hour)
when: "${env:RUN_SLOW:-}" # skips when empty or false/0 after resolution
saveAs: { recordId: global } # promote a capture to the global store
hurl: | # the payload — raw hurl, one or more entries
GET ${url:base}/api/v1/records/{{recordId}}
HTTP 200
- use: otherMacro # composition (cycle-checked, depth ≤ 32)
with: { term: "${name}" }
retry:/delay: are baked into the entry’s [Options] so the emitted
artifact replays with identical semantics under stock hurl. optional: steps
run as their own batch so a failure cannot poison neighbours. Raw [Options]
written inside a hurl: block are linted by the same rules — retry:/repeat:
finite and at most 10 000, delay:/retry-interval:/max-time: at most one
hour, and a key declared both in YAML and in the block is refused
(pack::option_declared_twice) — and whatever the values, a batch’s watchdog
budget never exceeds four hours (ADR-0007).
expect: macros carry no requests: status: 200 and/or raw hurl:
assert lines merge into the previous request entry (a Then before any
When is an error).
hurl: or ref: — two body forms, chosen by capability
A step’s body is an inline hurl: block or a ref: naming an entry in a
real .hurl file (ADR-0018). They are not two spellings of one thing, so the
choice is not a matter of taste:
hurl: | | ref: name | |
|---|---|---|
| variables | ${…} spliced before hurl parses | {{…}} bound via bind: |
| can substitute | anything, anywhere — including a whole multi-line docstring body | only what hurl can template |
| reuse | none: the block has no name | any number of macros, each binding differently |
runs under stock hurl | no | yes, unchanged |
| unknown variable | caught when the artifact is parsed | caught at --dry-run, by name |
Reach for inline for a request only this macro makes, and always when you
need to splice something hurl cannot template — ${docstring} as a request
body is the clearest case, since a bound value is a single-line scalar.
Reach for ref: when the same request serves several macros, when the hurl
was written by somebody else, or when the file must stay runnable on its own.
bind: # pack scope — every macro in this file
base: ${url:base}
apiToken: ${secret:apiToken} # injected at run time, never into an artifact
macros:
search:
params: [q]
match: "the operator searches for {q}"
bind: { q: "${q}" } # macro scope
steps:
- ref: admin.search
bind: { index: records } # step scope — the most specific wins
hurl: and ref: are chosen per step, not per suite. A step is one or the
other, but a macro mixes them freely — which is what adopting an existing corpus
looks like in practice: ref: the requests the corpus already has, write inline
for the ones it doesn’t.
archiveFirstResult:
match: the operator archives the first result
steps:
- ref: admin.search # the corpus already has this request
- hurl: | # this one is new, and splices ${…}
POST ${url:base}/api/v1/admin/records/{{recordId}}/archive
HTTP 204
recordId is captured by the fragment and read by the inline step: the World
threads captures across both forms, and contiguous same-engine steps batch
together whichever form they were written in. Pick per step by capability —
inline splices ${…} anywhere (including a multi-line ${docstring} body, which
a single-line bound scalar cannot express); ref: keeps the file runnable on its
own and checks its interface by name.
proef fragments lists the corpus: which entries exist, how many scenarios run
each, and — the two questions a listing exists for — which are annotated but
reached by nothing, and which carry no # @proef at all and so cannot be
referenced. --check exits 1 on the first; add --require-annotated to include
the second, which is opt-in because an unannotated entry is inert by design
(pointing at a corpus you did not write costs nothing), and only a team
mid-port means “not done yet” by it.
Set [run] fragments to the directory holding those files (see
CONFIG.md). Every {{variable}} a fragment reads must be bound in
one of the three scopes, captured by an earlier step, or supplied by the fragment
itself; nothing is implicit, because hurl’s per-entry variable: assigns into one
shared set rather than scoping, so an unbound name would quietly inherit an
earlier entry’s value.
A fragment supplies its own value with an ordinary [Options] variable: line —
which is how a corpus file stays runnable on its own, with fewer variables to
pass in:
# @proef admin.search
GET {{base}}/api/v1/admin/search/{{index}}
[Options]
variable: index=records # the file answers its own question
HTTP 200
Do not then also bind: that name. Both spellings reach the entry as
variable: index=, hurl takes the last, and the fragment’s own line is last — so
the bound value would never reach the request. Proef refuses the pair
(pack::option_declared_twice) rather than picking one silently; delete whichever
is not authoritative.
Bindings resolve once per scope instantiation — pack scope once per scenario,
macro scope once per invocation, step scope per step — so one bind: entry is one
value, and two are two. That is what makes a pack-scope ${fake:email} a single
identity for the whole scenario.
A binding nothing can read is refused rather than dropped, at two levels:
- The table — a
bind:with noref:in scope to read it (proef::pack::bind_without_ref), checked at all three scopes. - One key — a
bind:entry no fragment in that scope reads (proef::pack::unread_bind_key), with did-you-mean over the names that are read. This is the one a typo produces:bind: { token: …, toekn: … }binds one real key and one that never arrives.
The key check is a union over the scope, never against a single fragment — a pack-scope table is the plumbing every macro in the file needs, so a key serving one macro and not its siblings is correct usage.
Note the one that surprises people: a macro-scope bind: does not reach a
use: target — the target resolves its own pack and macro scopes — so the table
belongs on the macro that actually carries the ref:.
You do not have to memorise a corpus you did not write: with proef lsp running,
completing inside a bind: table offers the {{variables}} the fragments this
pack ref:s actually read, each labelled with the fragment that wants it. The
names are read off the .hurl file itself, so they cannot drift from it. If a
name still goes unsupplied, proef::lower::unbound_placeholder names it at lower
time — --dry-run is enough to surface that, no server needed.
The same check reads inside your bind values — with hurl’s own parser, so
what hurl calls a function is never mistaken for a variable ({{newUuid}} and
{{newDate}} need no supplier). A {{name}} that is a variable is templated
when the entry runs, so it must be supplied by then: an earlier step’s capture
(the usual shape: bind: { path: "records/{{recordId}}" }), the fragment’s own
[Options] variable: (its lines evaluate before the injected ones), a secret
in scope, or a sibling literal bind whose name sorts before this one — the
injected lines are written and evaluated in name order, so base can feed q
but not the other way around. And the reverse collision warns rather than
fails: a literal bind: that re-uses a name an earlier step captured wins
silently from that entry on (hurl’s variable: assigns into one shared set),
so proef::lower::bind_shadows_capture names it — rename the binding if the
capture was the point.
Where a file,…; asset lives
A file body — a request body, a multipart part, or a file,…; an assert
compares against — sits beside the source that names it:
| The step is | The file goes | Because |
|---|---|---|
hurl: | (inline) | beside the feature | the suite is one unit; tests/features/fixture.jpg serves tests/features/*.feature |
ref: <name> | beside the fragment | that is where stock hurl looks, so the file keeps running on its own |
This is hurl’s own rule (its --file-root defaults to the .hurl file’s
directory), and the one Karate, pytest and Jest use for fixtures. You never
set a root: proef stages each asset from its own source into the run’s
artifacts/assets/<slug>/, and the emitted artifact carries the matching
--file-root in its replay line, so the hand-off runs unchanged under stock
hurl.
Two consequences worth knowing. The path must be a plain relative one — no
leading /, no .. — because it names a file inside your suite, not a
location on the machine. And a missing file is refused at --dry-run,
naming the directory it was looked for in (proef::run::asset_unstageable) —
so the gate CI runs before standing an environment up answers it, rather than
a failing request minutes later, or hurl reporting an unreadable body against
the artifact. Staging resolves beside the file the
parser actually read, so the directory you run from does not matter; a symlink
already at a staging destination is replaced, never written through; and two
references that name one file on a case-insensitive filesystem (Data.json and
data.json on macOS or Windows) are refused (proef::run::asset_unstageable)
rather than silently last-writer-won.
Recipes — the three shapes every real suite needs
Log in, then use the token
The most common real-world flow: the API mints a token at a login endpoint,
and every later request carries it. A capture crosses steps through the
World, so the shape is one login macro capturing the token and any number
of authed macros reading it:
macros:
logIn:
match: the operator logs in
steps:
- hurl: |
POST ${url:base}/auth/login
{"user": "${vars:user}", "password": "${secret:password}"}
HTTP 200
[Captures]
token: jsonpath "$.token"
listRecords:
match: the records are listed
steps:
- hurl: |
GET ${url:base}/records
Authorization: Bearer {{token}}
HTTP 200
{{token}} is hurl run-time vocabulary (the second tier): the capture is
assigned when the login entry runs, and every later entry in the scenario
reads it — no bind:, no config. To reuse one login across scenarios, add
saveAs: { token: global } to the login step and read ${global:token};
note the promotion is refused if the value carries a secret (ADR-0005).
Wait for an eventually-consistent result
A POST answers 202 and the resource appears a moment later. Don’t sleep —
put a finite retry: on the step that polls, and let its asserts be the
condition:
awaitRecord:
params: [id]
match: record {id} is eventually visible
steps:
- retry: { count: 10, interval_ms: 300 }
hurl: |
GET ${url:base}/records/${id}
HTTP 200
The retry re-runs the entry until its asserts pass or the count runs out —
and the count must be finite: hurl cannot be interrupted mid-call, so an
unbounded retry is a hang the watchdog would have to abandon (the same
reason retry: -1 is refused at load). The retry also bakes into the
emitted artifact’s [Options], so a replay under stock hurl polls the
same way.
Seed data before, clean up after
Three scopes, three mechanisms — pick by lifetime:
- Per scenario: a
Background:in the feature file runs its steps before every scenario in that file — provisioning prose, bound to macros like any other step. - Per suite:
[run] setup/[run] teardowninproef.tomlname feature files that run once around the whole pool — seed a database, then delete the run’s residue. Teardown runs even when the suite fails or is interrupted; a teardown failure is exit 3, never silent. The keys are documented in Configuration. - Across scenarios: a
saveAs: { id: global }capture in setup (or any scenario) persists into later scenarios and runs as${global:id}— the handle teardown needs to delete what setup created.
Asserting responses — the hurl vocabulary
Assertions live inside a step’s raw hurl: block (or an expect: macro), so the
whole hurl 8.0 assert grammar is available untouched: proef resolves ${…}
before the run and hands the rest to the embedded engine verbatim (ADR-0005). The
block is parsed at pack-load time, so a grammar slip fails fast with a diagnostic
— only JSONPath semantics (a query that lexes but never matches at run time) can
slip through. An assert reads <query> [filters…] <predicate>; HTTP <status> is
the implicit status assert. The authoritative list is hurl’s own
asserting-response and
filters docs; the common shape:
GET ${url:record}
Authorization: Bearer ${secret:apiToken}
HTTP 200
[Asserts]
jsonpath "$.id" isUuid # type/shape checks — schema-lite
jsonpath "$.status" == "active"
jsonpath "$.createdAt" isIsoDate
jsonpath "$.tags" count == 3
header "Content-Type" contains "application/json"
duration < 1000 # response-time budget (ms)
- Queries (what to read):
status,header "<n>",cookie "<n>",body,bytes,jsonpath "<expr>",xpath "<expr>",regex "<pat>",sha256,md5,url,redirects,variable "<n>",duration,certificate "<f>". - Predicates (the check):
== != > >= < <=,startsWith,endsWith,contains,includes,matches "<regex>",exists/not exists,isEmpty, and the type familyisBooleanisIntegerisFloatisNumberisStringisCollectionisListisObjectisDateisIsoDateisUuidisIpv4isIpv6. - Filters transform the value before the predicate, chained left→right:
count,nth <n>,first,last,split "<sep>",replace/replaceRegex,toInt/toFloat/toString,toDate "<fmt>",format "<fmt>",base64Decode/base64Encode,urlDecode/urlEncode,daysAfterNow/daysBeforeNow,jsonpath "<expr>",regex "<pat>",utf8Decode.
[Asserts]
jsonpath "$.items" count > 0
header "Set-Cookie" split ";" nth 0 startsWith "session="
RFC 9535 JSONPath (hurl 8.0): filter expressions and functions are in scope —
jsonpath "$.books[?(@.price < 10)]", length(), count(), match(),
search(). Use the bracket form for names with hyphens ($['x-custom-id']), and
assert a missing path with not exists (a non-matching path yields no value,
not count == 0). The same query+filter grammar drives [Captures], threading a
value into later steps as {{name}}.
Built-in shape macros
For the most common single-value shape checks, the built-in Core pack ships a
small, product-neutral set of expect: macros so you rarely hand-write the
predicate. Each reads a JSONPath ({path}) into the previous response and merges
one assert:
When the record is fetched
Then the value at "$.id" is a uuid
And the value at "$.name" is a string
And the value at "$.tags" is a non-empty list
Available: the value at {path} is a string / … a number / … a boolean /
… a uuid / … an ISO date / … present / … a non-empty list. Quote the path
in prose when it contains spaces (the quotes are optional and shed). These are a
convenience layer over the predicates above — they deliberately do not cover
whole-body structural matching; reach for a raw expect: hurl: block for
anything they omit.
Negative cases — one macro per malformation, one shared expectation
A validation suite is naturally combinatorial, and its cases differ
structurally rather than by value: one omits a key, one empties it, one adds a
key the caller may not set. A single parameterised macro cannot express that —
the bodies are different shapes, not one shape with a hole in it. Outline
placeholders do substitute into docstrings, but an Examples cell cannot
practically hold JSON, and a raw body in the feature file defeats the prose the
design exists to protect.
Name each malformation. One macro per case, each sentence saying what is wrong in business terms:
macros:
createWithEmptyTitle:
match: creating a task with an empty title is refused
steps:
- hurl: |
POST ${url:base}/tasks
Content-Type: application/json
{"title": "", "priority": "high"}
createWithServerOwnedField:
match: creating a task that sets its own id is refused
steps:
- hurl: |
POST ${url:base}/tasks
Content-Type: application/json
{"title": "ok", "id": "caller-chosen"}
Then let one expect: macro serve the whole catalogue. Because an expect:
merges its asserts into the previous request entry, a single parameterised
expectation covers every case in the set — typically the largest de-duplicator
in a validation pack:
expectErrorCode:
params: [code]
match: "the error code is {code}"
expect:
- hurl: |
jsonpath "$.error.code" == "${code}"
The scenarios then read as the specification they are, and each one’s asserts land on its own request:
Scenario: An empty title is rejected
When creating a task with an empty title is refused
Then the response status is 422
And the error code is TITLE_REQUIRED
# emitted
POST http://127.0.0.1:8787/tasks
Content-Type: application/json
{"title": "", "priority": "high"}
HTTP *
[Asserts]
status == 422
jsonpath "$.error.code" == "TITLE_REQUIRED"
The trade, stated plainly: the pack grows with the malformation catalogue — one macro per case, where a value-driven test would have used one row. That is the cost of feature files that read as prose, and it buys a suite a non-engineer can review. The expectation side does not grow with it.
Variables — the two tiers
| Syntax | Resolves | When | Examples |
|---|---|---|---|
${…} | proef | at lowering (before execution) | ${param}, ${env:NAME:-default}, ${url:key}, ${vars:key}, ${run:id}, ${global:key}, ${secret:NAME}, ${fake:name} |
{{…}} | hurl | at run time | captures ({{recordId}}), secrets ({{apiToken}}) |
${…} is recursive (captured arguments may themselves contain ${…}, depth
≤ 8) and $${ escapes a literal ${. ${secret:NAME} never inlines the
value — it lowers to {{NAME}} and the engine injects it through hurl’s
redaction, so artifacts carry placeholders only.
Config variables — ${url:key} and ${vars:key} come from proef.toml’s
[url] / [vars] tables, deep-merged with the active --env profile. This is how
you keep URLs and settings out of the feature files entirely; a referenced-but-undefined
one is a lower-time error. See CONFIG.md (ADR-0012).
Captures ([Captures] in a hurl block) flow forward within the scenario
as {{name}}. saveAs: { name: global } additionally promotes the captured
value into the persistent global store (.proef-state.json), where later
scenarios and later runs read it at lowering time as ${global:name}.
Fakes are deterministic synthetic data: firstName, lastName, name,
fullName, email, username, phoneNL, postCode, city, street,
int, number, digits4, digits8, bool, word, uuid (unknown
generators fail at load). Each ${fake:kind} reference gets its own
value — an occurrence counter advances every time a scenario resolves one, so
independent ${fake:email} references in the same scenario never collide,
however many a step ends up resolving. A step’s name: label is the one
deliberate exception: it is not independent of its own payload, so it
replays from the start of the step’s own occurrence window instead of
minting new ones — the label’s Nth ${fake:…} reference reuses whichever
occurrence the payload’s (and when:’s) Nth reference consumed, matched by
position, not by generator kind. When the label’s ${fake:…} references
mirror the payload’s in kind and order — the common case, e.g. a label that
names the same field the payload sends — this reproduces the payload’s own
value exactly. When they diverge in kind (say the label’s first reference is
${fake:fullName} but the payload’s first reference is ${fake:email}),
the label instead shows whatever that occurrence generated for its own
kind — a value the request did not send. That includes a label
with more ${fake:…} references than its payload: each extra one still
reserves its own place in the sequence, so a later step can never be handed
a value the label already displayed. That whole sequence is a pure function
of ${run:id}: the same --run-id reproduces the same fakes, byte for
byte, across runs. The counter restarts at zero for every scenario —
known limitation: two different scenarios that each resolve
${fake:email} at the same position in their own step order (typically each
scenario’s first fake reference) get the same address, because both count
from zero independently. If two scenarios must not collide, key the value
yourself (fold in ${run:id} or a captured id) rather than relying on
${fake:*} alone.
Secrets
Reference with ${secret:NAME}. Values come from PROEF_SECRET_<NAME>
environment variables first, then the encrypted store (proef secret set NAME,
or … | proef secret set NAME --stdin for scripts (never argv — ps shows
it); proef secret list names,
proef secret rm NAME removes
— XChaCha20-Poly1305, key auto-created 0600 under ~/.config/proef/).
Values never appear in artifacts, events, logs, or reports; events carry
capture names only. The store file (.proef-secrets.json, mode 0600)
holds ciphertext only and is gitignored by default; the key
(~/.config/proef/keys/default.key) must never leave your machine.
Artifacts — the executed input
Every scenario emits <feature-path>--<scenario>.hurl — the feature’s
suite-relative path with its extension dropped, then the scenario, both
slugified (so two same-named features in different directories cannot collide)
and capped at 120 bytes with a hash tail — the exact bytes the engine executes
— plus .map.json (artifact lines ↔ feature lines, batch and
step indices) and .vars (referenced globals as values, secrets as names).
The header’s # replay: line is a complete stock-hurl command, including
--secret NAME=<value> placeholders for you to fill. proef artifacts <dir> -o out/ --run-id ci emits the same set deterministically for hand-off.
Runs, records, CI
- Exit codes:
0pass ·1test failure (including cancelled runs) ·2input error ·3system error. .proef-runs/<run-id>/events.jsonlis the record — one JSON event per line (scenario/step lifecycle, per-attempt progress, failuredetail);proef explainsummarizes it;--junitand the GitHub Actions summary derive from the same run.--format jsonprints one machine-readable summary object on stdout (the human report moves to stderr);--jobs Ncontrols parallelism;proef.tomlholds project defaults ([run] jobs,[http] timeout-ms).proef fmt suite --checkkeeps hurl blocks canonically formatted;proef flowslists scenarios.
IDE integration — one test per scenario
The harness crate bridges proef into any nextest/libtest UI (rust-analyzer,
RustRover, cargo nextest run): it lists scenarios via
proef flows --format json and runs each as its own test via
proef test --scenario <name>.
PROEF_HARNESS_SUITE=suite cargo nextest run -p proef-harness
Set PROEF_BIN=/path/to/proef when the binary is not on PATH. With
PROEF_HARNESS_SUITE unset the harness exposes nothing, so a plain
cargo test stays green. Each scenario appears as a separate test in the
IDE’s runner — click-to-run one scenario without touching the terminal.
Unset is not the same as unreadable. A variable set to bytes that are not
valid UTF-8 means you asked for something the harness cannot read, so it
exposes a single failing proef::config trial naming the variable — it will
not fall back to proef on PATH, and it will not report green having
listed no tests.
Secrets in CI
Two working setups:
- Values via env (simplest): set
PROEF_SECRET_<NAME>from your CI’s secret storage — no store, no key, nothing on disk. - Committed ciphertext store: commit
.proef-secrets.json(it holds ciphertext only) and supply the project key as thePROEF_KEYCI secret (base64 < ~/.config/proef/keys/default.key). A set-but-invalidPROEF_KEYis always an error, never a silent fallthrough.
Note: .proef-state.json (the persistent World) is plaintext and therefore
gitignored — captures promoted with saveAs: global land there. A capture
whose value equals a known secret is refused with a warning; it never
persists (proef doctor also reports store/key health).
Environment variables
| Variable | Read by | Purpose |
|---|---|---|
PROEF_ENV | proef test/flows/macros/artifacts/lsp | Active environment profile — the --env flag wins over it |
PROEF_SECRET_<NAME> | proef test | Secret value override (beats the encrypted store) |
PROEF_KEY | proef test/secret | Base64 project key override — decrypt a committed store without the key file |
PROEF_CONFIG_DIR | proef secret | Key-file location (default: XDG config dir) |
PROEF_HARNESS_SUITE | nextest harness | Suite directory the harness lists and runs |
PROEF_BIN | nextest harness | Path to the proef binary the harness invokes |
Set-but-unreadable is an error, never silence. A variable whose value is
not valid UTF-8 is something you asked for and proef cannot read, so it is
reported rather than treated as unset: the commands above exit 2 naming the
variable (proef doctor reports it as a failed check and exits 3, with its
other environment findings), and the harness exposes a failing proef::config
trial. Reading a malformed PROEF_KEY as absent used to mean decrypting with
the wrong key and reporting tampering; a malformed PROEF_ENV meant running
against the wrong environment. PROEF_CONFIG_DIR is a path and is read as raw
bytes, so it has no such failure mode.
Suite-defined env vars (like PROEF_BASE_URL in the guides) are a convention
of ${env:…} references in proef.toml, not built-ins — name yours freely.
Style
Keep prose at business level — no URLs, headers, or JSON in feature files;
that’s what packs are for. Prefer several small macros composed with use:
over one large one. Give steps name: labels — they anchor artifacts,
events, and failure output.
Label a macro’s steps whenever it has more than one. A macro with three steps
turns one feature sentence into three engine steps that share a step anchor
exactly — same file, same line, same text — so on the console, in the HTML
report, in JUnit, TAP and the job summary they arrive as the same sentence
repeated, one row per step. The name: is the only thing that distinguishes
them:
✓ tests/features/a.feature:9 — the workspace is provisioned › fixture warm-up probe (4ms)
✗ tests/features/a.feature:9 — the workspace is provisioned › provision the environment (7ms)
Without the labels those two lines differ only by their glyph.
proef in CI
Everything proef’s differentiators buy — typed exit codes, JUnit, sharding, the rerun overlay, the regression gate — pays off in CI, and each piece is documented on its own page. This page is the missing last mile: one paste-ready workflow, then the pieces it composes.
A complete GitHub Actions workflow
name: e2e
on: [push, pull_request]
jobs:
e2e:
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
shard: [1, 2, 3]
steps:
- uses: actions/checkout@v4
- name: Install proef
run: |
curl -LsSf https://github.com/cargo-bins/cargo-binstall/releases/latest/download/cargo-binstall-x86_64-unknown-linux-musl.tgz | tar -xz -C /usr/local/bin
cargo-binstall proef --no-confirm
- name: Run the suite
env:
# Secrets reach proef only through the environment — never files,
# never flags (values would land in the process listing).
PROEF_SECRET_APITOKEN: ${{ secrets.API_TOKEN }}
run: |
proef test tests/features \
--env staging \
--shard ${{ matrix.shard }}/3 \
--junit auto \
--meta commit=${{ github.sha }}
The pieces:
--junit autowritesreport.junit.xmlinto the run dir — underGITHUB_ACTIONSonly, so local runs stay clean. Failures also reach the job summary and the PR’s changed-files gutter as::errorannotations automatically; no extra step.--ctrf report.ctrf.jsonwrites the same verdicts as CTRF JSON — the format that carries whatJUnitXML cannot: retries and flakiness natively (a pass-after-retry lists its real failed attempts with messages), tags, and a file path per test. Point a CTRF consumer such asctrf-io/github-test-reporterat it. The two files never disagree: a quarantined failure is skipped with a message in both (ADR-0019), because a dashboard reading “failed” beside exit 0 would contradict itself. Like--junit, the file is written even when[run] setupaborts the run — a job gating on it must never see it missing.--shard I/Npartitions by a stable hash of(file, scenario): adding a scenario never re-buckets the others, so shard timings stay comparable across commits. Every matrix job runs the same expression and the shards partition exactly. To balance by time instead of by count, see below.--meta commit=…records provenance the run cannot harvest itself: proef never readsGITHUB_SHAor any CI variable (ADR-0020) — what the workflow hands over explicitly is what the record carries.PROEF_SECRET_<NAME>supplies${secret:name}values. They never appear in artifacts, events, logs, or reports — including base64/hex/ percent-encoded reflections.
Balancing the matrix by duration
A hash split balances by count, and a matrix finishes when its slowest shard
does — so one long scenario can leave three runners idle. Every run that reaches
its suite writes a small timings.json into its run directory; archive it once
and hand it to the next run’s matrix:
- name: Run the suite
run: |
proef test tests/features \
--shard ${{ matrix.shard }}/3 --shard-weights timings.json
# one job publishes the file the next run's matrix reads
- uses: actions/upload-artifact@v4
if: matrix.shard == 1
with:
name: proef-timings
path: .proef-runs/*/timings.json
Every job must read the same file. The tempting shortcut — letting each job
weight by its own local run history — is silently wrong: matrix jobs run on
different machines with different (usually empty) runs-dirs, so each would
compute a different assignment and scenarios would run twice or not at all while
the suite still reported green.
A scenario the file does not mention falls back to the frozen hash, so a test added after the timings were captured still runs exactly once. A missing or malformed file is exit 2 — never a silent fall back to the unbalanced split. Full detail in CONFIG.md.
Gating on regressions between runs
proef diff compares two run records; --fail-on-regression makes it a
gate. Download the base branch’s record (uploaded as an artifact by its own
run) and compare:
proef test tests/features --junit auto # today's run
proef diff path/to/base-events.jsonl # vs the downloaded baseline
# exit 1 on a regression; new flakiness and perf deltas print either way
Continuing a cancelled run
A run stopped by --max-fail, a runner timeout, or a docker stop records what
it never reached. SIGTERM and SIGHUP take the same graceful path as Ctrl-C:
in-flight batches finish, the rest record as skipped, [run] teardown runs, the
reports are written, and the record closes with run_finished + cancelled
(exit 1) — so a timed-out job still leaves a complete record and its JUnit. Only
a second signal (exit 130) or a SIGKILL truncates it.
proef test --rerun re-runs the last run’s failures and the
scenarios it never got to — and its JUnit and HTML report cover the whole
suite via the rerun overlay (rerun_of in the record), so one report stands
for the composed result, never a false green.
Flakiness over history
Keep a few records ([run] keep-runs in proef.toml sets the rotation) and
proef flaky renders verdicts over them — flapping, passes-only-on-retry,
always-failing — from the same records CI already produced.
It also audits @quarantine itself, which nothing else can: a quarantined
scenario failing every run is DISABLED (switched off — its failures gate
nothing, so no job ever reports them), and one green throughout is recovered
(the tag can come off). --format json carries quarantined and the verdict
key, so a scheduled job can gate on either without parsing the table.
--by <key> splits the same history per run context — --by env, or any
[meta]/--meta key such as --by runner. A scenario that flaps in one
environment and is solid in another is not flaky but context-dependent, and a
pooled history cannot tell those apart: it reports the one conclusion the
merged view can never reach, naming the scenarios whose verdict changes with
where they ran. A run that never set the key is its own (unset) bucket rather
than being folded in with the runs that did.
Verdicts are keyed by the run’s input fingerprint — inputs.json beside
each record, a hash of the feature sources, the loaded macros and fragments,
and the resolved ${url:…}/${vars:…} scope — so a pack or config edit starts
a fresh window instead of mixing runs of different inputs; --by splits within
it. Three guards, each a [flaky] key with a flag twin (CONFIG.md): a
scenario seen in fewer than min-samples runs (default 10) reads
insufficient-data rather than earning a verdict — a fresh matrix needs ten
records before the table says anything, and --min-samples 2 restores the
pre-0.18 floor; a flagged scenario stays flagged until recovery-runs (5)
trailing clean runs, so it cannot flip between adjacent runs; and a run in
which more than outage-rate (0.8) of the suite failed is an environment
incident, excluded wholesale, so one staging outage cannot mark the suite
broken.
Editor setup — proef lsp
proef lsp starts a Language Server Protocol
server that speaks generic LSP over stdio. Point any LSP-capable editor at
proef lsp and get, live as you type:
- Diagnostics — the same validation
proef test --dry-runperforms (unbound steps, unknown step kinds, malformed hurl blocks, unresolved${…}references), published across the whole suite as you edit. - Go-to-definition — jump from a Gherkin step to the macro that binds it, or
from a pack’s
ref:to the# @proefannotation in the.hurlfile (ADR-0018). - Completion — step completions offering the suite’s macro patterns; on a
pack’s
ref:line, the fragment names in scope; and inside abind:table, the{{variables}}those fragments actually read, each labelled with the fragment that wants it. A variable a fragment supplies itself ([Options] variable:) is left out: it needs nobind:, and binding it would be refused asoption_declared_twice. - Find-references — every step across the suite that a given macro binds.
- Document symbols — a feature outlines to its scenarios (tagged ones showing their tags), a pack to its macros (detailed by the pattern each matches). Which vocabulary applies is decided by what discovery found in the file, not by its extension.
- Hover — the macro a step binds, or a
use:targets, with its pack, pattern and params; on aref:, the fragment’s file and the variables still needing abind:. Every fact comes from the same analysis the diagnostics do, so a hover can never contradict the squiggle on the same line. - Semantic tokens — the two variable tiers, told apart on screen. Inside a
pack’s
hurl: |block,${…}(resolved at lower time, by proef, before any request exists) and{{…}}(resolved at run time, by hurl) look identical to every editor: a YAML highlighter sees a string, and a hurl highlighter never runs because the block is not a file. proef is the only party that knows which is which.${…}is reported as a macro — a substitution performed before execution, which is what a macro is — and{{…}}as a variable; every mainstream theme colours those differently already, so nothing needs configuring. A$${escape stays unhighlighted, because it is a literal${proef will not substitute. - Quick fixes — a misspelled name that already earned a “did you mean”
becomes an applicable edit:
use:andref:targets,with:andbind:keys, step kinds, Examples placeholders, and data-table columns. A fix is offered only when it is certain — the suggested name is near enough, and the misspelling occurs exactly once as a whole token in that same file — so applying one is never a guess. Reach it from the squiggle or from the token itself; ause:error underlines the macro’s name key, which is often several lines above the word you typed.
The server analyzes the configured suite — proef.toml’s [run] suite if
set, else the tests/ convention — resolved under the directory it is launched
in (its working directory), discovering every .feature file and every
packs/*.yaml / packs/*.yml macro pack beneath that root — the same
resolution proef test uses, so the two never diverge. Launch your editor from
the project root (or configure the server’s root/working directory to it).
When analysis fails
A panic inside analysis or a feature never ends the server. proef reports it
once through window/showMessage — the channel an editor actually surfaces —
and keeps serving; the next edit retries, and reports again only if the new
state also fails. A server that died would show nothing, which reads as “proef
has no opinion about this file” rather than as the failure it is.
Running proef lsp by hand also prints the panic to stderr, which is where the
detail lives.
Naming the config: --config
proef lsp --config <path/to/proef.toml> names the config to read instead of
searching for one, exactly as it does for every other subcommand (CONFIG.md).
Reach for it when discovery cannot find the right file — most often a config
that sits beside the suite rather than above it, which an upward search
launched from the repository root can never reach.
For proef lsp the flag also outranks the workspace root the client announces:
the flag names a file, and a named file is not a guess to be improved on.
Without it, an editor rooted somewhere else loads a different config than the
runner, and the diagnostics stop being trustworthy in exactly the layout the
flag exists for — proef test --config … runs green while every ref: reads as
unknown in the editor.
cmd = { "proef", "lsp", "--config", "/abs/path/to/proef.toml" },
Unlike the runner, an unreadable or absent file does not stop the server: it starts on defaults, because an editor offering less is better than one that will not boot. A relative path works, but prefer an absolute one — an editor’s working directory is rarely the one you assume.
--env <name> travels the same way (proef lsp --env staging, or PROEF_ENV):
it selects the [env.<name>] profile the analysis resolves ${url:…}/${vars:…}
against, so the editor reports the missing_config_var a staging run would hit
and not the one the default profile would. Before 0.18 the flag was accepted for
lsp and silently dropped.
File types served
| Kind | Pattern | Typical editor filetype |
|---|---|---|
| Feature files | *.feature | gherkin / cucumber / feature |
| Macro packs | packs/*.yaml, packs/*.yml | yaml |
Neovim
Built-in LSP client (Neovim 0.8+), no plugin required. Add to your config and
open a .feature file from inside the suite:
vim.api.nvim_create_autocmd("FileType", {
pattern = { "cucumber", "gherkin", "yaml" },
callback = function(args)
vim.lsp.start({
name = "proef",
cmd = { "proef", "lsp" },
-- The server scopes analysis to the configured suite under its launch
-- directory; anchor root_dir at the nearest proef.toml (or the current
-- file's directory) so the two agree.
root_dir = vim.fs.dirname(
vim.fs.find({ "proef.toml" }, { upward = true, path = args.file })[1]
) or vim.fs.dirname(args.file),
})
end,
})
With nvim-lspconfig you can instead
register it as a custom server via vim.lsp.config/configs, using the same
cmd = { "proef", "lsp" }.
Helix
Add a server and attach it to the languages you author. In
~/.config/helix/languages.toml:
[language-server.proef]
command = "proef"
args = ["lsp"]
[[language]]
name = "gherkin"
language-servers = ["proef"]
# Attach to pack YAML too, so pack diagnostics surface while editing macros.
[[language]]
name = "yaml"
language-servers = ["proef", "yaml-language-server"]
Run hx --health gherkin to confirm Helix found the proef binary.
Emacs (Eglot)
Eglot ships with Emacs 29+. Associate your feature-file major mode (for example
feature-mode) with the
server:
(with-eval-after-load 'eglot
(add-to-list 'eglot-server-programs
'(feature-mode . ("proef" "lsp"))))
;; M-x eglot in a .feature buffer opened from the suite root.
For lsp-mode, register a stdio client whose connection is
(lsp-stdio-connection '("proef" "lsp")) for the same major mode.
v1 limitations
This is the first release of the language server. Known boundaries:
-
proef.tomlconfig is a startup snapshot.${url:…}/${vars:…}values are read once when the server starts. After editingproef.toml(or switchingPROEF_ENV), restart the server to pick up the change; until then those references analyze against the old (or, if the file could not be loaded at startup, an empty) scope and may warn. -
Built-in macros have no jump target. The
expect*family lives in a pack compiled into the binary, not a file on disk, so there is nothing for go-to-definition to open. Hover still answers — the macro is in the analysis like any other, and reports its pack asbuiltin:…, which is why the jump is unavailable.proef macroslists the whole family with the sentence each binds. -
Completion ranking is best-effort. All of the suite’s macros are offered; ranking is a lightweight edit-distance heuristic. Full context-aware ranking is a follow-up.
-
External edits to closed files need a reopen. The server does not watch the filesystem; it re-reads a file’s bytes from its open editor buffer. If a
.featureor pack file is changed outside the editor (or by another tool) while closed, reopen it so the server sees the new bytes. -
No VS Code extension yet. v1 is a server-only generic-LSP binary. It works with any editor that speaks generic LSP (Neovim, Helix, Emacs, Sublime LSP, …); a VS Code wrapper is a possible follow-up.
If one is built, its
documentSelectormust match on path ({ scheme: "file", pattern: "**/*.feature" }), not on a language id. The ecosystem is split — the two established Gherkin extensions registercucumberandfeaturerespectively — so a selector naming either id attaches for some users and silently does nothing for the rest. The table under File types served is that split; a path selector is the only thing all of it has in common. -
No Zed support. Zed binds language servers to languages it has a tree-sitter grammar for, and there is no Gherkin grammar in it. That grammar is a prerequisite, not a configuration step, so Zed waits on work outside this repository.
-
Overlay lookup can still miss if the suite root is reached through a symlink. The root is deliberately left uncanonicalized (canonicalizing would resolve symlinks and desync source names from the client’s document URIs), so if an editor resolves a symlinked suite root differently than the server’s raw working directory, the overlay lookup can miss and the LSP analyzes the saved on-disk bytes instead of the unsaved buffer. An editor’s percent-encoding choice no longer matters here — the overlay matches open buffers by decoded source name, not the raw URI.
Troubleshooting
First stop, always: proef doctor (native libraries, environment, secret
store/key health) and proef test <suite> --dry-run (full validation with
located diagnostics, no network). Every diagnostic code is indexed in
DIAGNOSTICS.md.
Reading the output
Exit codes are a contract:
| Code | Meaning | Typical fix |
|---|---|---|
0 | everything passed (warnings allowed) | — |
1 | at least one check failed: a test assertion, a cancelled run — or a --check-style gate (fmt --check, fragments --check, diff --fail-on-regression) that found what it gates on | fix the system under test — or the expectation |
2 | your input is at fault: packs, features, flags, filters, secrets, bad {{var}}/JSONPath | the diagnostic names the file and line |
3 | the environment or proef is at fault: unreachable target, native libs, IO, output proef could not write (full disk, failing device) | check the target, proef doctor, disk |
130 | interrupted twice — the second signal (Ctrl-C, SIGTERM or SIGHUP) is a hard exit (128 + SIGINT), so cleanup and the record’s tail are skipped | the run record will read as incomplete; a single Ctrl-C — or a single SIGTERM/SIGHUP, e.g. a CI job timeout — cancels gracefully and still runs [run] teardown |
Step glyphs:
| Glyph | Status | Meaning |
|---|---|---|
✓ | passed | ran, all asserts held |
✗ | failed | ran, an assert failed (details + reproduce: line at the end) |
∅ | skipped | not run: when: guard, an earlier failure, cancellation, or an authored @skip on the scenario (the tag spelling prints as the reason) |
⚠ | warned | an optional: step failed, or a saveAs: global promotion was refused — the reason prints on the ↳ line |
Under --console dotted the tree collapses to one glyph per scenario —
. passed, F failed, s skipped, w warned (lowercase = non-gating) —
and failures still print in full after the pool; --console quiet keeps only
the run line and the summary. On a terminal the status vocabulary is
colored; NO_COLOR (or a non-terminal stream) turns it off, and run.log
never carries the paint either way. The record and the exit code are
identical in every mode. Every run’s last line names its run id and
wall-clock — the id is the reproduction key --shard, --shuffle and
${fake:…} all hang off.
Frequent situations
“missing secret value(s)” — the suite references ${secret:NAME} with no
value available. Fix: proef secret set NAME (or export PROEF_SECRET_<NAME>=…). In CI, see
AUTHORING — Secrets in CI.
An assertion fails on values that look identical — they differ only in whitespace, and proef says so:
Assert failure (f--s.hurl:9: actual: string <ok> expected: string <ok >)
[whitespace only — actual: string·<ok>, expected: string·<ok·>]
The bracketed note appears only when the two values differ solely in
whitespace, and repeats them with every whitespace character drawn: · for a
space, \t / \r / \n for the usual escapes, and \u{a0} for the exotic
ones — a non-breaking space pasted out of a browser being the classic. Usual
causes: a trailing space in the expectation, a CRLF fixture leaking \r into a
body, or a tab where the author typed spaces.
“unknown environment <name>” — --env <name> (or PROEF_ENV) names an
environment proef.toml doesn’t define; the error lists the known ones. Fix the
name or add the [env.<name>] section. Exit 2.
“<url|vars> variable <key> is not set” (resolve::missing_config_var) — a
pack references ${url:key} / ${vars:key} that neither the base [url]/[vars]
nor the active [env.<name>] defines. Add it, or select the environment that has
it with --env. See CONFIG.md.
“no path given and no default suite found” — proef test got no path and there
is no [run] suite in proef.toml nor a tests/ directory. Pass a path, set
[run] suite, or create tests/. Exit 2.
Exit 3 with connection errors — the target is unreachable. Check the URL your
${url:base} resolves to for the active --env (a
--dry-run prints nothing wrong because no request is sent), then the network. The
dev fixture (cargo run -p xtask -- fixture) binds the default ${url:base} port
(8787), so with no PROEF_BASE_URL set it becomes your local target automatically
(it falls back to an ephemeral port, printing PROEF_BASE_URL, only if 8787 is busy).
“batch budget exceeded — scenario thread abandoned” — the watchdog killed
a batch that outran its computed budget (timeouts × attempts + delays +
repeats + margin, ADR-0007). Usually a huge retry:/delay: (or a hurl-side [Options] repeat:/
retry-interval:/max-time:) value — the pack lint caps literals (counts at
10 000, durations at one hour), and the computed budget itself clamps at four
hours, so a lint-clean product that would otherwise run for days is abandoned
at the ceiling. A {{var}}-driven
retry:/delay:/[Options] repeat:/max-time: cannot be estimated at all — it resolves
inside hurl at run time — so the batch falls back to the default budget
([http] timeout × 4, at least 60s) rather than to an estimate that assumes no
retries. If a legitimately long templated retry is being abandoned, raise
[http] timeout.
“global state file .proef-state.json is not valid JSON” — the persistent
World is derived data (only saveAs: global promotions live there).
Deleting .proef-state.json is safe; the next run recreates it.
Corrupt .proef-secrets.json — proef doctor names it; the next
proef secret set moves the wreck to .proef-secrets.json.corrupt and
starts fresh. Values are re-enterable; nothing else references the file.
“no scenarios matched the filters” — --tags/--scenario selected
nothing; exit 2 by design so a typo’d filter can never produce a silent
green CI run.
A failure line carries (via tests/hurl/admin.hurl#admin.search) — not an error. The
step ran a named fragment (ADR-0018) rather than an inline hurl: block, and that is
the third file involved: the request lives there, not in the feature or the pack. The
spelling is the one ref: accepts, so it pastes straight back into a pack. Every failure
sink carries it — console, explain, TAP, JUnit, the GitHub summary, the HTML report.
“ref: names no loaded fragment” (pack::unknown_ref) — either the name is a typo
(the message suggests the closest) or no fragment files were loaded at all, which the
message says plainly. The usual cause is a missing [run] fragments in proef.toml:
without it nothing is scanned, so every ref: is unknown. Note the root resolves against
the config file’s directory, not your working directory.
“reads name, which nothing supplies” (lower::unbound_placeholder) — a
fragment’s variables are not implicit: what the .hurl file reads must be bound at pack,
macro or step scope, captured by an earlier step, supplied by the fragment’s own
[Options] variable:, or carried by a secret of that name — and inside a bind:
value, an earlier-sorting sibling literal counts too (injected lines evaluate in name
order). --dry-run reports it without a network. With proef lsp running, completing
inside a bind: table offers what the pack’s ref:ed fragments still need — the union
over every fragment the pack refs (ranked by the nearest ref:, each labelled with its
owner), minus the names a fragment supplies itself.
“index is supplied twice” (pack::option_declared_twice) — the fragment sets the
name in its own [Options] variable: and a bind: supplies it. Both reach the entry
as variable: index= and hurl takes the last, which is the fragment’s — so the bound
value would silently never be sent. Delete whichever is not authoritative: the bind:
if the file’s own default is right, the variable: line if the pack should decide. The
same rule already applies to retry:/delay:.
Go-to-definition on a ref: does nothing — check that [run] fragments is set and
that the editor’s workspace root matches where proef.toml lives; the server resolves the
corpus relative to that file’s directory.
“payload does not parse: contains no hurl entries” — the hurl: block is
comments/blank lines only (a temporarily commented-out request). proef
refuses it at load: a step that executes nothing must not report green.
A Then step shows ∅ not run (its request entry did not run) — the
request it asserts on was skipped (guard or earlier failure); the asserts had
nothing to attach to.
Windows/| head pipelines — closed pipes are tolerated everywhere; if a
pipeline misbehaves, check the consumer, not proef’s exit code.
Building from source
proef-engine-hurl links native libraries. Debian/Ubuntu:
apt install build-essential pkg-config libssl-dev libcurl4-openssl-dev libxml2-dev libclang-dev. macOS: Xcode CLT suffices. Verify with
proef doctor — it reports the embedded hurl, parser/libxml2 linkage, and
libcurl. Prebuilt binaries (brew/binstall/GitHub Releases) need none of this.
Digging deeper
Every run leaves .proef-runs/<run-id>/: events.jsonl (the machine record —
EVENTS.md), run.log (console mirror), and artifacts/ with the
exact executed .hurl files. proef explain summarizes the latest run,
proef diff [base] [new] compares two of them (regressions, fixes, flakiness,
perf), and proef report [run] writes a self-contained HTML page of a run; the
reproduce: hurl --test … line under a failure replays the artifact with stock
hurl, taking proef out of the loop entirely.
A cancelled run (Ctrl-C) is a complete record, not a truncated one —
proef explain/proef report never banner it as incomplete, because its
nonzero skipped count already says what didn’t get to run. proef diff --fail-on-regression disagrees on purpose: it still refuses to certify “no
regressions” against a cancelled run, since a regression could be hiding among
the scenarios it never reached. The three commands’ differing treatment of
cancelled is a deliberate choice, not an inconsistency.
proef diff takes each side as a run id, a record directory, or an events
.jsonl file — the stream is the record, so a file means the same thing
under any name. That third form is the CI baseline flow: upload
.proef-runs/<id>/events.jsonl as an artifact on your main branch, download
it in the PR job as (say) baseline.jsonl, then
$ proef test tests/features --run-id pr
$ proef diff baseline.jsonl pr --fail-on-regression # exit 1 on passed → failed
and a scenario that regressed against main fails the PR without a run record store shared between jobs.
proef diff keys steps by (text, ordinal) so that line shifts don’t lie and
repeated steps stay distinct. The trade-off is positional: if a scenario loses
an earlier duplicate of a step, every later instance shifts down one ordinal,
and the comparison lines up two different runs’ steps. Timing and attempt
counts for the shifted steps can then be attributed to the wrong one. Renaming
or reordering steps between the two runs being compared is worth a second look
at the numbers; adding or removing steps at the end is not affected.
proef.toml — the project configuration reference
proef.toml lives in the project root — found by searching up from the working directory, so any subdirectory works (a project is where its proef.toml is, not where your shell is)
and is committed project config — it describes the suite, not your
machine. Every key is optional; an absent file means all defaults. Unknown keys
are rejected (deny_unknown_fields), so typos fail loudly instead of being
ignored.
The file holds a few kinds of thing, kept in distinct sections: runner config
(how proef behaves — [run], [http], [sla]), suite variables (data your
tests reference — [url], [vars]), run metadata ([meta], recorded in the
run head) and report tag links ([tag-links]), plus per-environment overrides
([env.<name>]). Secrets never live here (see below).
Precedence (highest wins): built-in defaults < proef.toml base tables <
active [env.<name>] < command-line flags. For a suite variable that means
[url]/[vars] < [env.<active>]; for a runner setting like jobs, the
--jobs flag still wins over both.
Reference
# ── runner config (how the tool behaves) ────────────────────────
[run]
suite = "tests" # default suite path — `proef test` needs no argument
fragments = "tests/hurl" # root of the hurl files packs may `ref:` (optional)
jobs = 8 # parallel scenario workers
runs-dir = ".proef-runs" # where run records land
keep-runs = 200 # past records kept there (0 = only the run in flight)
setup = "tests/setup.feature" # run once before the pool (optional)
teardown = "tests/teardown.feature" # run once after the pool (optional)
exclusive-tags = "@serial" # scenarios matching this run with the pool to themselves
[http]
timeout-ms = 30000 # per-request timeout (batch-level default)
follow-location = false # follow redirects
max-redirs = 10 # redirect ceiling (only meaningful with follow-location)
user-agent = "proef" # sent with every request — how a server spots test traffic
cookie-store = true # false = run cookie-less (hurl's --no-cookie-store):
# prove a stateless API by refusing to carry a session
# TLS and proxy: the settings that differ *between environments*, which is why
# they belong here and not in a pack. Paths resolve against this file.
insecure = false # skip certificate verification (warns loudly when on)
cacert = "certs/ca.pem" # trust this CA instead of the system store
client-cert = "certs/client.crt" # mTLS: the certificate to present
client-key = "certs/client.key" # mTLS: its private key
proxy = "http://proxy.corp:3128"
no-proxy = "localhost,.internal"
[sla] # opt-in run-level latency budget (omit = no gate)
p95-ms = 250 # 95th-percentile per-step duration ceiling
max-ms = 1000 # slowest single-step ceiling
# ── suite variables (data your tests reference) ─────────────────
[url]
base = "http://127.0.0.1:8787" # → ${url:base}; the dev fixture's default port
[vars]
apiVersion = "v1" # → ${vars:apiVersion}
username = "dev@example.com" # → ${vars:username}
# ── per-environment overrides (mirror the base tables) ──────────
[env.staging.url]
base = "https://staging.example.com"
[env.prod.url]
base = "https://api.example.com"
[env.prod.vars]
username = "release@example.com"
[env.prod.http] # an environment may override a runner setting too
timeout-ms = 60000
| Key | Default | Notes |
|---|---|---|
[run] suite | (unset) | default path for proef test/flows/macros/fragments/artifacts (and what doctor and lsp inspect); falls back to the tests/ convention, then errors. An explicit path always wins |
[run] jobs | available parallelism | --jobs flag wins; live threads never exceed the scenario count |
[run] runs-dir | .proef-runs | run records land here; rotation only ever deletes uuid-named dirs — see the --run-id note below |
[run] keep-runs | 200 | how many past records runs-dir retains, besides the one being written; 0 keeps none but the run in flight |
[run] setup | (unset) | feature run once before the pool (suite setup); its saveAs: global reaches every scenario; a failure aborts the run |
[run] teardown | (unset) | feature run once after the pool (suite teardown), only if setup succeeded; its failure is a distinct exit 3 |
[run] exclusive-tags | (unset) | tag expression (same language as --tags) selecting scenarios that run with the pool to themselves; a malformed expression is exit 2 |
[http] timeout-ms | 30000 | per-entry [Options] in a hurl block override it; 0 is refused (exit 2) — libcurl reads zero as no timeout, the unbounded hang the default exists to prevent |
[http] follow-location | false | per-entry [Options] override it |
[http] max-redirs | (engine default) | redirect ceiling; only reached when follow-location is on |
[http] insecure | false | skip TLS certificate verification. The run prints a warning naming the profile whenever this is on — a green run that verified nothing is not the same result, and nothing else would say so |
[http] cacert | (system store) | CA bundle to trust instead of the system store. Path — resolves against proef.toml |
[http] client-cert | (unset) | client certificate for mTLS. Path. A combined PEM may carry the key too |
[http] client-key | (unset) | private key for client-cert. Path. Setting it without client-cert is exit 2 — a key alone presents nothing, and the failure would otherwise surface at the server as an unexplained auth error |
[http] proxy | (unset) | proxy URL for every request |
[http] no-proxy | (unset) | comma-separated hosts bypassing proxy (curl’s NO_PROXY form) |
[http] user-agent | (engine default) | User-Agent for every request. Run-wide only — hurl has no per-entry user-agent option, so one entry opts out with a User-Agent: header instead |
[http] cookie-store | true | false runs the suite cookie-less (hurl’s --no-cookie-store): no Set-Cookie is kept, none is replayed — how a stateless API is proven stateless. Run-wide only — hurl has no per-entry spelling for it at all |
[sla] p95-ms | (unset) | 95th-percentile per-step duration ceiling; unset = no gate |
[sla] max-ms | (unset) | slowest single-step ceiling; unset = no gate |
[flaky] min-samples | 10 | minimum observed runs before proef flaky classifies a scenario (else insufficient-data) |
[flaky] recovery-runs | 5 | trailing clean runs that resolve a flagged scenario to healthy (hysteresis) |
[flaky] outage-rate | 0.8 | a run failing over this share of suite scenarios is an environment outage, excluded from the history |
[url] <key> | (none) | URL variables, referenced as ${url:<key>} |
[vars] <key> | (none) | non-secret variables, referenced as ${vars:<key>} |
[meta] <key> | (none) | run metadata recorded in the run head; [env.<name>.meta] overrides it, --meta flags win (ADR-0020) |
[tag-links] "<glob>" | (none) | tag glob → URL template ({tag} substituted); matching tags link out from the HTML report and GitHub summary. Base table only |
[env.<name>.<section>] | inherits base | per-environment override of any base section (url/vars/http/run/sla/meta) |
A custom --run-id is kept, but never rotated
--run-id accepts any single path component, and a run named that way writes
a perfectly good record that explain, diff, flaky, report and
--rerun all find — a directory is a record because it holds an
events.jsonl, not because of the shape of its name (ADR-0021).
“The latest run” is resolved by time rather than by spelling, and where that time comes from differs by name: a generated id is a uuid-v7, which carries the moment proef minted it, while a custom id carries no time at all and orders by the directory’s modification time instead. Both are wall-clock, so the two compare directly — but a custom-id record that is later copied or restored sorts by that moment, not by when it ran.
What a custom id changes is rotation: proef deletes only uuid-named
directories, so [run] keep-runs will never reclaim .proef-runs/pr-1234/.
A caller minting a fresh custom id per build accumulates records without
bound. That is deliberate and the asymmetry is the point — runs-dir may be
., and guessing at user-named directories is the worse failure by a wide
margin. Reading a directory is safe; deleting one is not.
Use --run-id when you want a stable, externally-known name — a CI job
archiving .proef-runs/pr-1234/ by path — and clean those up the same way you
created them.
Environments
An [env.<name>] profile mirrors the base tables and deep-merges over them,
key by key — so an environment lists only what changes; everything else inherits
from the base (the Cloudflare-Wrangler / Cargo-[profile.*] model). Select the
active environment at run time:
$ proef test # base [url]/[vars]; discovers the default suite
$ proef test --env prod # [env.prod.*] merged over the base
$ PROEF_ENV=staging proef test # same, via the environment variable
The rule is almost uniform: under [env.<name>], url.* / vars.* override
variables and http.* / sla.* override runner settings wholesale — but
[env.<name>.run] accepts only jobs. Every other [run] key (suite,
runs-dir, setup, teardown, exclusive-tags, …) names what the suite is,
not how hard to run it, and letting an environment swap those made two
environments run different tests while claiming to be the same suite — so
listing one there is a hard parse error (exit 2), never a silent no-op. A named-but-undefined --env is a
user error (exit 2) listing the known environments. A ${url:key} / ${vars:key} referenced
in a pack but defined in neither the base nor the active environment fails at lower time
(proef::resolve::missing_config_var).
TLS, proxies and mTLS ([http])
These are the settings that describe an environment rather than a test, which
is why they live here and not in a macro pack. Staging presents a self-signed
certificate, CI traverses a corporate proxy, production speaks mTLS — and the
same suite has to run against all three without a single pack changing. Put the
differences in [env.<name>.http] and the shared parts in [http]:
[http]
user-agent = "proef-suite" # every environment identifies test traffic
[env.staging.http]
insecure = true # staging's certificate is self-signed
proxy = "http://proxy.corp:3128"
[env.prod.http]
cacert = "certs/prod-ca.pem" # a private CA, not the system store
client-cert = "certs/client.crt" # mTLS
client-key = "certs/client.key"
Three things worth knowing before you rely on this:
- Paths resolve against
proef.toml, not the working directory — the same one-path rule every other config path follows, soproef testfrom a subdirectory finds the same CA bundle it finds from the project root. insecure = truewarns on every run, naming the profile that set it. A suite that goes green without verifying a certificate has not proved what a green suite normally proves, and the run record deliberately carries no config, so the warning is the whole audit trail.- A per-entry
[Options]block still wins — for all of these but one. These are batch-level defaults, so one request that must skip verification, or present a different certificate, says so in its own hurl block and overrides the table for itself alone. The exception isuser-agent: hurl has no per-entry option for it, so[http] user-agentis run-wide and a single entry cannot opt out. Set aUser-Agent:request header in that entry instead.
Credentials stay out of this table on purpose. There is no [http] user or
netrc key: a password belongs in the secret store (proef secret set), where
it is encrypted at rest and masked out of artifacts, events, logs and reports.
A proxy URL is the one edge — if yours embeds credentials, it is plaintext in
a file you probably commit, and it belongs in an environment variable your shell
expands into the proxy’s own configuration instead.
Balancing a shard matrix by duration (--shard-weights)
--shard I/N assigns scenarios by a frozen hash of (file, scenario), so adding
one test never re-buckets the others. What it cannot do is balance by time —
and a CI matrix finishes when its slowest shard finishes, so a count-balanced
split routinely leaves runners idle.
Every run that reaches its suite writes timings.json into its run directory (a
run whose [run] setup aborted has no suite to weigh, and writes none). Archive
that one file and hand it to the next run’s matrix:
# each matrix job, all reading the same archived file
- run: proef test --shard ${{ matrix.shard }}/4 --shard-weights timings.json
Every job must read the same file. The tempting alternative — “let each job
weight by its own newest local record” — is silently wrong: matrix jobs run on
different machines with different (usually empty) runs-dirs, so each would
compute a different assignment and scenarios would run twice or not at all while
the suite still reported green.
Two rules place scenarios, and they partition rather than compete:
- a scenario the file mentions is placed by the balanced split (longest-processing-time-first: heaviest to the lightest shard, repeatedly);
- a scenario it does not mention falls back to the frozen hash.
So a test added after the timings were captured still runs exactly once. That property is pinned by a test that runs a whole matrix and asserts set equality both ways.
What you give up is the thing hash mode was chosen for: a balanced split is not stable under insertion, so adding a slow scenario can move others. That is what balancing means, and it is why the flag is opt-in rather than the default. A missing or malformed weights file is exit 2, never a silent fall back to the unbalanced split you passed the flag to avoid.
SLA gate ([sla])
The optional [sla] table is a run-level latency budget: after a run, proef folds
every step’s wall-clock duration into two aggregates and fails the run if either exceeds
its ceiling. p95-ms caps the 95th-percentile step duration (the typical slow request);
max-ms caps the single slowest step. Skipped steps (which never hit the network) are
excluded from the population.
The gate is opt-in and off by default — with no [sla] table a run behaves exactly as
before. A breach prints the offending metrics and the slowest steps on stderr and maps to
exit 1 (a test failure), the same code as a failed assertion. It never introduces a new
exit code and never downgrades a User/System fault (exit 2/3), so a breach can only turn
an otherwise-green run red. Tighten or loosen per environment the usual way:
[sla] # baseline budget
p95-ms = 250
max-ms = 1000
[env.staging.sla] # staging tolerates slower responses
p95-ms = 800
This is distinct from hurl’s per-request duration < <ms> assert (which fails one step
inside a hurl: block): the [sla] gate is an aggregate budget over the whole run, while
duration < is a targeted per-request check. Use the assert for a hard per-endpoint SLA and
[sla] for a suite-wide budget.
Hurl fragments ([run] fragments)
Names the root directory holding the .hurl files a pack may ref: (ADR-0018), scanned
recursively. Those files stay valid hurl — the # @proef <name> annotation is an
ordinary comment — so the same file runs under hurl and under proef, and a corpus you
already own is annotated once instead of transcribed into YAML where it drifts.
[run]
fragments = "tests/hurl"
# tests/hurl/admin.hurl — runs as-is under `hurl --variables-file …`
# @proef admin.search
GET {{base}}/api/v1/admin/search/{{index}}
Authorization: Bearer {{apiToken}}
[Query]
q: {{q}}
HTTP 200
# the pack supplies the fragment's variables; nothing is implicit
bind:
base: ${url:base} # pack scope — every macro in the file
apiToken: ${secret:apiToken} # injected at run time, never into an artifact
macros:
search:
params: [q, index]
defaults: { index: records }
match: "the operator searches for {q}"
bind: { q: "${q}", index: "${index}" } # macro scope
steps:
- ref: admin.search
Relative paths resolve against this file’s directory, not the working directory, so
the root may sit outside the suite — even outside the repo — and the config still works
from any subdirectory. There is no convention fallback: with the key unset a ref:
reports proef::pack::unknown_ref and says no fragment files were loaded, which beats
guessing at a directory. Entries with no # @proef annotation are inert, and the files
are not scanned at all until some pack names a fragment — so pointing at a corpus you did
not write neither changes what runs nor costs anything to parse. When a pack does name
one, the corpus is scanned once per command, not once per pack load: a run with
[run] setup/teardown loads packs four times against the same files, and on a
200-file corpus rescanning each time was most of the run.
proef fmt never touches these files: it normalizes hurl blocks inside packs by
locating them in YAML, and a corpus you do not own is not proef’s to rewrite.
The read is bounded. A single corpus file over 8 MiB is skipped with
proef::pack::oversized_fragment_file, and the reader stops once the corpus as a whole
passes 64 MiB (proef::pack::fragment_corpus_too_large). Both name the file, both
leave every other file loading, and a ref: into a skipped file then reports
unknown_ref beneath them.
These are not tuning knobs — they are the line past which the input is not a test corpus.
A hurl file is human-authored text, so 8 MiB is on the order of 200,000 lines; the largest
corpus anyone has reported is 15 files and 112 fragments. The bound exists because the
root is one config line: pointed one directory too high it reads whatever is underneath,
and a 279 MB file cost 601 MB of resident memory on proef flows — a command that never
looks at a fragment. If you hit either limit, the fix is almost always to narrow
[run] fragments to the directory that actually holds your hurl files.
Suite setup & teardown ([run] setup / [run] teardown)
[run] setup and [run] teardown each name a feature file run once around the whole
suite (the model Playwright/Jest globalSetup use, ADR-0014). setup runs before the
parallel pool and merges its saveAs: global promotions into the shared store before any
scenario lowers, so it is the place to seed a fixture or provision shared state that every
scenario then reads via ${global:…}. teardown runs once after the pool for cleanup.
[run]
setup = "tests/setup.feature"
teardown = "tests/teardown.feature"
Failure semantics are deliberate: a setup failure aborts the run before the pool and maps to a user (2) or system (3) fault — a broken fixture is not a failing test, so it never becomes exit 1. Teardown runs only if setup succeeded (it still runs after ordinary scenario failures, for reliable cleanup), and a teardown failure is a distinct exit 3 cleanup fault — the suite’s own verdict stands, but the failure is never silently swallowed. Both features are excluded from the pool, so a setup/teardown feature inside the suite directory never also runs as an ordinary scenario. Auth is already covered by pre-set secrets, so setup is for seeding/provisioning, not obtaining a runtime token.
Run metadata ([meta], [env.<name>.meta], --meta)
Key/value pairs recorded in the run head and shown by the HTML report, the
GitHub summary, explain and diff — never interpreted, never harvested.
Static facts live in config; invocation facts ride the flag; the active env
profile’s table deep-merges over the base, and flags win over both — the
same base < env < flags chain as every other setting:
[meta]
team = "payments"
[env.staging.meta]
tier = "staging"
proef test --meta commit=$(git rev-parse HEAD) --meta build=$CI_JOB_URL
proef itself never reads git, the hostname, or CI variables — if you want
the SHA recorded, your shell harvests it and proef records it (ADR-0020).
The active --env profile name is recorded automatically as its own
field: it is user-chosen input, and without it two records of one suite
against different [env.<name>] merges read as regressions in diff.
Tag links ([tag-links])
Tag glob → URL template; a matching tag becomes a link in the HTML report’s
by-tag table and the GitHub summary ({tag} is substituted). The pattern
language is the same anchored glob --tags atoms use — one matcher for
selection, exclusivity and links:
[tag-links]
"JIRA-*" = "https://jira.example/browse/{tag}"
Everything else ignores the table; a tag matching no pattern stays plain.
Scenarios that cannot share the pool ([run] exclusive-tags)
Some scenarios cannot run beside anything: one asserting absolute positions
(items[0]) needs a store no concurrent scenario writes to. exclusive-tags
is a tag expression — the same language --tags takes — selecting those.
Atoms may glob (* any run, ? one character, anchored — @serial-* is the
whole family; case-sensitive, unlike Robot Framework):
[run]
jobs = 8
exclusive-tags = "@serial"
@serial
Scenario: The report lists every record in order
A matching scenario waits for the pool to drain, runs alone, and the pool refills
after it. Ordering is unchanged: scenarios still start in discovery order, so an
exclusive one never loses its place — and that is also what the cost is made of.
Queueing is strict FIFO, so while an exclusive scenario sits at the head no new
scenario starts: the pool drains to empty, the exclusive one runs by itself,
and only then does parallelism return to jobs width. Expect a throughput dip
around each exclusive scenario roughly as long as the slowest scenario already
running, plus the exclusive one’s own duration. Put another way, the cost is
bounded and predictable rather than free.
Exclusivity is enforced against the dispatcher’s own bookkeeping. A scenario the watchdog abandons for blowing its batch budget (ADR-0007) leaves the active set immediately, but its detached thread keeps going until the request in flight returns — hurl cannot be cancelled mid-entry. In that window an exclusive scenario can start while an abandoned neighbour is still issuing requests. It takes a budget blowout in the same moment to happen, and it is the one caveat the isolation guarantee carries; a run that never trips the watchdog never meets it.
This is exclusion, not ordering. A scenario that must run before the
others — installing a fixture the rest depend on — belongs in
[run] setup, which already
runs once before the pool exists.
A tag expression in config, not a reserved tag name, deliberately: with a bare convention, a scenario added months later lands untagged in the parallel pool and breaks isolation intermittently — which reads as flakiness rather than as a missing declaration. The expression keeps the rule in one reviewable place. A malformed one is a user error (exit 2), never a silently-ignored key.
Editor support (JSON Schema)
TOML language servers validate against JSON Schema, so one file buys
completion, hover documentation and typo detection in proef.toml itself:
proef schema config > proef-config.schema.json
Then point your editor at it. Taplo and tombi (the engines behind VS Code’s Even Better TOML, and the usual Neovim/Helix TOML setups) both read a first-line directive:
#:schema ./proef-config.schema.json
[run]
suite = "tests"
The schema is generated from the same Rust model that parses the file, so it
describes the keys as they are written — kebab-case (runs-dir,
keep-runs, exclusive-tags), not the struct spelling — and inherits
deny_unknown_fields, so an editor flags a typo’d key before you run anything.
Each key’s hover text is its documentation from this page’s source.
Regenerate it after upgrading proef; there is no --add-to here (that
installs the pack schema beside packs), because a project has one
proef.toml and the #:schema line is a one-time edit to a file proef
otherwise only ever reads.
What does not live here
- Secrets —
proef secret set/PROEF_SECRET_<NAME>env, on their own encrypted channel (${secret:…}); never inproef.toml(see AUTHORING — Secrets). - Per-machine values —
proef.tomlis committed and shared; anything that differs between teammates belongs in an environment variable (${env:NAME:-default}) or an environment profile you don’t commit.
Feature files reference ${url:…} / ${vars:…} / ${secret:…} and declare none
of them — variables have exactly one home, this file, and test files stay pure prose.
proef discovers the nearest proef.toml by searching up from the working directory
(like cargo/git), so it is found from any subdirectory.
The search only goes up, so a discovered file must sit at or above the directory
you run from — that is a requirement, not a convention. Putting proef.toml
beside the suite (tests/proef/proef.toml) and running from the repository root
does not work by discovery: nothing searches downward, so the config is simply not
found and every ${url:…} reads as unset.
--config <path> names the file instead and removes the constraint entirely:
$ proef test --dry-run --config tests/proef/proef.toml
It is global — every subcommand accepts it, including proef lsp and
proef test --watch. The editor loading a different config than the runner is
the drift that makes diagnostics untrustworthy, and --watch watches the file
the run actually resolved through, so editing it retriggers.
For proef lsp the flag also outranks the workspace root the client announces:
the flag names a file, and a named file is not a guess to be improved on.
A named file that does not exist is a user error (exit 2) for the runner, not a
fall back to defaults: discovery finding nothing means “no project here”, but a
named path that is not there is a typo, and answering it with a silently
unconfigured run is the thing worth refusing. proef lsp makes the opposite call
and starts anyway — an editor offering less is better than one that will not
boot, and unlike a run it produces no results to be wrong.
It applies to every subcommand in the same way, including the ones that read
nothing from the file: proef fmt --config missing.toml exits 2 rather than
formatting and reporting success. The flag names a file, and a named file that is
not there is a typo whatever the command was going to do with it.
Which directory a path is relative to
Paths written in proef.toml resolve against the directory holding
proef.toml. Paths typed on the command line resolve against the working
directory. That is the whole rule, and it has no exceptions: suite,
fragments, setup, teardown and runs-dir all follow it, as do the two
files proef keeps beside them — .proef-state.json (the persistent World) and
.proef-secrets.json (the secret store). An absolute value is taken as written,
so a corpus or a record store may sit outside the project entirely.
The consequence worth stating plainly: a project is where its proef.toml is,
not where your shell is. Running from a subdirectory finds the same config by
walking up, and now resolves the same suite, writes the same run records, and
reads the same World and the same secrets as running from the root. It is the
convention Cargo (path/members are manifest-relative), tsconfig.json and
pytest’s rootdir all follow.
With no proef.toml in scope there is nothing to be relative to, so paths stay
relative to the working directory, unchanged.
A config that does not sit at the project root spells its paths from where it
sits — a proef.toml in tests/proef/ beside a suite in tests/features/
writes suite = "../features".
How a path is spelled in what proef writes
The same rule, read backwards. Nothing proef records names the machine it ran
on: a path that reaches an artifact, a sidecar, an event, a report or a
diagnostic is spelled relative to the directory holding proef.toml.
That matters because ADR-0010 makes the emitted .hurl a contract — the same
inputs must give the same bytes — and a run record is meant to travel from a
laptop to CI and back. So all four of these produce the same record:
$ proef test # path derived from [run] suite
$ proef test suite # typed
$ proef test /home/you/proj/suite # typed absolutely
$ cd sub && proef test # from a subdirectory
Two cases keep an absolute path, because no project-relative spelling of them
exists: a suite or fragment corpus that genuinely sits outside the project
(fragments = "/opt/shared-corpus"), and a run with no proef.toml in scope. A
path typed relative is recorded exactly as typed — it is machine-independent
already, and it is the spelling your terminal can open.
How long run records live
Each run writes one directory under runs-dir holding its artifacts, its
events.jsonl and its run.log. [run] keep-runs bounds how many are kept:
the default of 200 suits an archive, and a suite re-run on every save wants far
fewer. The artifacts are the largest part and are byte-identical between runs of
an unchanged suite — but they are what that run executed, and once the corpus
moves on proef artifacts no longer reproduces them, so the cost is bounded
rather than dropped.
Rotation only ever deletes directories named by a generated run id, because
runs-dir may be . or otherwise shared with your own files. A record written
under --run-id <name> is therefore never rotated: if CI mints a fresh name per
build into a persistent directory, prune it yourself.
A starter file ships as proef.toml.example in the repository root.
Diagnostic codes — the greppable index
Every proef diagnostic carries a stable code (proef::<area>::<name>), printed
above the rendered message. Codes are a contract: they never change meaning,
and searching this file (or tests/errors/) for a code you hit is the fastest
route to the cause. The corpus column marks codes with a seeded broken
example under tests/errors/<area>__<name>/ — dry-running one shows the exact
rendered output.
Severity is error (fails validation/exit 2) unless marked warning.
proef::feature::* — feature-file parsing and expansion
| Code | Meaning | Corpus |
|---|---|---|
feature::parse | The file is not valid Gherkin (parser error, located) | ✓ |
feature::empty_file | The file is empty — a Feature: header is required | |
feature::no_examples | A Scenario Outline has no Examples rows | |
feature::ragged_examples | An Examples row’s cell count differs from the header | |
feature::bad_examples_header | Duplicate or empty Examples column name | |
feature::unknown_placeholder | An outline step uses <name> not present in the header | ✓ |
feature::empty_scenario | A scenario has no steps (background included) | ✓ |
proef::pack::* — macro-pack loading and validation
| Code | Meaning | Corpus |
|---|---|---|
pack::yaml | The pack file is not valid YAML | ✓ |
pack::duplicate_macro | Two macros share a name | ✓ |
pack::empty_macro | A macro has neither steps: nor expect: | |
pack::steps_and_expect | A macro has both steps: and expect: | |
pack::empty_expect | An expect: item asserts nothing | ✓ |
pack::empty_step | A step has no payload and no use: | |
pack::multiple_payloads | A step has more than one payload key | |
pack::unknown_step_kind | The payload key names no registered engine | ✓ |
pack::invalid_hurl | The payload does not parse as hurl (incl. zero-entry payloads) | ✓ |
pack::payload_invalid | A structured payload fails the engine’s validator | ✓ |
pack::bad_reference | A ${…} reference in the payload cannot resolve at probe time | |
pack::retry_not_finite | retry/repeat is -1, 0, or above the 10000 cap | ✓ |
pack::delay_unbounded | delay exceeds the 1-hour cap | |
pack::option_declared_twice | retry/delay set both in the block’s [Options] and as the step’s own key, or a variable both supplied by a fragment’s [Options] variable: and given by a bind: | ✓ |
pack::pattern_braces | Unbalanced {/} in a match: pattern | |
pack::pattern_empty_capture | {} with no capture name | |
pack::pattern_no_anchor | A pattern with no literal word (capture-only) | ✓ |
pack::pattern_unknown_capture | A {capture} that is not a declared param | ✓ |
pack::pattern_duplicate_capture | The same {capture} written twice in one pattern | ✓ |
pack::adjacent_captures | {a} {b} with no literal between captures | ✓ |
pack::default_not_param | A defaults: key that is not a declared param | ✓ |
pack::bad_save_target | saveAs: target other than global | |
pack::unknown_ref | ref: names no loaded fragment (suggests the closest) | ✓ |
pack::duplicate_fragment | Two fragment files declare the same # @proef name | |
pack::bad_annotation | A fragment file the engine’s parser could not read, an annotation it could not attach, or a name containing # (which could never be referenced) | |
pack::unreadable_fragment_file | A fragment file that could not be read at all (encoding, permissions) — its siblings still load, and it is silent until something ref:s the corpus | |
pack::oversized_fragment_file | A fragment file over 8 MiB — skipped unread (the size comes from the directory entry), siblings still load | |
pack::fragment_corpus_too_large | The corpus as a whole passed 64 MiB — the read stops, naming the file it stopped at | |
pack::body_form_conflict | A step is both ref: and a payload (or use:) | ✓ |
pack::bind_without_ref | bind: with no ref: to read it — on a step (an inline block takes ${…} instead), or on a macro whose steps have none (a use: target resolves its own) | ✓ |
pack::unread_bind_key | a bind: key no fragment in that scope reads (did-you-mean over the readable names) — the finer half of bind_without_ref, and the one a typo produces | |
pack::unknown_use | use: names no known macro | ✓ |
pack::use_cycle | use: composition forms a cycle | ✓ |
pack::use_too_deep | use: nesting exceeds depth 32 | |
pack::use_with_modifiers | A use: step carries modifiers that belong on the target | |
pack::use_with_payload | A step is both use: and a payload | |
pack::with_without_use | with: on a step that has no use: | |
pack::missing_use_param | The use: target requires a param with: does not supply | ✓ |
pack::unknown_with_key | A with: key the target macro does not declare | ✓ |
proef::bind::* — matching prose to macros
| Code | Meaning | Corpus |
|---|---|---|
bind::unbound_step | No macro pattern matches the sentence (suggests the closest, plus a paste-ready macro stub) | ✓ |
bind::ambiguous_step | More than one macro matches (all candidates listed) | ✓ |
bind::missing_param | A required param has no capture, table value, or default | ✓ |
bind::bad_table | A data table has an unusable shape | |
bind::unknown_table_key | A table key that is not a declared param | |
bind::table_conflict | A table value collides with a pattern capture | ✓ |
bind::docstring_unused | A docstring the macro never references (warning) |
proef::lower::* — lowering to engine batches
| Code | Meaning | Corpus |
|---|---|---|
lower::then_before_when | An expect: step with no previous request to attach to | ✓ |
lower::bad_status | expect: status: is not an HTTP status number | |
lower::kind_unrouted | Internal safety net: a lowered step’s kind maps to no engine (registry drift — unreachable through the CLI, which is why it has no corpus case) | |
lower::expansion_too_deep | Macro expansion exceeded depth 32 at run time | |
lower::unbound_placeholder | A {{variable}} nothing supplies — read by the fragment, or inside a bind: value (as the engine’s parser reads it: a hurl function like {{newUuid}} is not a variable) — with no bind: in scope, no earlier capture, no fragment-own [Options] variable:, no earlier-sorting sibling literal, and no secret of that name | |
lower::bind_shadows_capture | A literal bind: re-assigns a name an earlier step captured, so the bound value silently wins from that entry on (warning) | |
lower::multiline_bind | A bind: value that resolves to a line break or control character (tab excepted) — a hurl [Options] variable: is a single-line scalar; the inline hurl: | form is what splices a multi-line body | |
lower::secret_in_composite_bind | A bind: value mixes ${secret:…} into a larger string — bind the secret alone and put the surrounding text in the fragment | |
lower::dry_run_unknown | A runtime-only global under --dry-run (warning) |
proef::resolve::* — ${…} variable resolution
| Code | Meaning | Corpus |
|---|---|---|
resolve::unknown_variable | ${name} found in no scope (suggests the closest) | ✓ |
resolve::missing_env | ${env:NAME} unset and no :-default given | |
resolve::missing_config_var | ${url:key} / ${vars:key} defined in neither proef.toml nor the active [env.<name>] (suggests the closest) | ✓ |
resolve::missing_global | ${global:key} absent from the World (strict mode) | |
resolve::unknown_namespace | ${ns:…} with an unknown namespace | |
resolve::unknown_run_field | ${run:…} other than ${run:id} | |
resolve::fake_unknown | ${fake:kind} names no generator (suggests the closest) | |
resolve::empty_reference | An empty ${} | |
resolve::depth_exceeded | Resolution passed depth 8 — a reference cycle |
proef::emit::* — artifact emission
| Code | Meaning | Corpus |
|---|---|---|
emit::invalid_artifact | The emitted artifact does not parse with the engine’s parser |
proef::run::* — preparing a scenario to execute
| Code | Meaning | Corpus |
|---|---|---|
run::asset_unstageable | A file,…; asset could not be staged into the scenario’s asset root: absent beside the source that names it, named by a path proef will not follow, or claimed by two sources at once |
proef::config::* — proef.toml loading
| Code | Meaning | Corpus |
|---|---|---|
config::toml | The file is not valid TOML for the config schema (located — the caret sits on toml’s own error span) | |
config::unreadable | The file (or an explicit --config path) cannot be read |
proef::source::* — source access (LSP whole-suite analysis)
| Code | Meaning | Corpus |
|---|---|---|
source::unreadable | A discovered feature or pack source could not be read (surfaced by analyze_suite; the CLI treats an unreadable file as a system fault instead) |
proef::tags::* — reserved-tag recognition
| Code | Meaning | Corpus |
|---|---|---|
tags::reserved_tag_typo | A tag looks like a reserved one (@quarantine, @skip) but is not exactly it, so it has no effect — a warning with the spelling it likely meant |
Coverage note
The fragment-file codes (pack::duplicate_fragment, pack::bad_annotation,
pack::unreadable_fragment_file, pack::unread_bind_key,
lower::unbound_placeholder, lower::multiline_bind,
lower::secret_in_composite_bind, run::asset_unstageable) are covered in
crates/proef-cli/tests/fragments.rs rather than tests/errors/: they need a
[run] fragments root, and the seeded corpus is deliberately config-independent.
The config::* codes are covered in crates/proef-cli/tests/cli.rs for the
same reason: a broken proef.toml cannot live in the config-independent
seeded corpus.
30 of the 77 codes carry a seeded corpus case today; the corpus guard asserts
a minimum, not parity. When you add a diagnostic, add its code here and prefer
seeding a tests/errors/<area>__<name>/ case alongside it.
tags::reserved_tag_typo is a warning (a near-miss is not an error), so it
cannot live in tests/errors/ — that corpus fails dry-run by design — and is
covered by a unit test instead.
Every row here is a code some code path actually emits. That was not free: this
file used to carry a pack::load row for a defensive case that never had a code
(a non-diagnostic core failure while loading packs renders as a bare error:),
so a reader who grepped for it found nothing and had no way to tell the index was
wrong. A row nothing emits is worse than a missing row.
Event schema — the run record for machine consumers
.proef-runs/<run-id>/events.jsonl is the run record (ADR-0008): one JSON
object per line, in emission order, no second format. proef explain,
diff, flaky and the console tree derive from this stream; --junit and
the GitHub summary are built from the same run’s in-memory outcome (identical
content, different plumbing — the stream stays the only persisted format;
the one exception is a --rerun’s JUnit, whose carried base scenarios are
reconstructed from the base record — see rerun_of below).
Consume it with jq, a log shipper, or anything line-oriented.
Two derived sidecars sit beside the record and are not records: timings.json
(per-scenario durations, re-derivable from the stream, read back by
--shard-weights) and inputs.json (the run’s input fingerprint — a hash of
proef’s own inputs, proef flaky’s equivalence class, ADR-0020 amendment;
absent on pre-0.18 records). Neither adds a variant or a field to the stream.
Wire shape: serde-tagged with "event", snake_case names. The first line is
always run_started and declares schema (currently 1); the last is
run_finished. The schema is additive-only: new variants and new
optional fields may appear, existing fields never change meaning or vanish.
Consumers must ignore unknown variants and unknown fields.
Truncated records. A record whose last line is not run_finished is a
run that did not finish — a hard interrupt (second signal, exit 130), a
SIGKILL, or a crash. (A single SIGTERM/SIGHUP is a graceful cancellation:
the record closes normally with run_finished + cancelled.) Treat it as partial: the scenarios present did happen, but the
totals were never written and no verdict was reached. explain, report and
diff classify these and say “run incomplete” rather than reporting a total
they cannot know. Consumers should do the same rather than inferring zero.
One exception to “fields never change meaning”, recorded rather than hidden.
run_finished’s passed/failed/skipped counted every phase before
0.6.0; from 0.6.0 they are the main-suite verdict and exclude
[run] setup/teardown (ADR-0014), so those totals agree with the exit code,
--format json/TAP, and the summary line. One deliberate divergence: a
failed phase additionally appears in JUnit as its own suite (a gated
pipeline must see the failure), so JUnit’s per-case counts can exceed these
totals on a run whose setup or teardown broke. A pre-0.6.0 record read by a
current consumer reports the older meaning.
Variants
run_started — head of every stream.
schema (u32) · run_id (string — uuid-v7 by default, but --run-id passes any name through verbatim, so consumers must treat it as opaque).
scenario_started — scenario (string) · file (string) · timestamp_ms
(u64, run-relative ms — only present with injected timing) · worker (u64,
0-based worker index — only present with injected timing). The timing pair is
stamped at the CLI sink on the worker thread (the sans-IO core leaves it absent),
and powers the HTML report timeline (ADR-0015).
batch_started — a contiguous same-engine step batch was dispatched.
scenario · engine (e.g. "hurl") · steps (count).
entry_running — live progress: one event per execution attempt of an
artifact entry, retries included. scenario · engine · entry (0-based
ordinal within the scenario’s artifact) · retry (0 = first attempt).
step_finished — scenario · engine · step ({file, line, text},
the authored feature anchor) · status (passed | failed | skipped | warned) · attempts (u32) · duration_ms (u64) · captures (capture
names only — never values) · fragment (file.hurl#name of the named
fragment the step ran, ADR-0018 — only present for a ref: step, so an
inline hurl: block and every pre-fragment record omit the key entirely) ·
label (the pack step’s authored name:, resolved — only present when
the step has one. One feature sentence commonly lowers to several engine
steps, and they share a step anchor exactly; the label is what tells them
apart, and it is the same text the emitter writes into the artifact’s entry
comment) ·
detail (string, only present on
failures/warnings/skips-with-reason) · attempt_details (array of strings —
the messages from earlier, failed attempts of a step that ultimately passed;
only present for a flaky pass, feeds JUnit <flakyFailure>) ·
reproduce_hint (string — the redacted curl of the failing request, the same
line the console prints as reproduce:; only present on failed steps,
absent in every pre-field stream, masked at the sink boundary like detail).
scenario_finished — scenario · file (feature path — with scenario,
the run-wide identity: names are unique only within one file; absent in
records that predate the field) · status · timestamp_ms (u64, run-relative
end ms — only present with injected timing) · worker (u64, in the
schema but never populated by proef’s own writer today — it is emitted from
the main dispatcher thread, not the scenario’s worker, so consumers must
accept the field without expecting it — ADR-0015).
env / metadata / shuffled — on run_started, all additive and
absent when unset. env is the active --env/PROEF_ENV profile name;
metadata is the explicit user-supplied map (--meta k=v, [meta],
[env.<name>.meta] — proef never harvests: no git, no hostname, no CI env
sniffing; ADR-0020); shuffled says the order was re-dealt (the permutation
is seeded by run_id, so the pair reproduces it). Keys and values pass the
sink-boundary mask like every text field.
rerun_of — on run_started, optional (additive; absent = not a
rerun). The base record’s run_id a --rerun re-ran failures from. The
rerun’s JUnit carries the base’s not-re-run suite scenarios as ordinary
testcases, and report overlays the base for a whole-suite page —
composition over records, the record files themselves never merge
(ADR-0008). Totals and the exit code stay the rerun’s own (ADR-0014).
tags — on scenario_finished, optional (additive; absent when the
scenario carries none and in every pre-field stream). The accumulated tags
(feature → rule → scenario → examples, deduped, authored order, @
stripped). On the finished event only: the cancel-skip path emits no
scenario_started, and per-tag skip counts need every scenario.
exclusive — on scenario_started, optional bool (absent = false,
which is what every pre-field record meant). The scenario ran with the pool
to itself ([run] exclusive-tags) — the bool the scheduler itself read,
recorded for timeline post-mortems (R11-6).
reason — on scenario_finished, optional (additive; absent in
pre-field records and on every non-skipped scenario). Why the scenario is
skipped: an authored skip carries the pasteable tag spelling (@skip /
@skip:reason), which always begins with @; a mechanical skip
(cancellation) carries proef-fixed prose, which never does. --rerun keys
its re-queue decision on exactly that split (ADR-0019).
phase — on scenario_started/scenario_finished, optional. "setup" or
"teardown" when the scenario belongs to a [run] lifecycle phase, absent for
an ordinary suite scenario. Consumers should use it rather than inferring phase
membership from the feature path: run_finished’s totals exclude phases
(ADR-0014), so a consumer that counts every scenario_finished will not match
them. Added additively — records without it have no phases, which is what they
had.
run_finished — tail. passed · failed · skipped — the main-suite
verdict: scenario counts for the primary suite only (ADR-0014). [run] setup/teardown scenarios still appear as their own scenario_started/
scenario_finished events earlier in the stream, but are excluded from these
totals, so they agree with the console summary: line, proef explain,
proef report’s HTML headline, --format json, TAP, the SLA gate, and the
exit code — JUnit agrees too except that a failed phase rides in as its own
suite (see above) · cancelled (bool, only present when true).
Example stream
{"event":"run_started","schema":1,"run_id":"019f…"}
{"event":"scenario_started","scenario":"reference","file":"suite/case.feature","timestamp_ms":0,"worker":0}
{"event":"batch_started","scenario":"reference","engine":"hurl","steps":2}
{"event":"entry_running","scenario":"reference","engine":"hurl","entry":0,"retry":0}
{"event":"step_finished","scenario":"reference","engine":"hurl","step":{"file":"suite/case.feature","line":4,"text":"the cookie session is exercised"},"status":"passed","attempts":1,"duration_ms":12,"captures":[]}
{"event":"step_finished","scenario":"reference","engine":"hurl","step":{"file":"suite/case.feature","line":5,"text":"the response status is 200"},"status":"passed","attempts":1,"duration_ms":0,"captures":[]}
{"event":"scenario_finished","scenario":"reference","file":"suite/case.feature","status":"passed","timestamp_ms":12}
{"event":"run_finished","passed":1,"failed":0,"skipped":0}
Guarantees
- Secrets never appear. Redaction applies once at the sink boundary
before any reporter (property-tested);
capturescarries names only. - Order: events of one scenario are internally ordered; scenarios
running in parallel interleave. Group by
(scenario, file)before assuming sequence across the stream —scenarioalone collides when two feature files reuse a name. - Flake-safe assertions: assert on
attemptscounts and normalized event order, never wall-clock (duration_msis engine-measured and varies).
Recipes
jq -r 'select(.event=="step_finished" and .status=="failed") | "\(.step.file):\(.step.line) \(.detail)"' events.jsonl
jq -r 'select(.event=="run_finished")' events.jsonl # the suite verdict (setup/teardown excluded)
jq -r 'select(.event=="entry_running") | .retry' events.jsonl | sort | uniq -c # retry pressure
# which fragment files a failing run actually exercised (ADR-0018); `// empty`
# drops the inline steps, which carry no `fragment` key at all
jq -r 'select(.event=="step_finished") | .fragment // empty' events.jsonl | sort | uniq -c
proef — Product Requirements Document
Status: approved; US-1…US-12 all in service (multi-engine architectural only), plus
post-M5 work: external config/environments (ADR-0012), the v0.6–v0.8 correctness series,
v0.9.0, named hurl fragments (ADR-0018 — see the §3 amendment), the adoption response,
the 0.11–0.14 hardening and CI-scale series, the RF capability audit and reserved tags
(ADR-0019/0020), and the deep-improvement waves plus round 19. CLAUDE.md’s Status
block is the running ledger; the goals and non-goals below are what has not changed.
· Date: 2026-07-28, status refreshed 2026-08-31 · Owner: Emre
Companion docs: ADRs for the why, TECH-SPEC for the how,
IMPLEMENTATION-PLAN for the when.
1. Problem & context
End-to-end tests written as Gherkin business prose — bound to a compiled-in vocabulary of YAML macro packs — let non-programmers author real tests. Meanwhile the backend team maintains a corpus of hand-written Hurl files as its API-testing lingua franca. There is no tool that joins the two: Gherkin prose on top, Hurl-grade API execution underneath — and none whose engine seam can later admit further engines behind the same prose (architectural readiness only).
proef is that tool: a declarative, modular, multi-engine e2e runner (mostly Rust).
One core parses/binds/lowers Gherkin; pluggable engines execute. The API engine embeds
hurl itself (ADR-0001), so API semantics are hurl’s by construction, and every scenario
also produces real .hurl artifacts the backend team can run and read (ADR-0010).
2. Goals
G1. Author API e2e tests as pure Gherkin prose in the 500-series style — no code,
no URLs in prose, data tables and Scenario Outlines supported.
G2. Execute with hurl’s exact semantics (asserts, captures, retries, templating) via the
embedded engine; identical results to the stock hurl CLI on the same artifacts.
G3. Emit canonical .hurl artifacts + sidecar maps for every scenario — debuggable with
the tools the backend team already uses, replayable via hurl --variables-file.
G4. Be modular: adding an engine is a new crate + a registry line, with
zero changes to proef-core (the structural acceptance test, ADR-0002).
G5. Track upstream hurl safely over time: exact pins, a zero-diff fork as patch vehicle,
and an upgrade-canary CI job so new hurl releases are absorbed deliberately (ADR-0003).
G6. First-class failure UX: every failure maps to the .feature line (and the artifact
span), rendered with miette-style labeled diagnostics; stable exit codes 0/1/2/3.
G7. CI-native: JUnit XML, GitHub job summary, JSONL run records, tag filtering, parallel
scenarios, --dry-run validation gate, and (M5) a libtest-mimic harness so
cargo nextest run and IDEs can drive scenarios.
3. Non-goals (v1)
Further engines (designed-for, not built — M6+); a desktop dashboard or server mode (precedent exists when needed); generating Gherkin, macros, or prose from hand-written hurl files (see the amendment below); OpenTelemetry export (semconv still immature — JSONL is the source of truth); dynamic plugin loading; Windows-static or musl-static binaries (dynamically linked like hurl’s own — consequence of ADR-0001); API mocking/contract testing; load testing.
Amendment (2026-08-11): the hurl non-goal is about generation, not direction
This non-goal previously read “importing/round-tripping hand-written hurl files into Gherkin (artifacts flow outward only)”. The clause and its parenthetical said two different things, and the parenthetical was the broader of the two: read literally it forbids hurl text being an input at all, which ADR-0018 (named hurl fragments) needs it not to.
What the non-goal protects is that proef never authors a test for you. A suite’s prose and its binding vocabulary are written by people, deliberately; a tool that derives them from an existing corpus produces scenarios nobody chose the words for, and the review that makes a feature file worth having never happens. That reasoning is untouched, and ADR-0016 (OpenAPI generation) stays declined on the same ground.
It does not extend to hurl text being an input source. A macro pack is already an
input written in a non-Gherkin language; a .hurl file naming reusable fragments is
another, and §1’s own framing — “the backend team maintains a corpus of hand-written
Hurl files as its API-testing lingua franca. There is no tool that joins the two” —
describes joining that corpus as the product’s purpose, not as a boundary. Nothing is
generated: features and macros stay hand-authored, and a fragment is inert until a
macro names it. The worklist reached the same reading independently for a different
item (OPEN-FINDINGS M2: “the non-goal forecloses a direction of data flow, not the
ability to check your own work”).
Recorded honestly: M3 asked that this charter be re-examined with a measured port cost, and that measurement does not exist. The narrowing is therefore argued from the non-goal’s own rationale rather than from evidence that pasting is expensive. If the measurement later shows pasting is cheap, that argues about priority — it does not restore a prohibition this amendment finds was never the point.
4. Users & personas
P1 — Test author (QA, PM, support engineer; not necessarily a programmer). Writes
.feature files against the existing step vocabulary. Cares about: prose that reads
naturally, --dry-run telling them exactly which step is wrong, failures pointing at
their line, not internals.
P2 — Pack maintainer (developer). Owns the macro packs: adds steps, params,
asserts. Cares about: pack readability (raw hurl blocks, ADR-0004), schema autocomplete,
load-time validation with did-you-mean hints, safe refactors (duplicate/cycle detection).
P3 — Backend engineer (owns the hurl corpus). Consumes emitted artifacts; pastes
between corpus and packs, or — since ADR-0018 — annotates a corpus file once and lets
packs name its entries, so the same file stays runnable under stock hurl. Cares
about: artifacts being idiomatic hurl, never containing secrets, runnable standalone;
and that proef reading their corpus never edits or reformats it.
P4 — CI pipeline (machine). Cares about: stable exit codes, JUnit/JSONL outputs,
deterministic behavior, bounded runtime (cancellation/budgets, ADR-0007), the canary job.
5. User stories & acceptance criteria
US-1 (P1) As a test author I write a scenario in prose and run it against a live
environment. AC: the four 500-series .feature files run via proef packs with
prose unchanged except agreed wording fixes; proef test runs them green against the
fixture; failures name feature file + line + step text.
US-2 (P1) I validate without executing. AC: proef test --dry-run binds every step,
expands every macro/outline, resolves ${…}, parses every generated artifact with hurl’s
parser, checks every file,…; asset those artifacts read is present, and exits 2 with
labeled diagnostics on any failure — no network I/O.
US-3 (P1) I pass data per step. AC: inline {captures}, | key | value | data tables,
and Scenario Outline <placeholders> all fill macro params; conflicts and missing
required params are parse-time errors naming the line.
US-4 (P1) Steps chain state. AC: a capture in one step (clientId) is usable in later
steps of the scenario as {{clientId}}; saveAs: global persists across scenarios and
runs (World, ADR-0005).
US-5 (P1) Slow backends don’t flake. AC: a step with retry: polls until its asserts
pass or the finite budget ends (maps to hurl [Options] retry); optional: steps warn
instead of failing.
US-6 (P2) I extend the vocabulary. AC: adding a macro with a match: pattern + raw
hurl block to a pack makes the new prose step available; proef schema reflects it; load
rejects ambiguous names, adjacent captures, literal-free patterns, infinite retries.
US-7 (P3) I get artifacts. AC: proef artifacts writes per-scenario .hurl +
.map.json (+ .vars when Worlds are referenced); stock hurl --test yields the same
verdicts (spike-proven); no secret values appear in any artifact.
US-8 (P4) CI consumes results. AC: exit codes 0/1/2/3 stable and integration-tested;
--junit auto under GITHUB_ACTIONS; JSONL event log written per run; --tags filters.
US-9 (P1/P4) Runs are observable. AC: console BDD tree with per-step timing/attempts;
proef explain [run] summarizes the latest failures from run records; proef diff [base] [new] compares two runs for regressions, fixes, flakiness, and perf deltas;
proef report [run] writes a self-contained HTML report of a run.
US-10 (P2) Secrets stay secret. AC: proef secret set stores encrypted values;
secret values never appear in artifacts, logs, reports, or events (property-tested).
US-11 (P4) hurl upgrades are safe. AC: the canary job builds against the next hurl
release and replays the suite; pins move only after it is green (runbook in
IMPLEMENTATION-PLAN §7).
US-12 (P1, M5) IDE/nextest integration. AC: the libtest-mimic harness lists one test
per scenario and cargo nextest run executes and reports them.
US-13 (P1) I can start from something that works. AC: proef init writes a
minimal suite (proef.toml, one .feature, one matching pack) that passes
--dry-run unchanged, installs the pack JSON Schema for editor completion, and
never overwrites an existing file.
6. Functional requirements (condensed; TECH-SPEC is normative)
Authoring: full gherkin-crate grammar (Feature/Rule/Background/Scenario/Outline/
Examples/tables/docstrings/tags/i18n); tags filter runs; variables come from
proef.toml (${url:}/${vars:}, ADR-0012), never the feature files. Packs: YAML
skeleton with match: patterns
({name} captures), params/defaults/tags/description; step bodies as raw
hurl: blocks (primary) or structured form (reserved for future engines); assert-only
expect: macros merge into the previous request (Then-steps); use:/with: nesting
with cycle/depth limits; schemars-derived JSON Schema; lint pass at load. Execution:
scenario = unit of isolation and parallelism (--jobs); contiguous same-engine batching;
engine-hurl via embedded run_entries with buffered I/O; per-entry [Options] override
batch defaults (verified); World seeding/merge-back; cooperative cancellation + budgets.
Artifacts: canonical emit, sidecars, vars files, # optional markers. Reporting:
event spine → console/JUnit/JSONL/GitHub-summary reporters; run-record rotation.
CLI: test ([path] --env --dry-run --tags --jobs --junit --format json|tap --watch --run-id --sarif --rerun --scenario[-file]; path optional — [run] suite then the tests/ convention), flows,
macros (call counts + dead-macro report), artifacts, schema [--add-to],
secret set|list, explain, diff [base] [new] --fail-on-regression,
report [run] -o <file>, doctor. Config
(proef.toml, ADR-0012): runner settings ([run] incl. setup/teardown suite
lifecycle features, ADR-0014; [http]/[sla]) + suite variables
([url]/[vars], referenced ${url:key}/${vars:key}) + per-environment overrides
([env.<name>]); precedence defaults < base tables < active [env.<name>] (via
--env/PROEF_ENV) < flags. Secrets via PROEF_SECRET_<NAME> env override → the
encrypted store (proef secret set — hidden prompt, or --stdin for
scripts), never in proef.toml.
7. Non-functional requirements
Correctness/fidelity: API semantics are hurl’s own binary-identical engine — no
reimplementation drift is possible (ADR-0001/0010). Portability: Linux (glibc),
macOS, Windows; documented build prereqs (libcurl/libxml2/libclang), doctor-checked;
dynamically-linked dist binaries (hurl’s own model). Performance: startup overhead
(parse+bind+lower for a 30-scenario suite) under ~1 s; execution dominated by the network;
scenario-level parallelism with worker threads. Reliability: deterministic lowering
(injected clock/run-id, sans-IO-lite core); bounded runtime under cancellation budgets;
state writes atomic (temp+rename). Security: secret redaction invariants
property-tested; artifacts guaranteed secret-free; encrypted-at-rest secret store;
0600 on sensitive outputs. Maintainability: exact pins + canary + thin-fork policy;
CI gates: fmt, clippy -D warnings, nextest, doc -D warnings, deny, machete,
zizmor, docs-check, public-api (audit nightly).
Extensibility: new engine = new crate implementing EngineFactory/EngineSession +
one registry line; step-kind schema contributed by the engine; zero core changes
(the M6 acceptance test).
8. Success metrics
M-1: the four 50x API features run under proef with prose unchanged (US-1) — the
port-fidelity bar. M-2: artifact parity — stock hurl CLI verdicts match proef verdicts on
every artifact in the integration suite (already spike-proven; kept as a CI invariant
until M4, then by construction). M-3: --dry-run catches 100% of the seeded
pack/feature error corpus with line-accurate diagnostics. M-4: one hurl upstream release
absorbed via the canary runbook with zero suite regressions. M-5: a future non-hurl engine lands (M6)
with git diff --stat proef-core empty.
9. Release phasing
v0.1 = M0–M3 (authoring, validation, embedded execution, artifacts, console+JSONL); v0.2 = M4 (upstream tracking hardened, JUnit/GitHub reporters); v0.3 = M5 (breadth: multipart/form/docstring bodies, watch, explain, libtest-mimic harness); v1.0 = stability declaration of pack schema + CLI + event schema; M6 engines version independently. Detailed task breakdown: IMPLEMENTATION-PLAN.md.
proef — Technical Specification
Status: normative · Date: 2026-07-28, current through ADR-0020 (run
metadata) · decisions referenced as ADR-XXXX. Post-M5 work — external config and
environments (ADR-0012), the v0.6–v0.8 correctness series, v0.9.0, fragments, and the
RF-audit series (reserved tags ADR-0019, run metadata ADR-0020) — is
specified here too; “M0–M5” described this file’s scope only until those landed.
Verified upstream facts cite hurl master @ 03fcb84c (2026-07-27) as file:line.
1. System overview
.feature files macro packs (YAML + raw hurl blocks)
│ │
▼ ▼
┌─────────────────────────────────────────┐
│ proef-core (pure: no IO, no clock, no │
│ rand — values injected) │
│ parse ─ bind ─ lower ─┬─ emit │
│ (gherkin) (matcher) │ (.hurl + │
│ │ sidecars) │
│ dispatch: contiguous │ │
│ same-engine batches ──┼── events ──────┼──► reporter stack
└────────────┬───────────┴────────────────┘ (console/JUnit/JSONL/GH)
▼ Box<dyn EngineSession>
┌──────────────────┐ ┌──────────────────┐
│ proef-engine-hurl│ │ future non-hurl │
│ parse_hurl_file +│ │ engine (seam- │
│ run_entries │ │ ready, none │
└──────────────────┘ │ scheduled) │
└──────────────────┘
World (typed vars + persistent global store) threads through every batch
Scenario = unit of isolation, parallelism, retry, and artifact emission. The orchestrator (in core, driven by cli) owns threads, the token, the World, and the event stream.
2. Workspace
proef/
rust-toolchain.toml # channel "1.97.1" (policy: latest stable at its x.y.1), rustfmt+clippy
Cargo.toml # virtual workspace, resolver="3", workspace.{package,dependencies,lints}
deny.toml .config/nextest.toml proef.toml.example
.github/workflows/ci.yml # fmt, clippy -D warnings, nextest, doc -D warnings,
# deny, audit, canary (ADR-0003)
xtask/ # automation as Rust (fixture, canary, docs-check, public-api); just aliases
crates/
proef-core/ # engine-agnostic: gherkin parse, packs, binding, lowering, IR,
helpers/ # emit, dispatch, World/state, events, errors, reporters
proef-engine-hurl/ # EngineFactory/EngineSession impl over embedded hurl
proef-cli/ # bin `proef`: clap, registry assembly, miette rendering
proef-fixture/ # dev-only: in-process synchronous fixture API server (ADR-0011)
proef-harness/ # libtest-mimic bridge: one Trial per scenario (US-12)
proef-lsp/ # language server: SourceProvider + collect-all analyze_suite over core
tests/ # .feature corpus + fixtures
docs/ # this corpus
Dependency rules: engines depend on core; core depends on no engine; cli depends on both
and is the only miette user (ADR-0009). Engines sit behind cargo features in cli
(engine-hurl default-on; any future non-hurl engine would be added the same way — none
scheduled). Only proef-engine-hurl
carries native build prereqs; proef-core is pure Rust. Lints/conventions: a strict
workspace lints table verbatim (clippy all=warn + curated pedantic slice), publish = false at the workspace root — overridden to true by the four crates that publish
(proef, proef-core, proef-engine-hurl, proef-lsp), MIT OR Apache-2.0.
3. Core domain types (sketches; signatures normative, field lists indicative)
#![allow(unused)]
fn main() {
// world.rs — typed variable scope (ADR-0005)
pub enum Value { String(String), Number(f64|i64…), Bool(bool), Null } // mirrors hurl Value subset
pub struct World { scenario: BTreeMap<String, Value>,
global: GlobalStore /* .proef-state.json, atomic temp+rename save */ }
// step.rs — lowered, engine-agnostic
pub struct StepRef { pub file: Arc<str>, pub line: usize, pub text: Arc<str> } // feature anchor
pub struct LoweredStep { pub step: StepRef, pub kind: StepKindId, pub payload: StepPayload,
pub optional: bool, pub when: Option<Guard>,
// `file.hurl#name` for a `ref:` step, None for an inline block
// (ADR-0018) — qualified at lowering so a record stands alone
pub fragment: Option<String> } // retry travels as baked [Options]
pub enum StepPayload { HurlEntries(String /* lowered hurl text */),
MergedAsserts { lines: usize /* expect: rows own the appended assert lines */ },
Structured(serde_json::Value) }
pub struct StepBatch { pub index: usize /* scenario-wide ordinal */, pub engine: EngineId,
pub steps: Vec<LoweredStep> }
// engine.rs — the seam (ADR-0002); see ADR text for EngineFactory/EngineSession
pub struct StepKindSpec { pub prefix: &'static str, pub schema: &'static str /* JSON-Schema frag */,
pub validate: Option<fn(&str) -> Result<(), PayloadProbeError>>,
// fragment files (ADR-0018): one Option, so extension and reader
// cannot disagree; discovery asks for the extension, never names one
pub fragments: Option<FragmentSupport>,
// raw [Options] keys → the core's budget policy, so option spellings
// live only in the engine that owns them (ADR-0007)
pub options: Option<fn(&str) -> Option<RawOption>>,
// which files a lowered payload sends, so the emitter records what an
// artifact reads without knowing the engine's body grammar (ADR-0002
// amendment, 2026-09-10 correction)
pub assets: Option<fn(&str) -> Vec<String>> }
pub struct FragmentSupport { pub ext: &'static str /* "hurl" */, pub scan: FragmentScanner,
// "which variables does one template value read", answered by the
// engine's parser — a hurl function ({{newUuid}}) is not a variable
pub template_reads: fn(&str) -> Vec<String> }
pub type FragmentScanner = fn(&str) -> Result<ScannedFile, FragmentScanError>;
// ScannedFile = { fragments: Vec<ScannedFragment>, unannotated: Vec<usize> } — the
// unannotated entry lines feed `proef fragments`' listing, not the pack loader
// Everything here is *read* from the entry — nothing is declared twice, so nothing can drift.
// A scanner reports ONLY annotated entries: an unannotated one is not a fragment, and a
// foreign corpus is mostly those, so building them only to be discarded is the bulk of a scan
pub struct ScannedFragment { pub name: String /* from `# @proef <name>` */, pub text: String,
pub line: usize, pub placeholders: Vec<String> /* reads */,
pub declared_options: Vec<String> /* ⊆ OPTION_FAMILIES */,
// `[Options] variable:` — what the entry supplies to itself.
// Answers its own placeholders (so the file runs standalone)
// and clashes name-to-name with a `bind:` of that name; kept
// apart from declared_options, which clashes family-to-family
pub supplied_variables: Vec<String> }
pub struct FragmentScanError { pub line: usize, pub column: usize, pub message: String }
// The vocabulary `declared_options` must use: matched by string equality against the pack's
// own option keys, so an engine-only spelling silences `option_declared_twice` rather than
// firing it. `MacroStep::declared_options()` derives the other half of that comparison.
pub const OPTION_FAMILIES: &[&str] = &["retry", "delay"];
// The one place `secret_bindings` (variable → secret) is joined with `secrets` (name → value).
// Yields borrows: an owned map would copy every secret value per scenario (ADR-0005)
pub fn secret_variables<'a>(bindings: &'a BTreeMap<String, String>,
secrets: &'a BTreeMap<String, String>)
-> impl Iterator<Item = (&'a str, &'a str)>;
pub struct DoctorCheck { pub name: &'static str, pub run: fn() -> DoctorResult }
pub struct BatchResult { pub steps: Vec<StepOutcome>, pub error: Option<EngineError> }
// `fragment` is carried here as well as on the event: JUnit, the job summary, the
// annotations, TAP and the console are built from RunSummary after the event stream has
// been written out, so they cannot read it back. Both are copies of one lowering-time
// source, so they cannot drift from each other.
pub struct StepOutcome { pub step: StepRef, pub status: Status, pub attempts: u32,
pub duration: Duration, pub detail: Option<String>,
pub attempt_details: Vec<String>, pub reproduce_hint: Option<String>,
pub fragment: Option<String> }
// events.rs — the spine (ADR-0008); serde, versioned
#[serde(tag = "event", rename_all = "snake_case")]
pub enum Event { RunStarted{..}, ScenarioStarted{..}, BatchStarted{..}, EntryRunning{..},
StepFinished{..}, ScenarioFinished{..}, RunFinished{.., cancelled} }
pub struct EventSink(Arc<dyn Fn(&Event) + Send + Sync>); // borrowed events
}
4. Pipeline (all in proef-core; pure — inputs include injected run_id, now, env snapshot)
4.1 Load packs. The fragment corpus ([run] fragments) is read once per command
into a pack::FragmentCorpus and scanned at most once, lazily — only when some pack
actually carries a ref:, which is what makes “pointing at a corpus you did not write
costs nothing” true of the scan (CONFIG.md). One proef test loads packs up to four
times (the suite, then [run] setup/teardown, each validated and then run) against
the same corpus, so the memo belongs with the corpus rather than the caller — a caller
that scanned eagerly to share the result would trade the promise for the speed.
Discover embedded helpers/ + project packs/; serde_norway with
deny_unknown_fields; validation passes: (1) match: guard rails — must contain literal
text, no adjacent captures, unclosed braces rejected; (2) params/defaults coverage;
(3) duplicate macro names across packs → error (qualify pack.yaml#name); (4) use:
cycle + depth ≤ 32; (5) unknown with: keys → “did you mean” (edit distance); (6)
finite-retry lint — retry: requires a finite count (ADR-0007); (7) hurl blocks:
lower a probe instantiation with placeholder params and parse_hurl_file it — syntax
errors reported with block-relative spans mapped to pack file/line; (8) engine kinds:
every step kind must be claimed by a registered engine’s StepKindSpec.
Fragments (ADR-0018) add five: (9) a ref: must name a loaded fragment → “did you
mean” over the scanned names, or a pointer at [run] fragments when none loaded;
(10) duplicate fragment name across files → error (qualify file.hurl#name, the same
suffix matching pack.yaml#name uses); (11) a step is ref: xor a payload/use:;
(12) an option family declared both in the fragment’s own [Options] and as the step’s
YAML key → error (pass 6’s twinned-option rule, applied across the file boundary), and
the same rule name-to-name for a variable both supplied by the fragment’s [Options] variable: and given by a bind: — the pair reaches one entry as two variable: k=
lines, where hurl’s last-wins would drop the bound value into every later entry too;
(13) a fragment file the engine’s FragmentSupport::scan could not read, or an
annotation it could not attach, positioned in the .hurl file itself.
Fragments skip pass 7: they parse as authored, so the probe instantiation has
nothing to guess. Whether every {{…}} a fragment reads is actually supplied — by a
bind: in scope, by an earlier step’s [Captures], or by the fragment’s own
[Options] variable: — is checked at lower time (§4.4,
proef::lower::unbound_placeholder) — only lowering knows what the preceding steps
captured, and a load-time half-check would be worse than one complete one. The
self-supplied source is scoped to the fragment that declares it: a name another entry
left in hurl’s shared set is the implicit inheritance this check exists to refuse, so
only [Captures] carries a value forward.
4.2 Parse features. gherkin 0.16 (Feature::parse); tags from
Feature/Rule/Scenario accumulate. Localized (# language:) features are
supported and test-covered: the crate strips the dialect keywords, proef
consumes the stripped step text, and a localized outline with Examples
expands like any other (outline detection keys on Examples presence, which is
dialect-independent). Caveat: a localized outline whose Examples block is
omitted cannot be told apart from a plain scenario (the crate’s dialect keywords
are private), so it surfaces as an unbound-step error rather than the crisp
no_examples — a worse message, never a silent pass. # comment lines are plain
gherkin comments (no # key: directive mechanism — variables live in
proef.toml, ADR-0012).
4.3 Bind. For each step (keyword stripped): first macro whose match: pattern
matches wins; ambiguity (2+ matches) is an error listing candidates. Captured {name}
values: trimmed, surrounding quotes shed (quotes preserve inner spaces/commas). Data
table rows | key | value | merge into args; key set by both capture and table → error.
Defaults fill; missing required params → error. Unbound step → exit-2 error with
closest-pattern suggestion (edit distance over pattern literals).
4.4 Lower. Outline expansion (parser does not do it): per Examples row, substitute
<col> in scenario name, step text, docstrings, table cells; ragged rows / unknown
placeholders → parse-time error with line. Background steps prepend to every scenario.
Macro expansion: params bound, ${…} resolved recursively, depth ≤ 8 (captured args
may contain ${…} — spike-verified necessity); $${ escapes; {{…}} passes through
untouched. Assert-only macros (expect:) merge into the previous request entry —
error if none (Then-before-When). Product: Vec<StepBatch> per scenario (contiguous
same-engine runs; batch maximally — split only at optional: boundaries and engine
changes, ADR-0010).
4.5 Emit. Canonical .hurl per scenario (stable formatting, snapshot-tested):
header comment per entry # <file>:<line> — <step text>; # optional markers; sidecar
<slug>.map.json (schema: { entries: [{hurl_lines: [a,b], feature: {file,line,text}, optional, captures: [names], batch: n, step: n}], schema: 1 }); <slug>.vars when ${global:}
or ${secret:} referenced (secrets as names only). Artifact dirs: .proef-runs/<id>/ artifacts/ (per-run) and proef artifacts -o <dir> (stable hand-off).
4.6 Dispatch. Per scenario thread: check token → factory.open(ctx) lazily per
engine on first batch → session.run_batch(batch, world, events, token) in order →
merge outcomes/World → finish() all sessions (reverse open order) → emit events.
optional: batch failure → warnings + continue; else fail-fast within the scenario.
5. proef-engine-hurl internals (verified seam facts inline)
Adapter. Per batch: seed VariableSet from World (insert; secrets via
insert_secret) → parse_hurl_file(&batch_text) → run_entries(&file.entries, &batch_text, Some(&input), &runner_options, &variables, &mut stdout_buf, Some(&listener), &mut logger) with WriteMode::Buffered terms (upstream’s own
threading mode: parallel/worker.rs:76,124-133) → map EntryResults to StepOutcomes
via the sidecar (SourceInfo spans → feature lines) → merge HurlResult.variables back
into the World (typed).
RunnerOptions mapping. Batch-level RunnerOptionsBuilder from config;
per-entry [Options] override batch defaults by clone-then-override
(runner/options.rs:43-58), variable= inserts persist for the rest of the
call — verified semantics, relied upon.
HttpDefaults carries what [http]/[env.<name>.http] express: timeout-ms,
follow-location, max-redirs, insecure, proxy/no-proxy, cacert,
client-cert/client-key, user-agent, cookie-store. Each reaches the
builder only when the project set it, so an unset key leaves hurl’s own
default rather than this engine restating a constant that could drift from
upstream on the next pin bump.
Path-valued fields arrive already resolved — the CLI applies the one-path rule
and core touches no filesystem (ADR-0012). Two keys are exceptions to the
per-entry override rule above, verified against OptionKind rather than
inferred: it has no UserAgent variant (run-wide; an entry opts out with a
User-Agent: request header instead) and no cookie variant at all, so
cookie-store is run-wide with no per-entry spelling whatsoever.
cookie-store = false maps to use_cookie_store(false) (hurl’s
--no-cookie-store, 8.0.0). Verified: hurl enables curl’s cookie engine only
under use_cookie_store (http/client.rs:390), which also means a
cookie_input_file handed over with the store off is silently ignored — so the
engine skips both halves of the §5 round-trip when the store is off. hurl’s own
FIXME there (a handle once given cookie storage cannot have it removed) never
reaches proef: the client is per-call (below), so a handle never transitions
on → off.
Client lifetime (verified). run_entries constructs http::Client::new()
internally per call (runner/hurl_file.rs:169) — fresh libcurl handle: connection
cache and cookie jar do not survive across calls. Consequences implemented: batch
maximally (§4.4); on forced splits, chain variables via HurlResult.variables
(lossless) and, when cookies are in play, round-trip HurlResult.cookie_store →
CookieStore::to_netscape() → temp file → next batch’s
RunnerOptionsBuilder::cookie_input_file (http/cookie_store.rs:66-72,
runner_options.rs:242-244) behind a SessionState struct. Upstream patch #1
(ADR-0003): run_entries(&mut Client) — two internal call sites
(hurl_file.rs:124, worker.rs:124); adopt when accepted, delete SessionState
cookie path.
Thread-safety (verified). No global mutable state in hurl (static mut/
lazy_static/OnceLock-mutable: zero hits); libcurl init is Once-guarded pre-main by
the curl crate; sole FFI global write is libxml2’s error handler set idempotently per
XPath eval — exercised concurrently by upstream itself. Scenario-per-thread is safe.
Cancellation & budgets (ADR-0007). No interrupt support exists upstream (verified);
engine computes batch budget = Σ(timeout × (retry+1)) + intervals + margin, clamped to a
four-hour ceiling (MAX_BATCH_BUDGET — lint-clean values still compose into an unbounded
product, ADR-0007 amendment); watchdog abandons over-budget scenario threads; token
checked between batches only.
Failure detail. Engine errors surface through hurl’s own
DisplaySourceError::description into StepFinished.detail (additive event
field), the console, JUnit, and the GitHub summary.
6. Pack schema v1 (normative field reference)
bind: # pack-scope fragment bindings (ADR-0018); macro and
<var>: "${…}" # step scope override, most specific winning
macros:
<macroName>: # unique across packs; qualify as pack.yaml#name on clash
params: [a, b] # required unless defaulted
defaults: { b: "x" } # optional params
match: "…{a}…" # step-definition pattern; omit → not Gherkin-reachable
description: "…" # docs + desktop palette later
tags: [Domain]
steps: # request steps (each lowers to ≥1 hurl entry)
- name: "…" # entry label (events/console)
optional: true|false # failure → warning (segments the batch)
when: "${expr}" # skip guard: skips when empty or literal false/0 after resolution
retry: { count: N, interval_ms: M } # finite only (lint); → [Options] retry
saveAs: { captureName: global } # promote capture(s) into the World
# (refused with a warning if the value equals a secret)
bind: { <var>: "…" } # step-scope fragment bindings (only with `ref:`)
hurl: | # PRIMARY form (ADR-0004): raw hurl, ${…} lowered first,
… # {{…}} left for run time; validated by parse_hurl_file
# OR structured payload (reserved for future non-hurl engines):
# <kind>: { … } (a future engine's structured payload)
- ref: file.hurl#name # ALTERNATE form (ADR-0018): one `# @proef <name>` entry
# of a scanned fragment file; `{{…}}` supplied by bind:,
# every one bound, captured by an earlier step, or set
# by the fragment's own [Options] variable: (which then
# may not also be bound — option_declared_twice)
- use: pack.yaml#other # composition, with:/inline args; cycle+depth checked
with: { a: "${a}" }
expect: # assert-only macro (Then-steps): merges into previous entry
- status: "${status}" # or raw hurl assert lines: hurl: |‐style fragment (M5)
JSON Schema is schemars-derived from these serde types plus engine-contributed
StepKindSpec fragments; proef schema --add-to injects the editor modeline (a proven
mechanism).
7. Gherkin mapping reference
One scenario = one flow/run-record/artifact set. Background: prepends. Rule: groups
pass through (tags accumulate). Outline/Examples per §4.4. Data tables per §4.3.
Docstrings: reserved for raw request bodies in generic steps (M5). Tags: @tag →
flow tags, --tags filters by a boolean expression (and/or/not/parens,
proef_core::tags); atoms may glob — * any run, ? one char, anchored, a
metachar-free atom staying literal equality — and the one matcher serves
--tags, [run] exclusive-tags and [tag-links]. @skip/@skip:reason and
@quarantine are reserved tags carrying behavior, recognized at the CLI edge —
core never reads tags (ADR-0019, ADR-0014). Step keywords: And/But resolve to
the previous primary keyword (gherkin crate StepType); keyword itself is not matched
against patterns.
8. Variables reference (ADR-0005)
Author-time (${…}, resolved in §4.4, recursive ≤ 8): ${param} · ${env:NAME} /
${env:NAME:-default} · ${url:key} / ${vars:key} (proef.toml [url]/[vars], base +
active [env.<name>] deep-merged; injected — ADR-0012) · ${run:id} (uuid-v7-derived, injected) · ${global:key}
(World read at lower time of the scenario) · ${secret:NAME} (encrypted store; emits
{{secret_name}} + insert_secret) · ${fake:kind} (deterministic from run id and an
occurrence index; the index is an incrementing counter scoped to one scenario — shared
across every ${fake:…} resolve in it, so independent references never collide regardless
of how many a step resolves; a step’s name: label resolves from a rewound copy of the
counter so its Nth ${fake:…} reference reuses the payload’s/when:’s Nth occurrence by
position, not generator kind — reproducing the payload’s own value exactly when the
label’s references mirror the payload’s in kind and order, otherwise surfacing that
occurrence’s own-kind value instead — then restores the real counter to the high-water
mark the replay reached (never below it), so an extra fake the label alone introduces
still reserves its slot and is never reissued;
the counter resets to zero at the next scenario, so the same generator at the same position
in two different scenarios currently coincides; port deterministic NL generators) · $${…}
literal escape. Run-time ({{…}}): hurl captures and
secret placeholders — resolved by the engine. Resolution order within a scope: step args
macro defaults.
9. Diagnostics (ADR-0009)
gherkin Span = 0-based byte offsets, end-exclusive → SourceSpan::new(start, end-start); parser appends a trailing newline when missing — attach the normalized
source text to diagnostics (or clamp); never use LineCol.column (char-counted) in byte
math. Pack YAML: serde_norway error locations; schema-path → YAML-location pass for
lint findings. Engine failures render: feature line + step text + assert detail +
artifact path:span (from sidecar). Every diagnostic carries a stable code
(proef::pack::adjacent_captures, …) for greppability.
10. CLI reference (v1)
proef [--config PATH] [--env NAME] <command> # global: the proef.toml to read, the [env.<name>] profile
proef init [dir]
proef test [file|dir] [--env NAME] [--dry-run] [--tags EXPR] [--jobs N] [--junit path|auto]
[--format json|tap] [--watch] [--scenario NAME] [--scenario-file FILE]
[--run-id ID] [--rerun] [--sarif PATH (with --dry-run)] [--max-fail N] [--shard I/N] [--shuffle] [--meta KEY=VALUE]... [--console MODE]
proef flows [file|dir] [--env NAME] [--format json]
proef macros [file|dir] [--env NAME] [--format json]
proef fragments [file|dir] [--env NAME] [--format json] [--check [--require-annotated]]
proef artifacts [file|dir] -o DIR [--env NAME] [--run-id ID]
proef schema [--add-to FILE…] proef secret set|list|rm
proef explain [run-id] [--format json] proef doctor [--format json]
proef diff [base] [new] [--fail-on-regression] [--format json] # each side: run id, record dir, or events .jsonl
proef flaky [--format json] [--by KEY] [--min-samples N] [--recovery-runs N] [--outage-rate RATE]
# verdicts over the retained run history, keyed by the input fingerprint
proef report [run-id] [-o FILE]
proef fmt <file|dir> [--check]
proef lsp
A path-less test/flows/artifacts resolves [run] suite then the tests/
convention (else exit 2). Exit codes: 0 ok · 1 test failure · 2 user error · 3
system error (typed enum, assert_cmd-pinned). Ctrl-C, SIGTERM and SIGHUP all
take the graceful cancel (ctrlc’s termination feature — a CI job timeout is
a cancellation, not a kill); a second signal while a test/watch run is
cancelling forces an immediate hard exit with code 130 (128+SIGINT, the
shell convention; the handler carries no signal identity, so every second
signal shares the code) — deliberately outside the graceful 0/1/2/3
taxonomy, so it is not an ExitCode variant. Config precedence: built-in defaults
< proef.toml base tables < active [env.<name>] (selected by --env/PROEF_ENV)
< flags; suite variables ${url:key}/${vars:key} resolve from [url]/[vars]
deep-merged with the active env (ADR-0012). Secrets additionally resolve
PROEF_SECRET_<NAME> env overrides before the store, and PROEF_KEY (base64)
overrides the key file — CI decrypts a committed ciphertext store without the key
ever touching disk. --dry-run = §4.1–4.5 including artifact parse-validation;
no engine sessions, no network.
11. State & files
.proef-runs/<run-id>/ → events.jsonl (the record, ADR-0008), run.log (console tee),
artifacts/*.hurl|.map.json|.vars (+ artifacts/assets/<slug>/, the staged file bodies),
timings.json (per-scenario durations, read back by --shard-weights) and inputs.json
(the input fingerprint, proef flaky’s equivalence class) — both derived sidecars, never a
second record — report.html, report.junit.xml (when requested);
[run] keep-runs-bounded
rotation, default 200 (only uuid-named run records rotate; the in-flight run never does).
.proef-state.json — persistent World: atomic temp+rename, 0600. proef.toml — project config:
runner settings ([run] jobs/runs-dir/keep-runs/suite/fragments/setup/teardown/exclusive-tags,
[http] timeouts, [sla] ceilings, [flaky] verdict thresholds) + suite variables ([url]/[vars]) +
run metadata ([meta], ADR-0020) + report tag links ([tag-links]) +
per-environment overrides ([env.<name>]); see docs/CONFIG.md, ADR-0012.
One path rule: a path written in proef.toml resolves against the directory
holding proef.toml; a path typed on the command line resolves against the working
directory. Absolute values are taken as written, and with no config in scope written
paths stay relative to the working directory. This covers suite, fragments,
setup, teardown, runs-dir, .proef-state.json and .proef-secrets.json — so a
project is where its config is, not where the shell is, and running from a
subdirectory reaches the same suite, records, World and secrets as running from the
root.
And its naming dual: a path that reaches an artifact, sidecar, event, report or
diagnostic is spelled relative to that same directory (front::SourceNaming).
Resolution makes a written path absolute; naming makes it project-relative again, so
nothing proef records names the machine that produced it — the four ways to point at
one suite (derived, typed, typed absolute, from a subdirectory) emit one artifact
byte-for-byte, which is what ADR-0010’s same-bytes contract means across machines. A
path that arrives relative is recorded exactly as it arrived; one that lies outside
the project keeps its absolute name, there being no project-relative spelling of it.
DiskSourceProvider (proef lsp) is deliberately outside this: it keys document
identity on absolute names.
12. Parallelism & cancellation
Scenario-per-OS-thread, --jobs bounded (default: available_parallelism, capped by
scenario count); events funneled through the sink (the console reporter buffers per
scenario and replays contiguously); one CancellationToken per run, child per scenario; Ctrl-C graceful /
second Ctrl-C hard (ADR-0007). Global-World writes serialize through the store lock;
scenario ordering within a file is preserved for artifact naming, not execution order —
except that a scenario matching [run] exclusive-tags waits for the pool to drain and
runs alone, which constrains concurrency rather than reordering the queue.
13. Security
Secrets: encrypted at rest (chacha20poly1305), surfaced only
via insert_secret; redaction invariants (never in artifacts/events/reports/logs)
property-tested; a saveAs: global capture whose value equals a known secret is
refused (warned) — .proef-state.json is plaintext and never receives
secret-derived material; sensitive files 0600; proef doctor reports store/key
health. context_dir confines file bodies (hurl’s own sandbox option) to the
scenario’s staged asset root — a directory holding only what proef put there,
so the sandbox is narrower than the suite it used to be. Assets are staged from
beside the source that referenced them (feature for inline, fragment for
ref:), which is hurl’s own per-file rule; references must be plain relative
paths, and one name may come from only one source per scenario. No telemetry.
14. Dependencies (exact at M0; policy: latest stable at adoption, Renovate-managed)
Engine: hurl =8.0.1, hurl_core =8.0.1 (--locked; ADR-0003). Core: gherkin 0.16,
serde 1, serde_json 1, serde_norway 0.9, schemars 1, thiserror 2,
tokio-util 0.7 (default-features = false; CancellationToken only). CLI: clap 4
(derive, env), miette 7 (fancy), uuid 1 (v7), notify =8.2.0, ctrlc,
chacha20poly1305 rpassword base64, quick-junit, toml. LSP: lsp-server 0.7,
gen-lsp-types 0.11 aliased as lsp-types with its url feature (proef-lsp’s stdio
transport, wired into proef lsp; ADR-0017 amendment). Fixture/harness (dev):
tiny_http (ADR-0011 — axum conflicts with the tokio-runtime ban), libtest-mimic.
Engine runtime: tempfile (Netscape cookie round-trip between batches, §5).
Dev: insta assert_cmd predicates proptest tempfile quick-xml + cargo-fuzz
targets; openssl-sys rides as the engine’s vendored-openssl feature carrier. Synthetic
data (${fake:*}) is a dependency-free SplitMix64/FNV implementation in-core — the
fake crate was not needed. Datetime uses jiff, never chrono, in our own code
(hurl’s internal chrono is its business) — currently jiff is a dev-only dependency
of the fixture’s /health identity; the sans-IO core still reads no clock (injected
timestamps only). Banned: serde_yaml/serde_yml, chrono (ours), reqwest
(superseded), maybe-async, async-trait (v1). Build prereqs (doctor-checked): Debian
build-essential pkg-config libssl-dev libcurl4-openssl-dev libxml2-dev libclang-dev;
macOS: Xcode CLT.
15. Conventions
Toolchain pinned to latest stable adopted at its x.y.1 point release —
~3-4 weeks after x.y.0, so a new minor waits out its first patch (1.97.1 at
writing; tools and third-party crates track latest immediately, and exact pins
like hurl outrank everything). rust-toolchain.toml cites this section as its
authority, so the policy is stated here rather than only in RELEASING.md,
where the correction first landed (R18-2): an unwritten policy contradicting
the written one is a docs defect, and fixing it in two files while the
normative spec still said “always latest stable” left the contradiction in the
source of truth. Edition 2024; resolver 3; workspace
lints (the legacy suite table); CI gates (PR): cargo fmt --check, clippy --all-targets --all-features -D warnings, cargo nextest run, cargo test --doc,
RUSTDOCFLAGS="-D warnings" cargo doc, cargo deny check, cargo machete, zizmor,
xtask docs-check, proef doctor, fuzz smoke + xtask public-api (nightly rustdoc);
cargo audit runs on the nightly schedule. Automation in xtask (+ just aliases); no shell scripts for logic.
proef — Implementation Plan
Status: M0–M5 delivered, and everything after them — external config and
environments (ADR-0012), the v0.6–v0.8 correctness series, v0.9.0, and named hurl
fragments (ADR-0018). M6 (future engines) remains unscheduled. CLAUDE.md’s Status
block is the running ledger; the milestone sections below describe the plan as it was
executed, not the current frontier. · Date: 2026-07-28, status refreshed 2026-08-12.
Normative design: TECH-SPEC; decisions: ADRs. Sizes are
t-shirt (S ≈ days, M ≈ small weeks, L ≈ multi-week) — deliberately not fake-precise.
0. Guiding principles
The porting rule: when the prior spike proved a feature, understand
the feature and implement it the cleanest way in this architecture — no copy-paste, no
compatibility with spike code. Every milestone lands green on all CI gates. Core stays
pure (no IO/clock/rand — TECH-SPEC §4). Nothing merges without its tests
(TESTING-STRATEGY). The spike (research/) is evidence, not a starting codebase.
Global definition of done (every milestone)
fmt, clippy -D warnings, nextest, doctests, rustdoc -D warnings, deny, audit all
green · new behavior has unit + (where applicable) snapshot/property tests · no unwrap/
expect in library code paths (CLI main excepted) · public items documented · CHANGELOG
entry · exit codes unbroken (assert_cmd suite).
M0 — Foundations (size M)
Objective: compilable, gated, empty-but-shaped workspace with the seam in place.
Tasks:
- Init repo
proef/; virtual workspace (resolver 3),workspace.package+workspace.dependencies+workspace.lints;rust-toolchain.tomlpinned to current stable (1.97.1 at writing);deny.toml, nextest config, README. - Crates:
proef-core,proef-engine-hurl(empty adapter),proef-cli(clap skeleton),xtask(+justfilealiases). Reserve crates.io names (0.0.0 placeholders,publish = falselocally thereafter). - Core scaffolding:
ExitCodeenum +CoreError/EngineErrortaxonomy (ADR-0009);Eventenum v1 +EventSink(ADR-0008);World/Valuetypes +GlobalStore(atomic temp+rename) (ADR-0005);EngineFactory/EngineSession/StepKindSpec/DoctorChecktraits (ADR-0002);CancellationTokenplumbing (ADR-0007). - CI: gates workflow + a stub canary job (builds
proef-engine-hurlagainst hurl=8.0.1— becomes the real canary in M4); Renovate config (pins grouped; hurl excluded from auto-bump). proef-cli:doctor(native-lib checks from enginedoctor()— first proof the capability hook works),--version, exit-code integration tests.
Acceptance: workspace builds on Linux+macOS CI; proef doctor reports libcurl/
libxml2 status via the engine-contributed check; assert_cmd pins exit codes; all gates
green. Proves: ADR-0002 seam compiles and registry assembly works.
M1 — Front end: packs, binding, Gherkin, lowering (size L)
Objective: .feature + packs → validated, lowered scenarios; --dry-run without
artifacts.
Tasks:
- Pack model (serde_norway,
deny_unknown_fields) incl.hurl:raw blocks,expect:,use:/with:,when:,retry:,saveAs:; schemars derivation + engineStepKindSpecfragment merge;proef schema [--add-to]. - Validation passes 1–8 (TECH-SPEC §4.1) with miette diagnostics + stable codes;
finite-retry lint; probe-instantiation parse of hurl blocks via
hurl_core. - Matcher:
{name}tokenizer + leftmost matcher + guard rails (cucumber-expression semantics; property tests: no-panic, quote round-trip, adjacent-capture rejection). - Gherkin: parse, tags, Background, Rule pass-through, outline expansion, data-table merge; binding with ambiguity detection + closest-pattern suggestions.
- Lowering: macro expansion (cycle/depth), recursive
${…}resolver (depth 8,$${escape; property + fuzz targets), Then-merge (expect:→ previous entry), batch segmentation (maximal; splits atoptional:/engine change). proef flows,proef test --dry-run(no emit yet), corpus test overtests/.
Acceptance: the four 500-series features (+ seeded error corpus) dry-run with line-accurate diagnostics; property/fuzz targets in CI (fuzz smoke = N seconds, full = nightly). Proves: PRD US-2/3/6 front-end half.
M2 — IR, emitter, artifacts (size M)
Objective: lowered scenarios → canonical .hurl + sidecars; --dry-run complete.
Tasks:
- Canonical emitter (stable formatting rules) +
# optionalmarkers + per-entry feature-ref comments; line-map construction. - Sidecar
<slug>.map.json(schema v1) +<slug>.vars;proef artifacts -o DIR; per-run layout under.proef-runs/<id>/artifacts/. --dry-rungains artifact parse-validation (parse_hurl_fileon every emitted file — the real parser as the validator).- insta snapshot suite: features+packs → artifacts + sidecars (golden corpus).
Acceptance: spike parity — the 500-series features emit artifacts that stock hurl parses (checked in CI via the canary toolchain image); snapshots reviewed. Proves: ADR-0010 emit half; US-7 static half.
M3 — engine-hurl: embedded execution (size L)
Objective: proef test runs scenarios end to end via embedded hurl.
Tasks:
- Adapter:
VariableSetseeding (World + secrets viainsert_secret),parse_hurl_file+run_entrieswith Buffered terms +EventListener→ events;EntryResult→StepOutcomemapping via sidecar spans. - RunnerOptions mapping (config → builder; per-entry
[Options]override relied on as verified);HurlResult.variablesmerge-back;saveAs: globalpromotion; typed Value bridging. - Segmentation runtime:
optional:warn-and-continue;SessionStatecookie round-trip (Netscape temp file) for split scenarios; variables chaining. - Parallelism: scenario threads +
--jobs; budgets + watchdog + token checks; Ctrl-C graceful/hard paths. - Reporters v1: console BDD tree (attempts/timings), JSONL event record, run-record
rotation;
--output json. - Fixture server (
tiny_http, ADR-0011) + integration suite: success/4xx/retry-delay/auth/ malformed-JSON/optional/World-chaining/cancellation-budget cases.
Acceptance: the 500-series features run green against the fixture with prose unchanged (US-1); failure demo maps to feature line + artifact span; exit codes correct under pass/fail/user-error/system-error; cancellation bounded-time test passes. Proves: ADR-0001/0005/0007 runtime; US-1/4/5/9.
M4 — Upstream tracking hardened + CI reporters (size S–M)
Objective: riding upstream is a runbook, not a risk; CI outputs complete.
Tasks:
- Canary job real: build+test against next hurl release (scheduled + on release);
failure opens an issue with the diff of
HurlResultbehavior. - Thin-fork rehearsal: apply a scratch one-commit patch via
[patch."crates-io"]from the fork tag, build, revert — documents the mechanics; draft upstream PR #1 (run_entries(&mut Client)) from the verified two-call-site change. - JUnit (quick-junit) + GitHub job summary reporters (
--junit autounder GITHUB_ACTIONS); reporter-stack decorators (Normalize/Summarize) formalized.
Acceptance: canary catches an artificially-pinned older/newer hurl mismatch in rehearsal; JUnit consumed by CI UI. Proves: ADR-0003; US-8/11; M-4 metric path.
M5 — Breadth + integrations (size M)
Objective: the conveniences that make proef the daily tool.
Tasks: multipart/form/docstring bodies through packs; expect: raw-hurl assert
fragments; more [Options] exposure (delays, location); --watch (notify 8.2.0);
proef explain; proef secret set|list (encrypted store port); ${fake:*} NL
generators; libtest-mimic harness (one Trial per scenario; nextest contract:
--list --format terse, --exact --nocapture) + docs for IDE use; proef fmt for
pack hurl blocks.
Acceptance: US-10/12 green; nextest runs the suite; watch-mode demo. Proves: ADR-0008 harness leg; PRD v0.3 scope.
M6 — More engines (future; sized when scheduled)
A future non-hurl engine — its step vocabulary carried as structured step payloads rather
than raw hurl. Structural acceptance test: git diff --stat proef-core is empty.
Mixed-engine 500-series suites become runnable.
(Note 2026-07-29: a driver bringing its own async runtime would conflict with the
tokio-runtime ban — pick a sync driver at sizing time. The core’s structured-payload paths
are already exercised by tests; no core work is expected.)
5. Sequencing & parallelization
M0 → M1 → M2 → M3 strictly ordered (each consumes the previous stage’s types). Within M1, tasks 1–3 parallelize with 4; M2.4 can start as soon as M2.1 emits. M4.3 reporters can start during M3.5. Docs (GETTING-STARTED + AUTHORING, delivered 2026-07-29) draft during M2–3 and finalize at M5 (deliberately after the schema stabilizes).
6. Risk register
| Risk | L | I | Mitigation / trigger |
|---|---|---|---|
| hurl minor release breaks the seam | M | M | exact pins; canary (M4); thin-fork shim; pinned seam integration test |
run_entries #[doc(hidden)] churn | M | M | same as above + upstream PR #1 conversation opens a stability dialogue |
| Native build prereqs trip up a machine | M | L | doctor first-run UX; README one-liner; CI images pre-baked |
| Segmented scenarios lose connections/cookies | L | L | batch-maximally; SessionState; upstream patch #1 erases it |
| Runaway scenario under retries | L | M | finite-retry lint; budgets + watchdog (ADR-0007); CI timeout |
| Pack schema churn post-v1 | M | M | schema: 1 field; additive-only until v1.0; snapshots catch drift |
| Event-schema consumers break | L | M | versioned events; JSONL replay tests |
| gherkin-crate stagnation returns | L | M | active again (0.16); parser is replaceable behind core’s parse stage |
7. Runbook — absorbing a new hurl release
- Canary red/green report arrives (scheduled job). 2. Read upstream CHANGELOG diff.
- Bump pins on a branch (
=X.Y.Z,--locked), run full suite + snapshots. 4. If breakage: fix adapter; if upstream regression or removed seam: add minimal patch on the fork branch, consume via[patch], open upstream PR, note in ADR-0003 log. 5. Merge; tag; update TECH-SPEC §14 versions. 6. If the fork carries patches: rebase them onto the new release tag; drop any that merged upstream.
8. Day-one checklist
git init proef && cd proef → commit toolchain+workspace manifests (M0.1) → cargo new
the four crates → copy workspace lints/deny/nextest configs → wire CI → first green
pipeline → open M0 tracking issue with this plan’s task list. Suggested first PR
sequence: M0.1+M0.2 together, then M0.3 split by module, then M0.4/M0.5.
proef — Testing Strategy
Status: normative · Date: 2026-07-28 · tools: cargo-nextest, insta, proptest, cargo-fuzz, assert_cmd, tiny_http fixture. Everything below is device-free and CI-green with no external network.
1. The layers
Unit (every crate): matcher tokenization/matching edge cases; resolver escapes and depth cap; Then-merge rules; batching/segmentation boundaries; sidecar math; World error→exit-code mapping.
Property (proptest): matcher — arbitrary patterns/text never panic, valid
pattern+generated text round-trips captures; resolver — $${…} escape round-trip,
resolution is idempotent once fully resolved, depth cap always terminates; secret-mask
invariant — for arbitrary events/reports containing a known secret value, rendered
output never contains it, nor any of its derived encoded forms (base64 both
alphabets ± padding, hex both cases, percent-encoding, JSON-string escape —
ADR-0005 as amended), with a companion property pinning that text free of the
secret and its forms passes through untouched; World — snapshot/restore is an involution;
fragment scanner (proef-engine-hurl) — over generated hurl files, every
reported line lies inside the file, entries are accounted for exactly once, starts
are ordered and distinct, and no fragment’s text runs into the entry after it.
That last one is the entry-boundary arithmetic’s whole job, and it is asserted
because a draft without it passed while the boundary was deliberately broken.
The scanner is proptested rather than fuzzed on purpose: it needs hurl_core, and
cargo dependencies are package-level, so putting it in fuzz/ would compile hurl
for every target there and drag native libraries into a job that has none.
Fuzz (cargo-fuzz, nightly job + PR smoke): fuzz_match_pattern (pattern×text),
fuzz_resolve (template strings), fuzz_tag_expr (tag expressions),
fuzz_pack_load (YAML bytes → loader must error, never panic), and
fuzz_fragment_binding (a pack against a real corpus: ref: resolution, unread
bind: keys, a bind: colliding with a variable the fragment supplies itself).
Parser-adjacent hand-written code is exactly where fuzzing pays.
Both loops take their target list from cargo fuzz list, never a list written
into a workflow. The names used to be spelled out in ci.yml and nightly.yml,
so a target ran nowhere until both were edited and nothing failed to say so.
fuzz_fragment_binding is structure-aware — it builds a well-formed pack and
corpus from the input rather than hoping the fuzzer discovers one. That is a
measured choice: a byte-oriented version never resolved a single ref: in 1.45
million runs, because reaching those rules means finding valid YAML and a matching
corpus name at once. When adding a target that needs structure, verify it reaches
the code by probe — panic on the condition under test, run briefly, confirm it
fires — because a target that compiles and finds nothing reads exactly like a
target that compiles and finds no bugs.
fuzz/ is its own workspace (the root Cargo.toml excludes it, since fuzzing
needs nightly), so no root-workspace command compiles it: a changed
proef-core signature breaks the targets while every root-workspace gate
stays green, leaving the fuzz jobs as the only signal.
cargo check --manifest-path fuzz/Cargo.toml --all-targets runs on the pinned
stable toolchain in seconds, so the gates job carries it — earlier than the
fuzz smoke, and on both gate platforms.
Snapshot (insta): emitter — golden corpus of (features + packs) → artifacts +
sidecars, byte-stable (the canonical-format compatibility surface, ADR-0010);
diagnostics — rendered miette output for the seeded error corpus (every validation pass
in TECH-SPEC §4.1 has at least one golden failure); proef schema output; event-stream
JSONL for a reference run (with injected clock/run-id — core purity makes this
deterministic).
Integration (fixture server): synchronous tiny_http dev crate
(proef-fixture) modeled on the spike’s fixture — not axum as originally
written: axum’s tokio requirement conflicts with the workspace’s no-async-runtime
ban (ADR-0006/0007 + deny.toml), and a sync fixture keeps that invariant
binary-wide (errata 2026-07-28, M3). Endpoints, extended: bearer-auth endpoints, search, create (201/422 paths), delayed
push-visibility (exercises retry for real), cookie-setting endpoints (exercises
SessionState round-trip), slow endpoint (exercises budgets/watchdog), malformed-JSON
endpoint. Suite covers: green path (the four 500-series features), capture chaining,
World/global across scenarios, optional: warn-and-continue, cancellation (token
cancel mid-run completes within budget, reports written), parallel --jobs determinism
(event Normalize), artifact↔execution same-bytes assertion (hash the emitted file and
the text handed to parse_hurl_file).
Fragments (ADR-0018): crates/proef-cli/tests/fragments.rs builds a self-contained
project per test — its own proef.toml, corpus and pack in a temp dir — because the
reference corpus under tests/ is config-independent by design (several tests run it from
a temp cwd with settings passed by environment variable and no proef.toml in scope), so
anything needing [run] fragments cannot live there. That is also why four diagnostic
codes are covered here rather than in tests/errors/ (DIAGNOSTICS.md says which).
The headline case runs one file under both runners: proef test against the fixture,
then stock hurl invoked on the same bytes with an equivalent variables file, asserting
the corpus comes back byte-identical. The engine is embedded, so a hurl binary is not a
build requirement — that half skips with a printed note when none is on PATH rather than
being faked. Provenance is asserted at both ends: the JSONL record for the event-driven
readers, and --junit for the RunSummary-driven ones, since those are fed by a second
copy that a green suite would not otherwise exercise.
LSP over stdio (crates/proef-cli/tests/lsp_stdio.rs): the real binary, spoken to as
an editor does. This is the only place proef.toml → DiskSourceProvider → document URI
is exercised end to end: the proef-lsp unit tests inject absolute source names through a
fake provider, so a config-layer change can break every go-to-definition while they stay
green — which has happened. These tests canonicalize their temp root, because on macOS
a tempdir is /var/… whose real path is /private/var/…, and without that any
cwd-relative path logic silently no-ops and the test passes without reaching the behaviour.
Corpus: --dry-run over every .feature in tests/ —
the suite’s own features are the regression corpus.
Documentation (xtask docs-check + crates/proef-cli/tests/docs.rs): the docs make
claims a machine can settle, so they are settled mechanically rather than by review.
docs-check reads files, and does six things: every workspace crate appears in
TECH-SPEC §2 and CLAUDE.md; every ADR file appears in the decision log; every diagnostic
code the workspace emits has a DIAGNOSTICS.md row and vice versa; every release names
each kind of change once (repeats accumulate by appending, which is how a changelog
gets written); every relative link resolves; and every fenced toml/yaml example parses
with the product’s own parsers, so the check means “proef would accept this”, not
“some parser would”. tests/docs.rs needs a built binary and therefore lives
with assert_cmd: it asks clap whether every documented command and long flag exists.
The split is a rule, not an accident, and each half states it in its own header — a check
that reads files belongs in docs-check even when a test would be easier to write, because
the doc-only CI step is the fast one and a file-reading check placed in the test suite
silently stops running there.
Both were written against defects that had already shipped — an ADR whose first example
could not load, and a row marked shipped that named a --html flag which never
existed. In each the surrounding prose was correct, which is precisely what a careful
reader does not catch. Two scoping rules keep them honest rather than noisy: only the
indexed corpus is linted (docs/superpowers/ is a dated archive, and editing history
to satisfy a checker is the wrong direction), and command detection is restricted to
code spans and fenced blocks — prose says “proef discovers packs”, and treating that
as an invocation produced sixty false positives against four real ones. Names the docs
discuss as proposals are listed explicitly in tests/docs.rs, so adding one is a
decision rather than the check going quietly soft.
CLI (assert_cmd): exit codes 0/1/2/3 pinned per command and failure class;
--format json schema-checked; --junit well-formed (quick-junit round-parse).
Canary (M4): scheduled + on-release job builds against the next hurl version and replays the integration suite; red = issue with behavior diff, pins never auto-move (runbook: IMPLEMENTATION-PLAN §7).
2. What is deliberately NOT tested here
Hurl’s own HTTP semantics (asserts, filters, templating execution) — that is upstream’s test surface; proef tests the adapter contract (options mapping, variable bridging, span mapping, segmentation) against the fixture instead of re-verifying hurl. This is a direct consequence of ADR-0001 and the reason the differential-oracle harness from the research phase was retired.
3. CI matrix & gates
Linux (ubuntu-latest, prereqs pre-baked) + macOS on every PR; Windows weekly (vcpkg
libs) while the port stabilizes, then per-PR (port green 2026-07-28: VCPKG_ROOT export,
hurl’s crates.io-missing icon supplied in CI, /-normalized path identifiers). Gates: fmt, clippy -D warnings, nextest (all crates),
doctests, rustdoc -D warnings, deny, cargo-machete, zizmor (workflow static
analysis), xtask docs-check, proef doctor smoke, public-api snapshot, fuzz smoke
(30 s/target), corpus dry-run, CLI suite, and — on a pull request — the changelog-entry
check (§8). The complexity ratios run as their own step, alone, for the reason §7 gives. Snapshot tests (insta) run inside nextest —
a drifted snapshot fails there, no separate step. Nightly: full fuzz (10 min/target),
canary, cargo-audit (advisories against unchanged code — deny covers PRs).
Coverage: measurable on demand, not gated in CI (P13’s local half). just cover
runs cargo llvm-cov nextest over the workspace (just cover-html for a browsable
report, just cover-lcov for a CI service’s lcov); the number today is ~90% line
coverage of the unit + integration suites (doctests excluded — nextest does not run
them). When a CI coverage job lands it must be a ratchet, not a fixed threshold —
the 2026 norm and the only kind that suits a pre-1.0 codebase: fail a PR only if
coverage drops, never on an arbitrary floor, and keep it informational (a PR comment)
rather than a hard merge gate. A fixed percentage gate is explicitly the wrong shape
here; it punishes honest additions of hard-to-cover error paths and invites coverage
theater. The xtask binary’s low number is expected — it is automation exercised by
running it, not by unit tests.
4. Test data management
tests/features/ — the real suite (also corpus input). tests/errors/ — seeded broken
features/packs, one file per diagnostic code, name = expected code (golden snapshots).
Insta snapshots live next to their suites (crates/proef-cli/tests/snapshots/),
reviewed via cargo insta review. Fixture data is
generated in-process (no committed binary blobs beyond one JPEG for multipart, M5).
5. Determinism rules (make flakes structural, not cultural)
Core purity (no IO/clock/rand — TECH-SPEC §4) means every non-integration layer is bit-deterministic by construction. Integration layer: fixture delays are token-driven (visibility timestamps), not sleep-raced; retry tests assert attempt counts, wall time only as generous upper bounds; parallel tests assert on Normalized event order, never raw interleaving. Any test needing “now” receives it as a parameter.
Retry-until-green is the anti-pattern, and that is why proef ships no
scenario @retry. A scenario-level retry is the headline feature of several
runners and is deliberately absent here: re-running a test until it passes
hides precisely the defects worth finding. A bug that fails one run in four
survives three retries 99.6% of the time (1 − 0.25⁴), so the suite reports
green while the product is broken for a quarter of its users. proef’s shape is
detect-then-quarantine: proef flaky returns a verdict over run history,
@quarantine stops a known flapper gating the build while keeping it visible
in every sink (ADR-0019), and per-step retry: covers the case that is
genuinely polling — a resource that becomes visible on the Nth attempt —
rather than rerolling a verdict. proef’s own suite is held to the same rule: a
red test here is reproduced and filed, never re-run until it cooperates and
then forgotten.
6. Every diagnostic code is named by a test
DIAGNOSTICS.md calls codes “a contract: they never change meaning”. A contract
with nothing holding it to it is a wish — 23 of 75 were in that state when the
rule was written: reachable in production, documented, and exercised by nothing
at all, not even an assertion on their message text.
source_guards.rs enforces it. A code counts as covered when either a seeded
tests/errors/<area>__<name>/ directory exists (the corpus driver dry-runs it,
so the rendered diagnostic is exercised end to end) or the literal code string
appears in a test. Naming the code, not matching the prose — the wording is
expected to improve, while the code is the part that promises not to change.
Two codes are exempted by name with recorded reasons (source::unreadable,
config::unreadable need a file the process may stat but not read, which CI
runners do not reproduce because they run as root). The guard checks its own
exemption list too: an exemption that outlives its code silently excuses
nothing.
Reaching a defensive guard is worth the effort rather than a reason to skip it.
lower::kind_unrouted fires only when the engine registry and pack validation
disagree, so its test makes them disagree; lower::expansion_too_deep sits
behind pack validation’s identical limit, so its test bypasses validation with
load_collecting — the only way to hand lowering a graph validation would have
stopped, and therefore the only way to prove the second line of defence still
works.
7. Complexity claims are asserted as ratios, never as benchmarks
A published performance claim is a claim like any other, and this project has now watched four separate ones decay in prose. The guard for a shape claim — “linear in the macro count”, “~2× per doubling” — is a ratio between two input sizes, not a stopwatch against a threshold:
- A ratio tests what was actually promised. The claim is a shape; a shape is a ratio.
- The separation is wide enough to be safe.
validation_cost_stays_linear_in_the_macro_countobserves ~2.05× against a bound of 3.0; restoring the pre-#138 quadratic shape measures 4.01×. Take the minimum of several interleaved samples — scheduler noise only ever adds, so the fastest observation is the closest to the work actually done — and assert the smaller load was slow enough to time at all, or the ratio is meaningless.
A timing test runs alone, or it does not run. These are #[ignore]d and have
their own CI step and just perf; nothing else shares the machine. The first
version of this section claimed the opposite — that a ratio “survives a shared
runner” because load inflates both sides and cancels — and shipped a test that
failed on its second full-suite run. Measurement: 2.05× alone, 3.09× under
nextest’s full parallelism. The larger input has the larger working set, so
memory-bandwidth contention penalises it more; the ratio drifts rather than
cancelling, and interleaving cannot fix a systematic effect. nextest’s
test-groups bound concurrency within a group, which does not isolate one
from the rest of the suite — so #[ignore] plus a dedicated invocation is the
only mechanism that actually delivers isolation.
Benchmark frameworks were considered and are deliberately absent. iai-callgrind
is the right tool for gating in CI, because instruction counts ignore runner noise
entirely — but it needs valgrind, making it a gate the maintainer cannot reproduce
on macOS. criterion and divan measure wall time, which is the same noise regime
as the ratio test while also adding a dependency tree to a workspace that audits
every edge.
8. A change that lands records itself
RELEASING.md states that every landed change adds an [Unreleased] line in the
commit series that lands it. Nothing enforced that, and the rule was broken exactly
once — by the series that added the guard for the changelog’s shape. A rule whose
only enforcement is a sentence in another document is a rule with a known decay rate,
which is the same finding this suite keeps re-deriving.
So a pull-request job asks one question: did any crates/** or xtask/** .rs file
change without docs/CHANGELOG.md changing too? If so it fails, naming the files and
quoting the rule. [no changelog] in the PR title waives it.
It was sized before it was written, because a gate with a high false-positive rate trains people to reach for the waiver and is then worse than nothing. Across the 21 source-touching merges preceding it the rule would have fired once — on the one commit that actually broke it. Pure-test and pure-performance changes all carried an entry already, so “source changed” tracks “worth recording” closely here. That is a measurement of this repository’s habits rather than a general law, and the waiver exists for where it stops holding.
The check reads a diff rather than files, so it is neither a docs-check task nor a
tests/docs.rs test — it lives in the workflow, which is the only place the base
commit is known. The PR title reaches it through env, never interpolated into the
shell body: a title is attacker-controlled text, and ${{ … }} inside run: is a
template injection that zizmor flags.
Runbook — thin-fork patching (ADR-0003 tier 2)
Rehearsed 2026-07-28 (M4). The steady state carries zero diff; this runbook is the mechanics for the moment a release breaks the seam or a small change is needed before upstream accepts it.
The drill (as rehearsed, local-path variant)
-
Obtain the pinned source (fork tag in the real flow; the vendored registry copy suffices for a drill):
cp -R ~/.cargo/registry/src/index.crates.io-*/hurl-8.0.1 /tmp/hurl-fork # apply the minimal patch (one commit on the fork branch in the real flow) -
Wire the override at the workspace root (
Cargo.toml):[patch.crates-io] hurl = { path = "/tmp/hurl-fork" } # drill # hurl = { git = "https://github.com/<org>/hurl", tag = "8.0.1-proef.1" } # real -
Verify cargo resolves the fork, then build and run the full suite:
cargo tree -p proef-engine-hurl | grep "hurl v" # must show the path/git source cargo build -p proef-engine-hurl && cargo nextest run -
Revert: remove the
[patch.crates-io]block,git checkout Cargo.lock, confirmcargo treeshows the registry source again.
Rehearsal result: override resolved, engine + suite built green against the patched path, revert restored registry resolution. Elapsed ≈ one hurl rebuild.
The real flow (when a patch is actually needed)
- Fork
Orange-OpenSource/hurl; branchproef-patches-<version>off the release tag; apply the minimal diff (one commit per logical patch). - Tag
X.Y.Z-proef.N; consume via the git[patch.crates-io]form above. - Open the corresponding upstream PR immediately (tier 3 — the fork’s diff must trend back to zero); note the patch in ADR-0003’s log.
- On each upstream release: rebase the branch onto the new tag, drop merged patches, re-run the canary, move the pins per IMPLEMENTATION-PLAN §7.
Pin-bump checklist — hurl 8.1 watch items (recorded 2026-08-16)
Behavior changes already on hurl master that the canary cannot see — compile-and-test stays green while a guarantee shifts. Work each of these when the pins move to 8.1 (IMPLEMENTATION-PLAN §7); each names the fixture to add.
The structural reason these need a list at all: the fragment scanner matches
OptionKind narrowly (if let OptionKind::Variable — one arm, by design,
so a semver-allowed variant addition is not a compile break). New [Options]
in a fragment file are therefore silently accepted, not flagged. That is the
right default for options with local effect, and exactly wrong for the first
item below.
variables-file:(upstream #2021) — check the sandbox before accepting it. Upstream opens the named file with a rawFile::openagainst process CWD, noContextDirconfinement. A fragment corpus is foreign by design, so a corpus file sayingvariables-file: ../../secrets.envwould read a file outside the project on proef’s behalf. On bump: decide refuse-or-flag (option_declared_twice’s family machinery fits), and add a fixture — a.hurlwith an escapingvariables-file:must not silently read the target.- Cross-host cookie strip (upstream #5118, landed) — a redirect across hosts stops forwarding cookies. Fixture: a fixture-server redirect pair asserting which cookies arrive, so the behavior flip shows up as a diff in our suite rather than as a user’s broken auth flow.
--file-rootresolution change (upstream #2830, watch) — multipart asset paths may move from CWD-relative to hurl-file-relative. This is proef’s multipart seam: artifacts are emitted to a different directory than the pack they came from, so relative asset paths are exactly the bytes that would change meaning. Fixture: a multipart scenario whose asset path only resolves under one of the two rules.Value::Duration(upstream #3519, watch) — a captured duration re-rendered into a template may change its string form. Fixture: capture a duration-typed value, splice it into a later request, snapshot the bytes.- New
[Options]variants generally (no-header,http2-prior-knowledge,fail-with-body,no-jsonpath-coercion, …): the scanner accepts them silently (above). On bump, sweep the new variants once and sort each into “local effect, fine” or “needs thevariables-file:treatment”.
Currently drafted patches
docs/upstream/0001-run-entries-reusable-client.patch—run_entriesaccepts&mut Client(verified two-call-site change; erases per-segment connection + cookie costs, deletes proef’sSessionStatecookie round-trip once adopted). Applies cleanly to 8.0.1 and compiles (verified in the M4 drill). PR text:docs/upstream/0001-PR-DESCRIPTION.md.
ADR-0001 — Embed hurl’s crates in-process as the API engine
Status: Accepted · Date: 2026-07-28
Context
The API engine must execute HTTP tests with Hurl semantics. Constraint history: the
original “entirely in Rust” rule (which forbade C-linked deps and favored a pure-Rust
reimplementation) was relaxed to “mostly Rust — C system libraries acceptable within
margin”, and the product intent was fixed as “a wrapper / Gherkin adapter over hurl,
riding upstream hurl as it improves”. Verified facts: hurl’s runner::run_entries is
the seam hurl’s own parallel workers use (buffered stdio, per-entry results with spans,
captures, libcurl timings); hurl links libcurl + OpenSSL + libxml2 and needs libclang at
build; hurl_core alone also links libxml2; the crates break API in minor releases
(issue #3846); the format itself is spec’d and stable. A working spike cross-ran
generated artifacts under a prototype pure-Rust runner and the stock hurl CLI: the first
cross-check caught a real semantic divergence (contains = element equality on
collections, not substring) — demonstrating the standing cost of reimplementation.
Decision
proef-engine-hurl embeds the hurl crates, pinned exactly: generate .hurl text from
the IR, hurl_core::parser::parse_hurl_file, then hurl::runner::run_entries with
WriteMode::Buffered terms, an EventListener for progress, and VariableSet in/out.
No pure-Rust HTTP reimplementation ships; no hurl subprocess is required at runtime.
Consequences
Positive: hurl semantics by construction (zero drift class); full assert/filter/XPath/
cookie/redirect/HTTP-2 surface available to packs on day one; rich in-process results
with source spans mapped back to .feature lines; ~a third less engine code than the
reimplementation plan. Negative: build prereqs on every machine (libcurl-dev libxml2-dev libclang pkg-config; macOS ships the libs) — mitigated by docs + proef doctor; dynamically-linked release binaries (no musl static); crate API instability —
mitigated by ADR-0003; run_entries is #[doc(hidden)] (semi-blessed seam) — covered
by a pinned integration test.
Alternatives considered
Pure-Rust native engine + hurl-CLI differential oracle (the original recommendation
under the zero-C rule): zero C deps and static binaries, but permanent semantic-drift
liability and a differential harness to maintain; recorded in research/ as the path
back if static distribution ever becomes a requirement. Transpile + hurl CLI
subprocess: full fidelity, least code, but an external binary dependency, no in-process
events, and clunky cross-scenario variable threading; its remnant lives on as the
upgrade-canary and proef artifacts hand-off. Divergent fork of hurl: rejected —
upstream keeps the format low-level by philosophy (issue #2090) and a divergent fork
defeats riding upstream improvements (see ADR-0003 for the thin-fork nuance).
ADR-0002 — Multi-engine core: factory/session seam, step-kind routing, batching
Status: Accepted · Date: 2026-07-28 (amended 2026-09-01 — the core’s entry grammar is a named closed set; see the Amendment below, and its 2026-09-10 correction)
Context
Requirement: all tests are Gherkin; the parser dispatches to pluggable engines — API
(hurl) now, a future non-hurl engine possible later behind the same seam — a
factory/session seam (multiple engine implementations behind one trait, a step’s
kind-prefix routes it to its engine, one shared variable scope). Balanced-architecture
stance: deliberate seams for a future engine, no gold-plating. Ecosystem survey: probe-rs’s ProbeFactory/DebugProbe
split is the closest production analog; sqlx registers compiled-in drivers explicitly;
dispatch cost at batch granularity is noise, so dyn vs enum is decided on coupling, not
performance (enum_dispatch would couple core to every engine crate — wrong direction).
Decision
Two traits in proef-core; engines implement both; the CLI assembles the registry.
#![allow(unused)]
fn main() {
pub trait EngineFactory: Send + Sync {
fn id(&self) -> &'static str;
fn step_kinds(&self) -> &'static [StepKindSpec]; // pack namespace + schema fragment
fn doctor(&self) -> Vec<DoctorCheck>;
fn open(&self, ctx: &ScenarioCtx) -> Result<Box<dyn EngineSession>, EngineError>;
}
pub trait EngineSession: Send {
fn run_batch(&mut self, batch: &StepBatch, world: &mut World,
events: &EventSink, cancel: &CancellationToken) -> BatchResult;
fn finish(&mut self) -> Result<(), EngineError>;
}
}
Routing: a macro step’s kind names its engine (http: → engine-hurl; other kind prefixes
reserved for a future non-hurl engine). A lowered scenario is an ordered heterogeneous step list; the core dispatches
contiguous same-engine batches in order. The World is the interop bus between
batches and engines. Sessions are per-scenario, opened lazily, torn down in finish
(+ Drop backstop); engines may hold sessions concurrently within a scenario.
Registry: Vec<Box<dyn EngineFactory>> in proef-cli, engines optionally behind cargo
features (one feature per engine). Lifecycle is enforced by ownership shape (only a session runs
batches), not typestate generics (which would break dyn).
Consequences
Adding an engine = one crate + one registry line; pack schema and doctor extend via
step_kinds()/doctor() without core edits — the acceptance test: a future non-hurl
engine lands with zero proef-core diff. Core stays free of engine-specific types. Costs accepted:
Box<dyn> indirection (irrelevant at batch granularity); two traits instead of one
(justified: lifecycle safety + capability discovery). Engines own their artifacts
(hurl files / screenshots / HAR).
Alternatives considered
Single Engine trait with runtime lifecycle state (v3 draft) — weaker lifecycle
guarantees; enum dispatch — inverts the dependency direction; dynamic loading (dlopen/
WASM) — rejected as over-architecture, compiled-in covers every stated future; typestate
generics — fights dyn, ownership shape gives most of the safety.
Errata
2026-07-28 (M1/M5): The routing example above names the API step kind
http:; ADR-0004’s examples and TECH-SPEC §6’s normative pack schema use
hurl: (the raw-block key doubles as the routing kind). The implementation
follows the tech spec: the step kind and the engine id are hurl, so
“a step’s kind names its engine” holds verbatim. Read http: in the Decision
above as hurl:. Other kind prefixes remain reserved for a future non-hurl engine as written.
Amendment — the core’s entry grammar is a named closed set
2026-09-01 · Accepted. “Core stays free of engine-specific types” is true and stays true. “Core stays free of engine-specific syntax” was never true, and the worklist carried the gap for two rounds without resolving it. This amendment states the real boundary and makes it enforceable.
Why the core knows any hurl at all
The core performs text surgery on entries: bake_entry_options splices an
[Options] block into each entry after its header block, and an expect: macro
merges asserts into the previous request entry (ADR-0004). Both operations have
to find an entry boundary in text the engine will later parse. That is structural,
not incidental — the surgery is what the pack format is built on — so a boundary
recogniser has to live somewhere, and pushing it behind the seam would move the
literals without making the algorithm engine-independent.
The measurement
Not the “~290 lines, all in lower.rs” the worklist recorded — that figure
counted #[cfg(test)] fixtures, where a core test exercising the pipeline
necessarily writes some engine’s payload. The vocabulary is thirteen distinct
literals across four files:
| Group | Tokens | Where |
|---|---|---|
| written — the core generates this hurl | [Options], [Asserts], HTTP *, variable:, retry:, retry-interval:, delay: | lower.rs |
| recognised — read to find an entry boundary | ``` (body fence), HTTP / HTTP / HTTP/ | lower.rs, emit.rs, pack/validate.rs |
| quoted — a hurl snippet shown to an author | GET ${url:base}/PATH, HTTP 200 | bind.rs |
The four boundary recognisers (is_method_line, is_section_header,
is_response_line, is_header_line) are already one canonical pub(crate) set
shared by three of those files. That half is done.
The third group is the one this measurement nearly missed, and it is worth
naming why. bind.rs renders a did-you-mean help string for an author whose
sentence bound no macro, and that string contains a small hurl example. It
generates nothing and parses nothing, but it is engine syntax living in the
core, and it drifts like any other copy. The guard’s first version could not see
it — the literal spans lines, and a per-line scan discards a run that never
closes — while this amendment claimed the set was closed. Multi-line literals
are where a larger piece of engine syntax would naturally be written, so the
blind spot sat exactly where the risk is highest. The guard now lexes whole
files.
The same row cost a second correction (2026-09-02). Lexing whole files
surfaced HTTP 200; the GET ${url:base}/PATH line directly above it in the
same literal stayed invisible for another round, because the guard classified
four shapes — fence, response line, section header, option line — and a
method line was not among them, though this section names it as one of the
four recognisers. A guard is closed only over the shapes it can classify, so
the two claims have to be checked against each other rather than assumed to
agree. The classifier now knows method lines, which is what added the row
above. In the same pass the scan stopped truncating at a file’s first
#[cfg(test)] mod and began excising every test module instead: production
code placed after one was silently unscanned, and html.rs and
pack/validate.rs already carry a second test module.
proef’s own pack keys (macros:, match:, secret:, steps:, use:) are
shaped like option lines and are excluded by name rather than listed as
sanctioned rows: an inventory that is a third exceptions stops reading as a
closed set.
The asymmetry this exposes
StepKindSpec::options exists, in its own words, as “the seam that keeps option
spellings out of proef-core” — added because matching "retry-interval:" as a
literal meant “one rule lived at two altitudes.” It covers recognising options.
The core still writes retry:, retry-interval:, delay: and variable: as
literals, so the same rule still lives at two altitudes, in the other direction.
Correction (2026-09-10) — the set was fourteen, and the fourteenth was unclassifiable
The measurement above says thirteen literals across four files. It was
fourteen. The one it missed is "file," in emit.rs, where file_refs_in
found the assets an artifact reads by scanning for that literal and a closing
; — hurl’s body grammar, in proef-core, for the entire life of asset
staging.
It went unrecorded for the same reason the method line did, and this section
had already written the rule that predicts it: a guard is closed only over the
shapes it can classify. engine_grammar_kind knew fences, HTTP, [Section]
headers, method lines and key: value options. A body constructor is none of
those — no colon, no brackets, no uppercase — so the literal was never
classified, never entered the inventory, and was never reported missing from
it. The set was not measured and found closed; it was measured through a
classifier that could not see this member.
That is the third decay of this section’s own claim: once by an order of magnitude in the count, once by a multi-line literal, and now by a shape. Each time the count was wrong in the direction of the guard’s blind spot, which is the only direction it can be wrong in.
Resolved by the second remedy, not the first. Decision 2 below sends the
author to one of two options: widen the sanctioned set on the record, or put
the syntax behind the seam. This is the first time the second was taken. The
scan is now StepKindSpec::assets, a fourth engine-contributed hook beside
validate, fragments and options — so the recognised group loses its
body-reference member entirely rather than gaining a sanctioned row.
Moving it also fixed the reading. A text scan cannot tell a real file,…;
body from the same six characters inside a JSON or assertion body; the engine
reads its own AST and can. It also has to avoid hurl’s shared
visit_filename hook, which carries the [Options] file paths (output,
cacert, client-cert, client-key, netrc-file, unix-socket) alongside
real bodies — output: names a file the run writes, and staging it would
demand a source that cannot exist. Only the two body positions are read. None
of that distinction is expressible in core, which is the argument for the seam
stated as a capability rather than as a rule.
The classifier gained a body arm in the same change, so the blind spot is
closed independently of the literal that exposed it: a bare lowercase keyword
followed by a comma (file,, hex,, base64,) is now classified, and
reintroducing one into core fails the guard with body "file," in emit.rs.
Decision
- The set above is the sanctioned core entry grammar. It is closed: a token outside it, or an existing token appearing in another core module, is a defect against this ADR.
- It is pinned by
crates/proef-cli/tests/source_guards.rs(hurl_grammar_in_core_is_the_closed_set_the_adr_names), which lexes every production literal inproef-coreand fails on growth, on relocation to another core module, and on shrinkage — then sends the author back here. A claim of this shape decays the moment it is only prose. This one already had, twice: once by an order of magnitude in the count, and once in this very section, which asserted a closed set while the guard behind it could not read a multi-line literal. - Migrating the written group behind the seam (an emitter beside
StepKindSpec::options) is deferred, not rejected. It buys nothing today: hurl is the only engine and no other is scheduled, so the migration would add a fn pointer, a trait obligation and a public-API break to relocate seven literals that exactly one implementation will ever supply. Trigger: a second engine being scheduled. That is also when ADR-0002’s acceptance test — a new engine lands with zeroproef-corediff — first has anything to say about them; until then it is unfalsifiable here either way.
Consequences
The acceptance test is narrowed on the record: a second engine lands with zero
proef-core diff except the written group, which is a known, enumerated,
guarded debt with a named trigger rather than an open question. Anyone reaching
for new hurl syntax in the core hits a failing test that names both remedies.
ADR-0003 — Upstream tracking: exact pins, thin zero-diff fork, upgrade canary
Status: Accepted · Date: 2026-07-28
Context
Requirement (stated): “use a hurl fork with as few changes as possible — I want to keep
using hurl as it improves over time; a wrapper/Gherkin adapter over hurl.” Verified:
hurl ships breaking crate-API changes in minor releases (#3846; maintainers advise
cargo install --locked); the release cadence is multiple per year; run_entries is
#[doc(hidden)]. One concrete patch need is already identified (ADR-0010 / TECH-SPEC
§5): run_entries creates its HTTP client internally, so accepting &mut Client would
erase per-segment connection costs — a verified two-call-site change.
Decision
Three-tier policy, in order of preference. (1) Steady state: depend on published
crates with exact pins (hurl = "=8.0.1", hurl_core = "=8.0.1"), build --locked;
the GitHub fork exists but carries zero diff. (2) Patch vehicle: when a release
breaks the seam or a small change is needed, carry a minimal-diff branch on the fork,
consumed via Cargo [patch."crates-io"] (or a git-tag dep), rebased onto each upstream
release. (3) Upstream everything: every patch is PR’d upstream so the fork’s diff
trends back to zero. An upgrade-canary CI job (weekly + on upstream release) builds
against the next hurl version and replays the full suite; pins move only after it is
green, via the runbook in IMPLEMENTATION-PLAN §7.
Consequences
“Keep using hurl as it improves” becomes a scheduled chore, not a gamble; the wrapper never becomes a divergent fork; breakage is discovered pre-pin-bump. Costs: upgrade PRs are deliberate work per release; the fork must be rebased when (and only when) it carries a patch; MSRV follows upstream (hurl master already 1.97.1 — neutralized by the project’s always-latest-stable toolchain rule).
Alternatives considered
Track master via git dependency — unvetted breakage flows in continuously. Vendor the
hurl source into the repo — a divergent fork in disguise; loses provenance and cadence.
Caret/tilde version ranges — semver is demonstrably not honored for the library surface;
exact pins are the only safe mode.
ADR-0004 — Pack format: YAML skeleton + embedded raw Hurl blocks
Status: Accepted · Date: 2026-07-28
Context
Requirement (stated): macro packs must be human-readable — “is there a better alternative than macro YAML?” Analysis (architecture review §8) evaluated YAML+schema, KDL, TOML, Pkl/CUE/Dhall, RON/JSON5, Rhai/Lua scripting, a custom DSL, and Karate-style Gherkin-native macros against readability/writability, comments, multiline bodies, templating interplay, editor tooling, serde support, and team familiarity (existing packs are YAML with schemars-driven autocomplete). Key insight: the unreadable part of packs was never YAML itself — it was HTTP-as-YAML-trees, while hurl’s own plaintext format is the human-readable HTTP DSL, and the backend team already reads/writes it fluently.
Decision
Packs stay YAML (serde_norway; schemars JSON Schema; comments; block scalars) but only
as the thin binding skeleton: macro name, match: pattern, params, defaults,
tags, description, composition (use:/with:), step modifiers (optional:,
when:, retry:, saveAs:). The HTTP payload of a hurl step is a raw Hurl block:
steps:
- name: Resolve the record name to its id
hurl: |
GET ${url:base}/api/v1/admin/search/records
Authorization: Bearer ${secret:apiToken}
[Query]
q: ${name}
HTTP 200
[Captures]
recordId: jsonpath "$[0].id"
Blocks are validated at pack load by parse_hurl_file after ${…} lowering — real hurl
syntax errors with real spans. Structured step trees are reserved for a future non-hurl
engine, which would have no native text DSL. Assert-only macros use
expect: (merged into the previous request entry — the Then-step rule).
Consequences
Pack bodies are literally hurl: copy-paste flows both ways with the backend corpus; no
bespoke assert/capture schema to maintain for the API engine; the emitter for hurl steps
approaches the identity function. Costs: autocomplete inside the block is plain-text
(mitigated: load-time parse errors are immediate; editors have hurl highlighting; a
proef fmt pass can normalize blocks); one lowering pass must run before parse (already
required for ${…}).
Alternatives considered
KDL — pleasant syntax but no schema/LSP story comparable to YAML, zero team familiarity;
recorded as the fallback if YAML friction materializes. TOML — wrong shape for nested
step lists (kept for proef.toml config). Pkl/CUE/Dhall — second language + toolchain,
over-architecture at this size. Rhai/Lua — packs become programs; kills static
validation and --dry-run guarantees. Custom DSL — a parser/LSP/formatter to own
forever. Karate-style callable feature files — collapses the macro/test distinction;
a typed params/defaults/validation model is strictly stronger.
Amendment (2026-07-30): the top-level key is macros:
The pack root key was renamed templates: → macros: to end a three-way naming
split (the YAML key said templates, the docs and internal model said macro, the
file/dir said pack). The entry is now uniformly a macro; a pack is a file of
macros. Pure rename — format, schema, and semantics are unchanged (error-corpus
snapshots regenerated, diff verified as templates:→macros: only). No templates:
alias is kept — one canonical spelling (golden rule: one way to do one thing).
ADR-0005 — Two-tier variables (${…} / {{…}}), the World, and secrets
Status: Accepted · Date: 2026-07-28
Context
The author-time variable system (${var}, ${env:NAME:-default}, ${run:id}, ${global:key},
${fake:*} seeded from the run id, $${} escape, recursive expansion) is proven with
test authors. Hurl has its own runtime templating ({{name}}) fed by captures. The spike
validated running both tiers side by side — including the bug it surfaced: step-captured
args can themselves contain ${…} and need recursive (depth-capped) resolution.
Decision
Two explicit tiers. ${…} is author time: params, env (with defaults), run id,
World reads (${global:key}), fake data, secrets references — resolved during lowering,
recursively with a depth cap of 8, and baked into artifacts (except secrets). {{…}}
is run time: hurl-native templates for captures, left verbatim in artifacts so the
embedded engine and the stock CLI resolve them identically. World: one typed variable
scope per scenario plus a persistent global store (.proef-state.json, atomic
temp+rename). Engine
bridging: seed hurl’s VariableSet from the World before each batch; merge
HurlResult.variables back after; saveAs: global promotes a capture into the
persistent store. Secrets: ${secret:NAME} resolves from the PROEF_SECRET_<NAME> environment
override, else the encrypted store .proef-secrets.json (chacha20poly1305 + rpassword); values are injected via
VariableSet::insert_secret (hurl redacts them in logs/reports); artifacts carry
{{secret_name}} placeholders, never values; our reporters additionally redact by value
(property-tested invariant).
Amendment (2026-08-16): “redact by value” includes each secret’s common
encoded forms — base64 (both alphabets, with/without padding), hex (both
cases), RFC 3986 percent-encoding, and the JSON-string escape — derived inside
Redactions::new so every sink is covered by construction. Demonstrated live
before the amendment: a server reflecting a bearer token base64-encoded put a
trivially-decodable string into an assert-failure detail, the raw needle never
fired, and the encoded credential reached the console and events.jsonl. The
needle set covers the reversible transforms that occur at HTTP boundaries; a
secret reflected hashed or re-encrypted matches no needle list, and the ADR
does not claim otherwise. Over-redaction is the accepted failure direction.
Consequences
Artifacts are runnable by both toolchains with identical meaning; authors keep a familiar
mental model unchanged; secrets are structurally absent from every persisted output.
Cost: two syntaxes coexist in packs — mitigated by the strict rule of thumb (“$ =
before the run, {{ = during the run”) documented in the pack authoring guide.
Alternatives considered
Single-tier (resolve everything at author time) — breaks capture chaining and makes
artifacts non-parametric. Single-tier (everything hurl {{}}) — loses env defaults,
fakes, and World reads. String-only World — kept
typed here (hurl Value model) because captures cross engines; stringly-typed
round-trips would lose numbers/bools at engine boundaries.
Errata
2026-09-07 (0.18): the invariant’s reach was found unenforced on one of
its two paths. The event stream masks through Redactions::apply_event, an
exhaustive destructure; the CI sinks that render from RunSummary (JUnit,
CTRF, TAP, timings.json, the GitHub summary and annotations) masked failure
detail, but five of them bypassed the masker for the identity fields
(scenario, file, tags, the skip reason) — no live leak, since secrets
lower to {{name}} and the engine pre-redacts details, but a boundary held by
convention. Closed per sink (#171), then made structural (#178):
Redactions::apply_outcome destructures ScenarioOutcome/StepOutcome
without .., so a new text field fails to compile until it is masked, and each
sink redacts one outcome after matching @quarantine on the raw identity.
Per-sink rather than a wholesale apply_summary, because RunSummary also
feeds exit_code_excluding, whose quarantine matching needs the unredacted
identity.
2026-07-29 (v0.3.1): the “secrets reach no sink” invariant now explicitly
covers the persistent World: a saveAs: global capture whose value equals a
known secret is refused (the owning step warns) — .proef-state.json is
plaintext at rest and must never receive secret-derived material. PROEF_KEY
(base64) may supply the project key via the environment for CI use of a
committed ciphertext store.
2026-07-28 (post-M5 hardening; API removed 2026-07-29 per YAGNI):
“snapshot/restore across scenario retries” originally described a mechanism whose
trigger was never specified anywhere in the corpus — no
CLI flag, tag, or pack directive schedules a scenario-level retry (US-5’s step-level
retry: is implemented and is the flake tool in practice). Decision: scenario-level
retries are deferred indefinitely, and the unused GlobalStore::snapshot/restore
API has been removed (YAGNI: no dead promised-behavior code). Whoever implements
scenario retries specifies the trigger surface and its mechanism in a superseding ADR. Also note: the scenario merge-back is write-set-only (World
tracks its saveAs promotions) — merging a whole snapshot back would lose concurrent
scenarios’ updates.
ADR-0006 — Engine traits are sync + dyn; no async machinery in v1
Status: Accepted · Date: 2026-07-28
Context
All planned engines are blocking at the edge: hurl is synchronous libcurl, and a future
non-hurl engine’s driver can be driven blocking — even one that also offers an
async API. Verified ecosystem facts (mid-2026, stable 1.97): async fn in traits is stable
for static dispatch but still not dyn-compatible (AFIDT is nightly-only, no
timeline; RTN unstable); maybe-async’s sync/async toggle is a non-additive cargo
feature with documented ecosystem breakage. Calls across the seam are coarse
(batch-level), so async buys no throughput inside the runner itself.
Decision
EngineFactory and EngineSession are synchronous traits, used as Box<dyn …>.
Parallelism is scenario-per-OS-thread. If a future host (server mode, desktop UI) is
async, it wraps engine calls in spawn_blocking at its edge; the core never learns
about executors (reinforced by the sans-IO-lite rule: core does no IO at all).
maybe-async is explicitly banned. Traits meant to be implemented externally
(Engine*) are not sealed; internal traits that must stay evolvable are sealed.
Consequences
Simple, dyn-compatible seam today; no async runtime in the dependency tree
(tokio-util’s CancellationToken is runtime-independent — ADR-0007); a future async
migration is additive (an AsyncEngineSession adapter or AFIDT adoption when stable),
not a rewrite. Cost: a long-running async host pays one thread per in-flight scenario —
acceptable at e2e-suite scale.
Alternatives considered
Async-first trait via async-trait — boxes every call, forces an executor decision on
every consumer, and models nothing real while engines block. maybe-async dual API —
non-additive feature hazard. Callback/actor-per-engine threading model — more machinery
than batch dispatch needs; revisit only if an engine genuinely multiplexes (e.g. a driver
with concurrent event streams), and then inside that engine crate, invisible to the seam.
ADR-0007 — Cancellation: cooperative at batch boundaries, with budgets
Status: Accepted · Date: 2026-07-28
Context
Source-verified: hurl has no cancellation mechanism anywhere — no signal handling in
the workspace, no abort check in the entry loop; delay/retry-interval are
uninterruptible thread::sleeps; retry/repeat accept Count::Infinite; the only
bounds are libcurl’s per-request timeouts (default 300 s), which do not bound total
entry time under retries. The standard here is structural cancellation
(CancellationToken threaded everywhere). Verified: tokio_util::sync::CancellationToken
is runtime-agnostic (tokio sync primitives are documented runtime-independent;
default-features = false — no tokio runtime enters the tree); is_cancelled() polling
and child_token() work from plain threads.
Decision
Cancellation is cooperative at batch boundaries: the orchestrator checks a per-run
CancellationToken (child token per scenario) before opening sessions and before each
batch; EngineSession::run_batch receives the token so engines may honor it at finer
grain when they can (a future non-hurl engine might; engine-hurl cannot mid-run_entries).
Stuck-batch policy, layered: (1) the pack lint rejects infinite retries (retry:
must carry a finite count) and unbounded repeat; (2) engine-hurl clamps per-request
timeouts and computes a batch budget = Σ(entry timeout × (retries+1)) + retry
intervals + margin; (3) a watchdog marks a scenario thread abandoned when its budget
expires — the runner records a System failure with full context and detaches the
thread (process exit reaps it) rather than blocking the run on an unjoinable thread.
Ctrl-C: first signal cancels the token (graceful: finish current batches, run
teardowns, write reports); second signal hard-exits.
Consequences
Bounded, explainable runs; no dependency on hurl gaining cancellation; a clean seam for engines that can do better. Costs: a cancel can wait out one in-flight batch (bounded by its budget); abandoned threads leak until process exit (accepted: the process is short-lived by design). If interrupt support ever lands upstream, adopting it is an engine-internal change (candidate for the ADR-0003 patch pipeline).
One consequence reaches a later feature, so it is recorded here rather than
discovered again. [run] exclusive-tags promises a matching scenario the pool
to itself, and the dispatcher enforces that against its own active set. An
abandoned scenario leaves that set the moment the watchdog fires, while its
detached thread keeps issuing requests until the entry in flight returns — the
whole reason abandonment exists. So an exclusive scenario can start while an
abandoned neighbour is still talking to the target: exactly the interference the
key promises away, in the one window this ADR knowingly leaves open. It takes a
budget blowout in the same moment, and a run that never trips the watchdog never
meets it, so this is documented rather than engineered away — closing it means a
bounded grace on the exclusive fill gate while detached threads report alive,
which buys a rare guarantee with a per-exclusive-scenario delay every run pays.
Revisit if the isolation guarantee ever has to be absolute.
Amendment (2026-09-07) — the budget family is closed over its inputs, and bounded as a product
The 0.18 survey found three holes in the value-cap regime, one of them the exact shape this ADR exists to prevent:
max-time:was read by the budget calculator and invisible to the lint.entry_timeouthas always taken a literalmax-time:as the entry’s timeout, while the option recogniser did not know the key — so[Options] max-time: 100000hwas lint-clean and produced a multi-year batch budget the watchdog dutifully honoured. It now carries the duration cap like every budget input, and a test pins the rule the hole broke: every option the budget reads must be one the lint can see (every_budget_input_carries_a_value_rule).retry-interval:multiplied into the budget with no value cap. It was in the recogniser for double-declaration purposes only; it now carries the duration cap.- Individually capped values compose into an unbounded product.
retry: 10_000(at the count cap) times a 30 s timeout is ~83 hours, lint-clean; saturated arithmetic reachesDuration::MAX, whoseInstant + budgetaddition panics — contained by the dispatcher’scatch_unwind, but reported as a phantom “scenario thread panicked” system fault from a user-authored value. Two closures: the computed batch budget clamps to an absolute ceiling of four hours (MAX_BATCH_BUDGET— generous for any batch of API calls with finite retries, and a truthful watchdog abandonment for a runaway product), and the dispatcher’s deadline arithmetic useschecked_addwith a far-future fallback, so an engine that ever hands core an unclamped budget degrades to “no deadline” rather than a panic.
In the same family: [http] timeout-ms = 0 was accepted and means no
timeout to libcurl — the exact unbounded hang the default defends against,
opted into by a value that reads like “immediately”. Refused as a user error
now.
Alternatives considered
Killing scenario threads — unsound in Rust (no safe thread kill). Running each batch in a subprocess for killability — reintroduces the subprocess architecture ADR-0001 rejected, per-batch. Relying on timeouts alone — unbounded under retry loops (verified), and no graceful-report path on Ctrl-C.
ADR-0008 — Serde event spine, decorator reporters, libtest-mimic harness
Status: Accepted · Date: 2026-07-28
Context
a live-event seam (EventSink(Arc<dyn Fn(RunEvent)>), domain events decoupled
from wire format) is proven; its run record is a separate reporter path. The two best
in-domain designs both converge on “typed event stream + composable consumers”:
cargo-nextest (runner emits structured events; reporter fans out to human/JUnit/machine
outputs) and cucumber-rs (decorator Writer stack: Normalize → Summarize → leaves,
with Tee, marker traits for ordering guarantees). Verified: nextest officially
supports libtest-mimic custom harnesses (documented CLI contract; libtest-mimic 0.8.x);
libtest’s JSON format itself is still unstable nightly territory; OTel test semconv is
“development” maturity.
Decision
One serde-able event enum in proef-core is the spine: RunStarted,
ScenarioStarted, BatchStarted, StepFinished (engine id, StepRef
feature/line, status, attempts, duration, capture names), ScenarioFinished,
RunFinished. The JSONL run record is the appended event stream — explain,
history, and any future UI replay from disk; no second record format. Reporters are
event consumers composed decorator-style: Normalize (repairs interleaving from
parallel scenarios) → Summarize → leaves: console BDD tree, JUnit XML (via
quick-junit), GitHub job summary, JSONL appender. Secret values never enter events
(capture names only — redaction invariant, ADR-0005). M5: a libtest-mimic harness
binary exposes one Trial per scenario, making cargo nextest run and IDE test UIs
drive proef with zero custom protocol work; libtest JSON remains an output adapter,
never the native schema. OTel export: deferred; if added, a thin optional reporter
mapping to test.* semconv names.
Consequences
Single source of truth for live progress and persistence; reporters are ~a page each;
new outputs are additive leaves. Replayability makes run records diffable and testable
(insta snapshots over event streams). Cost: event schema becomes a compatibility
surface — versioned with a schema field from day one.
Alternatives considered
a split design (live events + separate JSON record) — two sources of truth to keep
consistent. tracing as the event bus — wrong tool: tracing is operator telemetry, not
a typed result stream (kept for diagnostics). Cucumber’s writers verbatim — async trait
- World coupling we don’t need; the decorator shape is what’s adopted.
Errata
2026-07-29: the variant set has grown additively since acceptance:
EntryRunning (live per-attempt engine progress) joined the six original
variants, RunFinished gained a cancelled flag, and StepFinished gained a
detail failure field — all serialized only when present, so pre-existing
streams parse unchanged. The additive-only rule held; this note keeps the
variant inventory honest.
2026-08-24 (RF wave 2): scenario_finished gained reason (why a
scenario is skipped — authored spellings start with @, mechanical prose
never does; ADR-0019) and tags (the accumulated tag set, finished-event
only because the cancel-skip path emits no start); scenario_started gained
exclusive (the scheduler’s own bool, R11-6); step_finished gained
reproduce_hint (the failing request’s redacted curl). All
serialized only when present; EVENT_SCHEMA_VERSION stays 1.
run_started further gained env/metadata/shuffled (ADR-0020) and
rerun_of (the E2 rerun overlay) — same additive discipline.
2026-09-06 (0.18): the record is now protected the way the console
already was. events.jsonl was handed a bare File, so a disk filling
mid-run truncated the record while the run exited by its verdict; the record’s
writer now latches its first failure and the exit funnel turns it into a system
error (exit 3) through the same fold as the JUnit/CTRF and GitHub-summary write
failures (escalate_environment_failures). A single SIGTERM/SIGHUP is a
cancellation — the record closes normally with run_finished + cancelled
— so only a second signal, a SIGKILL, or a crash leaves a truncated record
(EVENTS.md). The sidecars that sit beside the record (timings.json,
inputs.json) are derived aids, never a second record: the stream stays the
only persisted format, and EVENT_SCHEMA_VERSION stays 1.
ADR-0009 — Error taxonomy by fault, stable exit codes, miette at the edge
Status: Accepted · Date: 2026-07-28
Context
The error model categorizes by who is at fault — User /
TestFailure / System(anyhow) — with a total mapping to its stable exit-code scheme
(0 ok · 1 test failure · 2 user error · 3 system error), integration-tested via
assert_cmd. 2026 ecosystem consensus (blessed.rs et al.): thiserror in libraries, anyhow
at the application edge; miette’s Diagnostic adds codes/help/labeled source spans for
user-facing errors. Survey of backend traits (tower/sqlx/rustls/probe-rs): behind dyn,
a unified error enum with boxed sources beats associated type Error. gherkin 0.16
spans are byte offsets (verified), directly convertible to miette SourceSpan (with an
EOF-trailing-newline clamp; LineCol.column is char-counted — never mixed into byte
math). snafu/error-stack: adopted by some large codebases, unnecessary at this crate
count.
Decision
A fault-category model, extended for engines: proef-core defines
#![allow(unused)]
fn main() {
pub enum CoreError { User(..), TestFailure(..), System(..) } // → exit 2 / 1 / 3
pub struct EngineError { pub class: EngineErrorClass, // Infra | AssertFailed | Setup
pub message: String,
pub source: Option<Box<dyn Error + Send + Sync>> }
}
AssertFailed folds into TestFailure; Infra/Setup into System. thiserror 2
everywhere; no anyhow in library crates (only inside System’s boxed source at the
edge). miette lives only in proef-cli: parse/bind/validation errors wrap into
Diagnostics with labeled spans into .feature files (gherkin byte spans) and pack
YAML (serde_norway locations); engine failures render the feature line + artifact span
from the sidecar. Exit codes are a typed enum, pinned by CLI integration tests.
Consequences
Every failure has a fault category, a stable exit code, and a source-located rendering;
engine crates stay miette-free (usable headless); the explain command reuses the same
classification. Cost: two error layers (core vs engine) — justified by the seam:
engines can’t know exit codes, core can’t know engine internals.
Alternatives considered
Associated type Error on engine traits — erased behind dyn anyway. anyhow
everywhere — loses matchable categories that exit codes require. snafu — per-crate
context ergonomics we don’t yet need; revisit if the workspace grows past ~10 crates.
Amendment — 2026-08-04 (exit code 130 documented, not a variant)
test and watch hard-exit with code 130 (128+SIGINT) on a second Ctrl-C
while a run is already cancelling (ADR-0007) — the shell’s own convention for a
signal-terminated process. This is a sanctioned OS-signal escape hatch, not a
graceful outcome the fault-category model classifies, so it is intentionally
not an ExitCode variant; ExitCode stays the total 0/1/2/3 mapping above.
Amendment — 2026-09-06 (every signal, and undelivered output)
Ctrl-C, SIGTERM and SIGHUP all take the graceful cancel (ctrlc’s
termination feature), so a CI job timeout or docker stop is a cancelled
run — exit 1 with a complete record — not a kill; the 130 hard exit above
fires on a second signal of any of the three, and the handler carries no
signal identity, so one code covers them all. And output proef could not
deliver never looks like success: a failed write of the run record, of a
JUnit or CTRF file, or of the GitHub step summary re-classifies the exit to
System (3) through one fold, escalate_environment_failures, beside the
stdout latch this taxonomy already covered.
ADR-0010 — Artifacts as contract: emitted .hurl is the executed input
Status: Accepted · Date: 2026-07-28
Context
The backend team’s hurl corpus makes .hurl the interop format. The spike proved
generated artifacts run identically under the prototype engine and the stock CLI — and
that keeping two implementations honest requires differential testing. Embedding hurl
(ADR-0001) enables something stronger: the artifact and the executed input can be the
same bytes. Verified seam facts that shape the mechanics: run_entries creates its
HTTP client per call (fresh connections; cookie jar seedable only via Netscape-format
file; variables chain losslessly via HurlResult.variables); per-entry [Options]
override batch-level RunnerOptions defaults (clone-then-override, verified).
Decision
For every scenario, the emitter produces canonical .hurl text; that exact text is
what parse_hurl_file + run_entries execute — drift between artifact and execution is
structurally impossible. Alongside each artifact: a sidecar map (<slug>.map.json:
entry ↔ feature file/line/step text, optional flags, capture names, batch boundaries)
and, when World/global values are referenced, a generated <slug>.vars file so the
backend team can replay with hurl --variables-file. optional: entries carry an
# optional marker comment (no hurl equivalent — the runner segments around them).
Execution batches maximally: one run_entries call per scenario unless optional:
boundaries or interleaved other-engine steps force a split; when a split occurs,
variables chain via HurlResult.variables and cookies (if used) round-trip via a
Netscape temp file behind a SessionState struct. Queued upstream patch #1 (per
ADR-0003): run_entries accepting &mut http::Client — verified two-call-site change —
which erases per-segment connection/cookie costs entirely. Artifacts never contain
secret values (ADR-0005). Artifact output: per-run under .proef-runs/<id>/artifacts/
plus proef artifacts for a stable CI hand-off directory.
Consequences
Debugging = opening the artifact with tools the team already knows; the backend corpus
and proef packs stay mutually copy-paste-able; --dry-run validation includes parsing
the real artifact with the real parser. Cost: canonical formatting is a compatibility
surface (snapshot-tested); segmented scenarios pay reconnection costs until patch #1
lands upstream.
Alternatives considered
Internal-only IR execution with optional export — loses the same-bytes guarantee and
demotes artifacts to lossy exports. Differential-oracle architecture (v1 plan) —
superseded by ADR-0001; preserved in research/ as the static-distribution fallback.
ADR-0011 — Fixture server is synchronous tiny_http, not axum
Status: Accepted · Date: 2026-07-28 (decision made at M3; recorded as an ADR during the post-M5 hardening pass — it had been noted only as inline errata in TECH-SPEC §14 and TESTING-STRATEGY §2)
Context
TESTING-STRATEGY as originally written named axum for the integration fixture
server (proef-fixture). Axum requires a tokio runtime. The workspace bans the tokio
runtime outright (ADR-0006/ADR-0007: engines and core are sync; only tokio-util with
default-features = false enters the tree, for CancellationToken). A dev-dependency
would not leak into shipped binaries, but it would put a full multithreaded async
runtime into every cargo nextest process, contradict the “no async machinery” line we
enforce with cargo deny bans, and normalize exactly the dependency the ADRs exclude.
The fixture’s needs are modest: a handful of JSON endpoints, per-env state isolation, deterministic token-driven delayed visibility, cookies, one deliberately slow route and one deliberately malformed one — all exercised by at most a dozen concurrent scenarios.
Decision
proef-fixture is built on tiny_http (synchronous, dependency-light) with a
plain std::thread accept loop and an Arc<Mutex<_>> state map keyed by the
X-Proef-Env header. It starts on an ephemeral port (Server::http("127.0.0.1:0")),
reports its base URL, and shuts down via an AtomicBool + recv_timeout poll.
Portability note: address introspection uses ListenAddr::to_ip() — the Unix
variant of ListenAddr exists only on unix targets, so matching on it breaks the
Windows build.
Consequences
- The workspace stays runtime-free end to end; the deny-list ban on tokio stands without a dev-dependency exception.
- The fixture is one file, debuggable with a thread dump, and starts in microseconds — each integration test spawns its own isolated instance.
- No HTTP/2, no TLS, no streaming in the fixture — acceptable: the engine’s HTTP behavior is hurl/libcurl’s concern, not the fixture’s; the fixture only scripts responses.
Alternatives considered
- axum (as originally specced): rejected — drags in the banned tokio runtime; every capability the fixture needs is available synchronously.
- hyper in blocking mode / raw
std::net: more code for no additional fidelity. - Out-of-process fixture binary: slower startup, port coordination, and a
lifetime-management problem tests would have to solve; in-process
tiny_httpgives free isolation per test.
Amendment — 2026-07-31 (dev-loop CLI binds the advertised default port)
Fixture::start() stays ephemeral (127.0.0.1:0) — the integration suite spawns a
dozen concurrent instances and each needs its own port. But the shipped proef.toml
advertises base = http://127.0.0.1:8787, so a first-time cargo run -p xtask -- fixture on a random port left the default unreachable and forced a PROEF_BASE_URL
export. The dev-loop CLI (xtask fixture, one instance at a time) now calls the
new Fixture::start_on(port) to bind 8787 by default (override: ... -- fixture <port>), falling back to an ephemeral port — with the PROEF_BASE_URL line printed —
only when 8787 is busy. The library API and the per-test isolation described above are
unchanged; only the human entry point picks a stable, documented port.
ADR-0012 — Project configuration & environments in proef.toml
Status: Accepted · Date: 2026-07-30
Context
Test files must stay pure prose — no URLs, no environment data, no variable
definitions (operator requirement: “test files for testing, not for variable
definitions”). Before this, the only non-secret variable source was the in-feature
# baseURL: directive plus ${env:…}, which put configuration inside the
.feature files. Suites also had to be given an explicit path on every
proef test invocation.
Decision
proef.toml (already the config file — ADR-0004 kept TOML for it) gains
variable-bearing sections and per-environment override profiles, modeled on the
Cloudflare-Wrangler wrangler.toml [env.<name>] pattern (and Cargo [profile.*]):
[url]and[vars]— non-secret variables, referenced in packs as${url:<key>}and${vars:<key>}(new lower-time resolver namespaces, within the ADR-0005${…}tier).[env.<name>.<section>]— per-environment overrides that deep-merge over the base tables, key by key:[env.prod.url]/.varsoverride variables,.http/.runoverride runner settings. Unlisted keys inherit the base.--env <name>/PROEF_ENVselects the active environment.[run] suite— the default suite path, soproef testneeds no path argument; thetests/directory is the zero-config fallback convention.
Core stays sans-IO: the CLI loads proef.toml, validates the --env name,
deep-merges, and injects the resolved scope as LowerCtx::config_vars (keyed
"<namespace>:<key>"). The resolver reads that injected map — it never touches a
file. Secrets remain on their own encrypted channel (${secret:…}) and never
appear in proef.toml.
Consequences
- Feature files become pure prose; URLs / credentials / env data live in one
external file — the single variable-definition mechanism (the legacy
# key:directive was later removed; see the 2026-07-31 amendment). - A referenced-but-undefined
${url:…}/${vars:…}is a user error at lower time (proef::resolve::missing_config_var) — the same strictness as${env:…}. - Deep-merge (not Wrangler’s non-inheritable
vars) means an environment lists only deltas; this is the deliberate divergence from Wrangler’s known footgun. proef test/flows/artifactsaccept an optional path plus a--envflag.
Alternatives considered
- Per-environment files (
proef.<env>.toml, Spring / dotenv style) — cleaner git diffs per env, but more files; the single-file[env.<name>]model keeps one mental model and reuses the existing loader. Recorded as the fallback if one env grows large. - Wrangler-exact non-inheritable vars — rejected: forces re-listing every var per env (the documented Wrangler footgun); deep-merge is the more ergonomic default.
- A
${baseURL}magic bare name — rejected: namespaced${url:base}is collision-free and consistent with${secret:…}/${env:…}(one way to do one thing).
Amendment (2026-07-31) — the # key: directive mechanism is removed
The original decision kept the in-feature # key: value directive (e.g.
# baseURL:) working “for one-off per-file overrides.” That left two ways to
define a variable — a directive inside a .feature file, and [url]/[vars]
in proef.toml — violating one-way-to-do-one-thing (operator: “variables should
be defined in 1 way, not multiple ways”). The directive mechanism is therefore
removed:
FeatureFile::directives,collect_directives,lower::resolve_directives, and theResolveCtx::directivesscope are deleted.${…}plain-name resolution is nowargs > defaultsonly; feature files carry no variable definitions.- Config becomes the single variable source. Config values may themselves embed
${env:NAME:-default}(resolved recursively), so the env-override + default that# baseURL: ${env:PROEF_BASE_URL:-…}provided is preserved asbase = "${env:PROEF_BASE_URL:-…}"under[url]. proef.tomlis now discovered by walking up from the working directory (like cargo/git), so config is found from any subdirectory (needed once the value moved out of the self-contained feature file — e.g. the libtest-mimic harness invokesproeffrom its own crate dir).
# comment lines before Feature: remain valid gherkin comments; they are simply no
longer parsed as directives.
ADR-0013 — Typed macro parameters
Status: Proposed — recommendation: defer (the shape below is recorded for when a real need appears) · Date: 2026-08-02
Context
A macro declares params and binds {capture} values from prose, data-table rows,
defaults, and use:…with:. Every arg is an untyped string today: the matcher
only checks that a capture names a declared param (matcher.rs), never that the value
looks like what the step expects. A mistyped value (the record abc is fetched where a
UUID was meant) is caught only at hurl run time — if at all — not by --dry-run.
Cucumber Expressions solve this with typed parameters ({int}, {uuid}, custom types)
validated before execution. proef wants the same shift-left check, expressed within its
architecture. Round-2 code validation (IMPROVEMENT-PLAN §12, item N1) established three
hard constraints:
- Placement must be the declaration site, not the pattern. Three of the four arg
sources are not captures (data-table rows,
defaults,with:), anduse:-only macros have no pattern at all — so an inline{name:type}form could not type them, and a type buried in thematch:string is invisible toproef schema(schemars). - The two-tier variable rule caps the value. An arg may legitimately be
${…}/{{…}}that only resolves at lower/run time (ADR-0005). Type-checking such an arg at bind time would false-positive, so the check must skip any raw arg containing${or{{— it is a best-effort lint over literal args, not a runtime type system. - One-canonical-way forces a single
paramsspelling.paramsis a YAML sequence today ([q, index]); adding a typed spelling alongside it would be two ways to declare params — the golden-rule violation that killed thetemplates:alias and the# key:directive.
Decision
paramsbecomes a name → type map:params: {q: uuid, index: int, note: any}. A missing/anytype means “declared, unchecked” (today’s behaviour). Parsed with a customDeserialize(not an untagged enum — CLAUDE.md bans those near thearbitrary_precisionfootgun). This is a breaking pack-shape change — every existing pack migratesparams: [q, index]→params: {q: any, index: any}, with no alias (consistent with thetemplates:→macros:and# key:removals).- The type set is a closed registry —
any,int,number,uuid,email,iso8601,word— modelled on thefake::GENERATORSregistry (fake.rs), validated at load so an unknown type is a pack error. No user-defined types in v1 (revisit if a real need appears). - Validation lands at three sites: bound captures + data-table rows at bind time
(
proef::bind::param_type_mismatch), and staticdefaults/with:values at load time (next to the existingdefault_not_param/unknown_with_keychecks). All checks are skipped when the raw arg contains${or{{(constraint 2). Datetime types parse viajiff(sans-IO — literal parse reads no clock). proef schemareflects the typedparamsshape (constraint 1 satisfied).
Consequences
- Breaking: every pack’s
paramsmigrates to the map form. A one-time mechanical change, called out in CHANGELOG and GETTING-STARTED; the error-corpus gains abind__param_type_mismatchcase. - Best-effort, not a guarantee: literal args are checked;
${…}/{{…}}args are not. This must be documented so authors do not read it as a type system — its value is catching typos in literal prose at--dry-run, not enforcing runtime types. --dry-rungains a real new class of caught mistake;proef schemaautocomplete gets richer.- The honest cost/benefit: a breaking migration for a best-effort literal-args lint. Recorded so the trade is deliberate, not incidental.
Best-practice basis & recommendation
Research into the field (2025–2026) is decisive on form and honest about worth:
- Form (if built): the industry norm is declaration-site typing referenced by name,
never a full type spec inline. Cucumber Expressions put only a bareword type-name in the
pattern (indexing a registry); SpecFlow/Reqnroll type by the binding method’s return
type; Bruno (v4, 2024) added declaration-site typed variables with a string default. A
name → typemap is the canonical shape (JSON Schemaproperties, OpenAPI). The inline{name:type}form (behave/pytest-bdd) only works because their “declaration” is code, not a schema — with a schema it loses tooling. So the decision above (single-shape map, closed vocabulary) is the correct form, and the list-or-map union is the worst option on every axis (permanent double surface, weaker autocomplete, bifurcated examples — the exact ambiguity “one canonical way” forbids). ESLint’s flat-config hard break (v9→v10, bounded deprecation then cut) is the precedent for a single-owner format choosing a clean break over an indefinite dual-shape. - Worth (candid): a best-effort lint over literal values is the mypy/TypeScript
gradual-typing bargain — genuinely useful for shallow typos, provided it degrades to
“unchecked,” never “pass,” on the parts it can’t see. But an API test runner’s args are
deferred-heavy (base URLs, captured ids, run-time tokens — all
${…}/{{…}}), the exact population the lint cannot check, which shrinks the realized benefit. proef’s own corpus bears this out: today’s params are overwhelmingly string-ish, so the high-value types (uuid/int/iso8601) would rarely fire. Karate — the nearest-neighbour API tool — never types inputs at all; it types responses.
Recommendation: defer. The clean, best-practice-aligned move for proef right now is not to add a low-current-value lint behind a breaking format change; it is to record the correct shape (above) and adopt it if a pack corpus emerges where structured literal args are common. Building it earlier is defensible only in the disciplined form above — never as a list-or-map union. (If the operator wants it now regardless, implement the single-shape map with the honest-degradation discipline; the migration is mechanical.) (Cucumber Expressions, Reqnroll conversions, Bruno typed variables, OpenAPI 3.1, ESLint 10 removal, Karate schema validation)
Alternatives considered
- Inline
{name:type}(Cucumber-Expression style) — rejected: invisible toproef schema, mis-binds in the tokenizer ({q:uuid}becomes a capture literally namedq:uuid), and covers only 1 of the 4 arg sources. paramsaccepts either a sequence (untyped, non-breaking) or a map (typed), via customDeserialize— the non-breaking alternative. Rejected under one-canonical-way (two shapes for one field), but it is the fallback if a breaking migration is judged too costly for a best-effort lint. (This is the open fork for the operator.)- Keep params untyped (do not do N1) — the status quo; the shift-left gap stays. A legitimate choice given the cost/benefit above.
- User-defined type registry — deferred; a closed set is simpler and covers the common shapes.
ADR-0014 — Suite-level setup & teardown
Status: Accepted · Date: 2026-08-02
Context
Gherkin Background runs steps once per scenario, and a persistent global store
threads saveAs: global captures across scenarios and runs (ADR-0005). What is missing is
a once-per-suite phase: authenticate once and share the token, seed a fixture before
the run and tear it down after. Cucumber solves this with @BeforeAll/@AfterAll hooks;
Karate with callSingle. Round-2 code validation (IMPROVEMENT-PLAN §12, item N3) fixed the
constraints:
- Tags never reach the sans-IO core runner.
ScenarioSpec/ScenarioOutcomecarry no tags;@quarantineis computed at the CLI edge (exec.rs) as anon_gatingset. So a lifecycle phase must be orchestrated at the CLI edge, not inside the runner. - Only
saveAs: globalpromotions cross scenarios, and each pooled scenario snapshots the global store at prepare time — so a setup phase must finish and merge its globals before the parallel pool starts, or early scenarios miss the seeded value. - An assert-failed setup must not be masked. A setup step that fails an assertion
classifies as a test failure (
fault: None), which the worst-wins fold would let through as exit 1 without aborting the pool — every scenario would then run against un-seeded state and cascade confusing failures.
Decision
- Two new
proef.tomlkeys —[run] setupand[run] teardown— each naming a feature file.setupruns (its scenarios, in order) once before the suite pool;teardownonce after. State reaches the suite through the existingsaveAs: globalstore; setup completes and merges before the pool is built. - Orchestrated in
execute()around the parallel pool (CLI edge), never in the sans-IO core. The setup/teardown feature is excluded frombuild_specsso it never also runs as an ordinary scenario, and it is invisible to--tags/--scenario/--rerun. - Failure semantics (exit-code contract, ADR-0009) — validated against Playwright,
Jest, k6, and pytest (see basis below):
- A setup failure of any kind aborts the run before the pool launches and is never masked. It maps to a user (2) or system (3) fault, not a test failure (exit 1) — a broken fixture is not a failing test, the same distinction Playwright draws between a “clear setup error” and a “cryptic test failure.”
- Teardown runs only when setup succeeded (gated on setup-success, not pool-success):
k6 and pytest-
yieldboth skip teardown when setup threw, because tearing down un-created state is itself a fault source. The setup feature is responsible for cleaning its own partial state on failure. - Teardown does run after the pool even when scenarios failed (Playwright/pytest/k6 all do), so cleanup is reliable.
- A teardown failure is loudly reported and yields a distinct non-zero signal (a system/cleanup fault, exit 3) — never a silently-green suite (no mainstream tool masks a teardown failure). It stays distinct from a test failure (exit 1): a cleanup hiccup does not mean the API under test is broken, but it is not hidden.
--dry-runis unaffected: it validates only (never calls the runner), and the setup/teardown features are validated like any other feature but never executed.
Consequences
- A genuine once-per-suite lifecycle, expressed declaratively (Hurl prose in a feature), with no code hooks — the glue-code path proef exists to avoid.
- Exactly one setup and one teardown mechanism — a
[run]construct, not a secondBackgroundconcept and not a tag. The setup feature is authored like any suite feature and reuses the whole pipeline (bind → lower → emit → execute). - Touches the pinned exit-code tests: a failing setup gates the run; new
assert_cmdcases cover the short-circuit and the “setup fault is never masked” invariant. - Product-neutral: the setup steps are ordinary prose bound to macros; nothing about the mechanism assumes a particular backend.
- Auth-once boundary (explicit): the pattern that motivates suite-setup elsewhere —
“authenticate once, share the token” (Playwright
storageState, KaratecallSingle) — is largely already covered in proef by a pre-set secret (proef secret set/PROEF_SECRET_*) resolved per scenario. Sharing a runtime-obtained secret through setup would collide with the invariant thatsaveAs: globalrefuses secret-valued captures (ADR-0005). Promoting a runtime capture into the secret channel is therefore out of scope for this ADR: the global store carries only non-secret setup state (seeded ids, fixture config). Recorded so the boundary is explicit rather than discovered later.
Amendment — 2026-08-10 (cancellation, and what --dry-run validates)
This ADR was specific about a failing setup and a failing teardown, and silent on the
operator interrupting the run. That silence read as considered when it was not, and the
behaviour it left was the opposite of this ADR’s premise: on Ctrl-C the teardown phase ran
with the already-cancelled token, so every teardown scenario resolved Skipped,
phase_failed ignored a phase that only skipped, and cleanup silently never happened.
Decision — cleanup outlives the interrupt. Teardown runs on its own, independent token, never the run’s. On Ctrl-C the pool stops at its batch boundary, the operator is told cleanup is running, and teardown completes — so an interrupted run does not strand whatever setup created.
Note the word: independent, not “child”. A child token cancels when its parent does,
which is precisely the behaviour being fixed; child_token() would have re-implemented
the bug.
ADR-0007’s responsive interrupt is preserved by the escape hatch that already existed: a second Ctrl-C hard-exits (130) out of teardown as out of anything else, and the announcement says so. A hung teardown is bounded by the same batch budgets and watchdog as any other phase, so this needs no timeout of its own.
This is the standard graceful-shutdown shape — first signal begins bounded cleanup, second
forces exit — rather than the test-runner norm, which is worse: Jest does not call
globalTeardown on Ctrl-C (#6029) and Go
does not run t.Cleanup on SIGINT (#41891),
both long-standing complaints rather than settled design.
Corollary — a phase that only skipped is a failure. A skipped phase carries no fault, so the worst-wins fold passed it silently; that is the shape that hid cancelled cleanup. A setup that completes no scenario now aborts the run (the suite would otherwise execute against state setup never created), and a teardown that completes no scenario is reported and fails the run. The setup abort is also what keeps teardown gated on setup-success as this ADR requires: the early return is the gate, so teardown never dismantles what was never built.
--dry-run now does what this ADR already claimed. The Decision above says the phase
features are “validated like any other feature but never executed”. They were validated by
nothing — --dry-run never read the keys — so a broken [run] teardown surfaced only
after a full suite had run, while the identical mistake in [run] setup failed in
milliseconds. Both phases are now validated by one loader shared with execute, which also
pre-flights teardown before the pool, so the same mistake costs the same either way and
a bad path is a user error (2) rather than a blanket system fault (3).
Best-practice basis
Config-key-names-a-file is the dominant model (Playwright globalSetup/globalTeardown,
Jest globalSetup, Vitest) — the [run] table already supplies the “once per run”
qualifier, so [run] setup reads like globalSetup without repeating global. Passing
state out-of-band through a serialized shared store (not live memory) is universal —
Playwright’s storageState file, Jest’s env-var workaround, Karate’s callSingle cache,
k6’s returned data — and proef’s global store is the direct analog. Setup-completes-before-
workers and setup-failure-aborts are unanimous. The refinements above (teardown skipped on
setup failure; teardown failure is non-zero, not silent) come straight from k6/pytest and
the Playwright/pytest “teardown is not silently swallowed” norm.
(Playwright global setup,
Jest config,
k6 lifecycle,
pytest #2508,
Karate callSingle)
Alternatives considered
@setup/@teardowntags — rejected. A tag means “filter/select”; overloading it with “change execution phase” is a second meaning, and tagged setup scenarios would entangle with--tags/--scenario/--rerun/flows/name-dedup with undefined ordering between two@setupscenarios. A[run]construct is single, explicit, and ordered.- Code hooks (
Before/Afterfunctions) — rejected. Arbitrary-code hooks are exactly the imperative glue the declarative-macro model avoids, and they break the sans-IO boundary. The legitimate need (setup/teardown IO) is met with Hurl steps instead. - Per-scenario
Backgroundonly (status quo) — does not cover once-per-suite work; authenticating in every scenario’s Background is wasteful and cannot seed shared state that must exist before the first scenario. - A dedicated
[setup]/[teardown]top-level table — rejected as heavier than needed; these are run-orchestration knobs, so they belong under[run]besidesuite/jobs.
ADR-0015 — Injected observability timestamps (run-level timeline)
Status: Accepted · Date: 2026-08-02
Context
The HTML report has a per-scenario timing waterfall (IMPROVEMENT-PLAN §12, N6a) derived
purely from each step’s duration_ms. It cannot show cross-worker occupancy — which
scenarios ran concurrently, on which of the --jobs workers — because the sans-IO core
reads no clock and the event stream carries no wall-clock timestamp or worker identity.
The core’s purity is deliberate (deterministic snapshots/properties). ADR-0012 established
the escape valve: values the core must not compute (config) are injected at the CLI
edge. run_id is already such a value — the core carries it (RunStarted.run_id) but
the CLI generates it. The same pattern extends to timing.
Decision
- Add two optional, additive fields to the
ScenarioStartedandScenarioFinishedevents:timestamp_ms: Option<u64>(milliseconds since the run began) andworker: Option<u64>(0-based worker index). Both are#[serde(default, skip_serializing_if = "Option::is_none")], so old records parse unchanged and single-threaded/None cases serialize identically (ADR-0008 additive-only). They are kept offRunStarted, whose exact wire bytes are pinned. - The core leaves them
None— it carries fields it never fills, exactly as it carries arun_idit never generates (sans-IO preserved). The CLI wraps the event sink in a stamping closure:EventSink::new(move |ev| inner.emit(&stamp(ev))).stampruns on the worker thread (emit is synchronous, called from the scenario worker), so it reads the run-startInstantand mapsthread::current().id()→ a stable 0-based index there, at the edge. The core never sees a clock or a thread id. - Errata (2026-08-11): only
ScenarioStartedcarries aworker. As implemented,ScenarioFinishedis emitted from the main dispatcher thread, not the worker that ran the scenario, so stamping a thread index there would name the wrong one; it carries the end timestamp andworker: None. The worker identity comes fromScenarioStarted, which is emitted on the worker thread, and the timeline pairs the two.EVENTS.mdhas always described it this way — the bullet above did not. - The HTML report gains a run-level timeline: a lane per worker, each scenario a bar from its start to its finish timestamp — the Gantt/occupancy view. It renders only when the stamps are present; a record without them falls back to the N6a per-scenario waterfall alone. The view stays a pure function of the (now richer) event record.
Consequences
- Core stays sans-IO; the injection is the established
run_id/config pattern. - The JSONL record gains observability fields, additively (ADR-0008 — still the one record; no sidecar timing file).
- Snapshot honesty: old records parse unchanged, but every event in a new run now
carries a (non-deterministic)
timestamp_ms, so thereference_event_streamsnapshot changes — it needs a new insta filter ("timestamp_ms":\d+→0) plus a deliberatecargo insta review. This is allowed (the same deliberate-acceptance policy as emitter changes), and is called out rather than glossed. The HTML snapshot likewise regenerates. - Enables the cross-worker timeline; the derived HTML view remains pure over the record.
Alternatives considered
- Stamp inside the core — rejected: reading a clock in
proef-corebreaks the sans-IO invariant that makes snapshots and property tests deterministic. - A second sidecar timing file — rejected: ADR-0008 makes the JSONL event stream the record; a parallel timing artifact is a second record format.
- Derive occupancy from
duration_msalone (no injected fields) — impossible: without an absolute clock there is no way to know which scenarios overlapped in wall-clock time; only the sequential per-scenario waterfall (N6a) is derivable, which is why it shipped first. - Wall-clock (unix ms) instead of run-relative — rejected: run-relative starts the timeline at 0, is cleaner to render, and avoids putting absolute wall-clock in the record.
ADR-0016 — OpenAPI → suite generator (scope decision)
Status: Proposed — recommendation: defer; the oracle/drift mode is permanently rejected regardless · Date: 2026-08-02
Context
A recurring market expectation of a “serious” API-testing tool is to generate tests from an OpenAPI spec (Schemathesis, Dredd, Step CI). proef has none, and the Round-2 scope stress-test (IMPROVEMENT-PLAN §12.5-A) flagged it as the single biggest capability a reviewer would call missing — while also being the one that sits on proef’s permanent charter line.
PRD §3 names, as a permanent non-goal, “API mocking/contract testing”, and the operative
expansion (IMPROVEMENT-PLAN §3) spells it: “contract testing (OpenAPI drift / Pact /
Schemathesis).” Yet the validation established that a generate-then-freeze framing — read
the spec once at the CLI edge, emit concrete editable .feature + macro packs, then execute
them through the normal deterministic pipeline — technically clears the sans-IO and
determinism objections (ADR-0012 is exact precedent for IO-at-the-edge). So the question is
genuinely open and needs an ADR to settle the boundary rather than let it erode.
Decision
The bright line (normative, decided either way): an OpenAPI spec may be a one-shot
seed — read once to emit prose + packs the author then owns, edits, and maintains —
but it may never become a recurring oracle: re-read on every run, used to
drift-check the live API against the spec, or gated on a generated-vs-committed diff. That
is OpenAPI-drift contract testing (PRD §3), and it is permanently rejected. Concretely, a
generator, if it ever exists, must have no --check/--verify/--diff mode and must
never be consulted after the initial emission.
On the narrow scaffolder itself: defer (recommendation). A one-shot
proef generate --openapi spec.yaml -o suite/ that scaffolds editable prose + packs (the
shipped bind::unbound_step stub-gen at suite scale) is defensible in-charter under the
bright line, but is not worth building now for the reasons below. If a concrete need
emerges, it may be adopted only as: CLI-only (never proef-core), deterministic given a
seed, output fully owned by the author after emission, and bound by the line above.
Consequences
- Deferring records the boundary so the omission is an intentional, documented call — and so the oracle mode is now explicitly foreclosed, not merely absent.
- Output-quality tension (the strongest argument against). OpenAPI describes an API in
endpoint terms (
POST /records/{id}/notes → 201); proef’s value is business prose (“a member posts a note to a record”). A generator produces the former, which the author must rewrite into the latter — so it saves little over hand-authoring and risks a corpus of mechanical prose that undercuts proef’s prose-first premise. - Dependency + direction cost. A robust OpenAPI 3.0/3.1 parser (the dialect split,
$ref,oneOf/allOf, discriminators) is a heavy, churny supply-chain surface to pin and clear throughcargo-deny/cargo-audit. It also introduces a new inward-ingestion direction — proef has only ever flowed outward (ADR-0010: artifacts are emitted, never imported). - One-canonical pressure. The first emission is
cargo new-style and harmless; the risk is re-generation, which would give packs a second maintenance path (regenerate vs hand-edit). The one-shot-seed discipline is not machine-enforced, so it must be a documented rule. - Already partly covered. The
#9stub-gen convenience emits a paste-readymatch:+hurl:macro skeleton for an unbound step today — the same idea at step scale, without any spec dependency.
Alternatives considered
- Full OpenAPI-drift checker (Schemathesis/Dredd-style: spec re-consulted per run to catch divergence) — permanently rejected. It is verbatim the PRD §3 non-goal, non-deterministic against a live API, and the reason the bright line exists.
- Property/fuzz-case generation from the spec (Schemathesis-style negative cases) — out. Runtime fuzzing breaks the sans-IO/deterministic-artifact invariants; frozen generated fuzz cases are a variant of the one-shot scaffolder and share its deferral.
- Build the narrow scaffolder now — declined on cost/value (output quality, dependency weight, inward direction) despite being technically in-charter under the bright line. Buildable later without a new ADR, provided it obeys this one.
- Status quo (no generator) — the recommended near-term state; authors write prose against the macro vocabulary, aided by stub-gen for missing steps.
Best-practice basis
The generate-then-freeze vs re-consult-the-oracle distinction, the sans-IO-clearance via the
ADR-0012 IO-at-the-edge precedent, and the “one --check flag from the non-goal” risk are
from the Round-2 scope validation (IMPROVEMENT-PLAN §12.5-A). Schemathesis and Dredd are the
reference generators; both re-consult the spec as an oracle, which is exactly what this ADR
forecloses. (Schemathesis,
Dredd, PRD §3, IMPROVEMENT-PLAN §12.5-A.)
ADR-0017 — proef lsp language server
Status: Accepted · Date: 2026-08-03 (implemented 2026-08-03) Design spec: docs/superpowers/specs/2026-08-03-proef-lsp-design.md
Context
Feature/pack authoring has no editor support — no live diagnostics, no jump-to-macro, no
step completion. Weak IDE support is a documented gap in Karate and the BDD field
(IMPROVEMENT-PLAN §10), and it is the one substantive unbuilt item on the Round-1 roadmap
(§5 item #11, graded ✅ FITS). proef is unusually well-placed for it: its analysis is already
headless and sans-IO — front::run yields the same Diag objects (stable code, byte
span, severity, help) regardless of driver — so an LSP is a second front-end over
existing analysis, not new analysis.
Decision
Build a proef-lsp crate (surfaced as the proef lsp subcommand) as a server-only,
generic-LSP stdio binary delivering the full v1 feature set — diagnostics,
go-to-definition, completion, and find-references. Key choices:
- Sync
lsp-server+lsp-types(rust-analyzer-family), not an async stack — honouring ADR-0006’s tokio ban. - Whole-suite model, recomputed wholesale on change (debounced), not an incremental (salsa-style) index. Mature LSPs index the whole project because cross-file features are core; they carry incremental machinery only because their scale is large. proef’s suite is tens of small files and the pipeline is milliseconds, so wholesale recompute buys the cross-file capabilities (live cross-file re-validation, find-references) at per-document simplicity. Incremental is YAGNI until a real perf ceiling.
- Second front-end over sans-IO core. The one enabling refactor: an injectable source
provider so file discovery + reading go through a trait (disk for the CLI, overlay-then-
disk for the LSP), and a collect-all mode (accumulate every diagnostic instead of
fail-fast). Both keep
proef-coresans-IO — the IO is injected, the ADR-0012 pattern. - Server-only v1; a VS Code extension is deferred (off proef’s pure-Rust brand; a thin wrapper can follow).
Consequences
- A new crate + two deps, pinned as shipped:
lsp-server 0.7.9andlsp-types 0.97.0(MIT/Apache — clean undercargo-deny); the provider trait moves theproef-corepublic-apisnapshot (a deliberate, reviewed change). lsp-types 0.97models document URIs as its ownUritype (RFC-3986), noturl::Url. The 0.97 line dropped theurldependency, so the converter and every handler key documents onUri(parsed/compared as an RFC-3986 string), neverurl::Url— a change from the pre-0.97 API that would silently fail to compile against the old assumption. Superseded — see the amendment below.- The front-end refactor (injectable provider + collect-all) touches
front.rsand its callers; the CLI path must stay behaviourally identical, guarded by the existing integration- snapshot suites.
- Highest bug-risk surface is the byte↔UTF-16 + source-normalization (BOM, trailing newline)
converter; mitigated by property tests and by reusing
tests/errors/(real spans/ranges). - A genuine competitive differentiator, and the natural depth move after the v0.4.0 breadth work — but a multi-week (L) effort; sequenced so each step (handshake → provider/collect-all/ converter → diagnostics → definition → completion → references) is independently testable.
Amendment — the types crate is gen-lsp-types, and Uri is url::Url again
lsp-types stopped receiving releases after 0.97; gen-lsp-types is the maintained
successor, generated from the LSP metamodel. proef depends on it under the original
name — lsp-types = { version = "0.11", package = "gen-lsp-types", features = ["url"] }
— which is rust-analyzer’s own aliasing pattern and leaves every lsp_types:: path in
the crate untouched. The decision above is unchanged: still sync, still the
rust-analyzer family, still server-only.
Two consequences change:
Uriisurl::Url. The generated crate gates its URI type behind features (url,fluent-uri, or a bareStringnewtype). Choosingurlcosts nothing — the embedded hurl engine already pullsurlinto the workspace graph — and buys backfrom_file_path/to_file_path, the native-path bridge the 0.97Urihad no equivalent for and thatdocuments.rstherefore hand-rolled (drive-letter prefixes, segment joining, percent-encoding). That bridge is deleted; the wrapper that remains exists only to pin the pipeline’s source-name identity rule. The consequence the original ADR recorded is retired with it, and the swap removes three crates (lsp-types,fluent-uri,serde_repr) while adding none.- Methods are enums, not string constants.
Request::METHODis now anLspRequestMethod<'static>whoseFrom<&str>falls back toCustom, so dispatch compares enum values and an unrecognised method lands in a variant rather than matching nothing.
One behaviour moved: what counts as a malformed document URI. fluent-uri rejected a
raw space; url percent-encodes it. The malformed-params test therefore asserts on a
schemeless URI, which url genuinely rejects — the guarantee under test (a bad URI
is answered with InvalidParams, never a dead server) is unchanged.
Alternatives considered
- Incremental (salsa) index — rejected for v1: it is the complexity mature LSPs accept for large scale, which a test suite does not have. Buys nothing at proef’s scale; can be added later behind the same interface.
- Per-document analysis (no cross-file model) — rejected: it would leave stale diagnostics when a pack changes and cannot do find-references; the wholesale-recompute model gives the cross-file behaviour without the index cost.
- Async LSP stack (tokio + tower-lsp) — rejected: violates ADR-0006;
lsp-serveris the sync, rust-analyzer-proven alternative. - Ship a VS Code extension in v1 — deferred: a TypeScript/npm deliverable off the pure-Rust brand and a separate release surface; generic-LSP config covers Neovim/Helix/Emacs/Sublime today, and the extension is a thin follow-up.
- Diagnostics-only or diagnostics+go-to-def MVP — considered; the operator chose the full feature set for v1 (the complete authoring experience), sequenced internally so value still lands incrementally.
Amendment — 2026-08-04 (go-to-definition gaps closed)
v1 go-to-definition resolved only feature step → macro, landing on the macro’s name key; two
narrower targets were cut and recorded only in a source comment. Both are now implemented: a
use: reference inside a pack jumps to the macro it names, and either path lands on the
macro’s match: line when one is locatable (falling back to the name key for use-only
macros). Both are best-effort text-scan locators in proef-core::pack::locate, indexed at
analyze time, following the existing sans-IO/text-scan idiom already used there — no parser
change.
ADR-0018 — Named hurl fragments: a second macro body form
Status: Accepted · Date: 2026-08-11
Context
ADR-0004 made a macro step’s HTTP payload a raw hurl block embedded in the pack YAML.
Field evidence says that was right: a real 844-line, 14-file hurl corpus was ported onto
proef through the paste path with 100% coverage — all 844 lines used only
[Asserts] (75) and [Captures] (15), with seven ordinary predicates, every one
passing through untouched (OPEN-FINDINGS §“Positive evidence”). Nothing here weakens
that.
Two things the embedded form cannot do, both structural rather than incidental:
- The block is not valid hurl. It carries
${url:…}/${secret:…}, soparse_hurl_fileonly ever sees it after a probe substitution that tries{{probe}}, then1, and accepts whichever parses (pack/validate.rs,probe_lower). Editors, hurl’s own tooling, andhurlitself see a file they cannot read. - A block has no name, so it cannot be shared.
use:composes macros, not payloads; two macros wanting the same request duplicate its text. For a corpus somebody else owns, the only route in is transcription — and a transcript drifts from its original the day after it is made.
The adoption question the worklist is now on (M1–M3) is not “can a suite be ported” — it can — but whether a team can keep both suites alive long enough to trust the new one. That needs one source of truth, not two copies.
Decision
A macro step’s body is hurl: | (as today) or ref: <fragment>, never both. A
fragment is one hurl entry in a real .hurl file, named by a comment directly above
it:
# tests/hurl/admin.hurl — runs under stock hurl, unmodified
# @proef admin.search
GET {{base}}/api/v1/admin/search/{{index}}
Authorization: Bearer {{apiToken}}
[Query]
q: {{q}}
HTTP 200
[Captures]
recordId: jsonpath "$[0].id"
bind: # pack scope
base: ${url:base}
apiToken: ${secret:apiToken}
macros:
searchRecords:
params: [q, index]
defaults: { index: records }
bind: { q: "${q}", index: "${index}" } # macro scope (quote in flow style)
steps:
- ref: admin.search
-
The annotation carries a name and nothing else, permanently. No
retry=, no key/value growth. A comment holding one identifier and zero behaviour cannot become a second configuration language, needs no parser or schema of its own, and cannot drift from the YAML. All orchestration stays in the pack. -
Names are free-form dotted, globally unique (as macro names already are), with
file.hurl#nameas a disambiguator (aspack.yaml#namealready is). -
One annotation claims exactly one entry. Forced, not preferred: a fragment reused by several macros must be composable into any position of any of them, and a multi-entry region is a fixed sequence only its original neighbours can reuse.
-
bind:maps a fragment’s{{names}}to proef values, at pack, macro and step scope, most specific winning, nothing implicit. A foreign corpus names its own variables; convention-matching would only work on files we wrote. Values resolve once per scope instantiation — one binding is one value, two bindings are two values. -
Non-secret bindings emit as per-entry
[Options] variable:;${secret:…}never does, and routes toinsert_secretas it always has, so no secret value enters an artifact (ADR-0005 intact). -
Every
{{placeholder}}must be bound, produced as a[Captures]name by a preceding step, or supplied by the fragment’s own[Options] variable:. A fragment’s interface needs no declaration anywhere: what it reads is read off hurl’s own AST and crosses the seam asScannedFragment::placeholders.The third source was missing from this list until it was found by audit, and its absence contradicted the decision above it: a file that answers its own question needs fewer variables passed in, which is precisely what makes it runnable on its own — so refusing those files refused the ones this ADR exists to accept. What a fragment supplies itself crosses the seam as
ScannedFragment::supplied_variables.It is a supplier, so it also collides with a
bind:of that name. Both reach the entry asvariable: <name>=, hurl takes the last, and the fragment’s own line is last — so the bound value would never be sent, and hurl’svariable:assigning into the run-level set rather than scoping means the loss persists into every later entry. Refused aspack::option_declared_twice, the same rule a doubly-declaredretry:gets, rather than resolved by an implicit precedence nobody wrote down.What earlier steps produce is derived differently, and deliberately: the core scans the emitted text of the steps it has already lowered (
emit::capture_names). It has to, because a preceding step may be an inlinehurl:block, which no scanner ever saw — there is noScannedFragmentfor it. So the two halves of this check reach the core by different routes, and the produced half is the one that knows hurl’s[Captures]syntax insideproef-core. A second engine would need produced names on the seam instead; nothing is scheduled, and the note is here so the asymmetry is a recorded decision rather than a discovery.
Both body forms stay, because they are not two spellings
Inline does lower-time text splicing; ref: does run-time binding. Neither
subsumes the other, and the boundary is demonstrable rather than stylistic:
tests/features/packs/breadth.yaml’s postCustomNote splices ${docstring} — a
multi-line Gherkin docstring — in as a request body. A hurl [Options] variable: value
is a single-line scalar (VariableValue is Null/Bool/Number/String), so no binding can
express it. Conversely no inline block can be named, shared, or run by stock hurl.
Splicing can substitute anything anywhere and is private to one macro. Binding is limited to what hurl can template, and buys a name, reuse, standalone runnability, and a static interface check. Choose by capability, not taste.
Consequences
The same file runs under proef test and under hurl file.hurl --variables-file … —
one source of truth, no transcription, and ADR-0004’s “copy-paste flows both ways”
becomes “no copy at all”. A fragment’s required and produced variables are machine-known,
so an unbound placeholder is a load/lower-time error with a real file:line — a check
the inline form structurally cannot perform. probe_lower’s two-candidate guessing does
not apply to fragments: they parse as authored.
Costs, stated rather than discovered later:
-
A test spans three files (
.feature→ pack →.hurl) instead of two.explainand LSP go-to-definition have to earn that back — and both now do. Go-to-definition on aref:line lands on the annotation; aref:step records the fragment it ran asfile.hurl#nameinstep_finished(an additive event field — ADR-0008 — absent for inline steps, so no pre-existing record changes a byte), whichexplainprints under a failure asvia …. The qualified spelling is the oneref:itself accepts, so a post-mortem line pastes straight back into a pack.The record carries the name because it must stand alone: by the time anyone reads it, the pack that named the fragment may say something else. That is also why the path is shortened to a project-relative spelling before it is stored —
[run] fragmentsresolves against the config file’s directory, and an absolute root would put a machine-specific path in a durable artifact and stop two checkouts’ records from comparing equal. -
A bound value gets exactly one expansion pass, and no more. hurl’s
eval_templateis a single non-recursive pass, so a rendered variable’s value is never re-parsed as a template. But a[Options] variable:value is itself evaluated as a template before it is stored (hurl-8.0.1/src/runner/options.rs:508-511), which is one pass more than the entry body gets. Sobind: { recordUrl: "${url:record}" }, whererecordis"${url:base}/api/v1/records/{{recordId}}", does work:{{recordId}}expands at option-eval time from a preceding capture, andGET {{recordUrl}}then renders the finished URL.[url]’s path table keeps working for fragments, and ADR-0012 is unchanged.The limit is the second level: a value that expands to text still containing
{{…}}will emit those characters literally. In practice that is the same requirement the bound-or-captured rule already enforces — every placeholder must be in scope at the entry that reads it.(An earlier draft of this ADR concluded that fragment paths had to move into the
.hurlfile and that[url]would narrow tobase. That was wrong: it read the non-recursiveeval_templateas applying to the binding path too. Recorded because the wrong version would have forced a needlessproef.tomlmigration.) -
Two ways to write a request body — accepted deliberately above, and the reason is recorded so the 844-line evidence is not forgotten by someone later tempted to deprecate the inline form.
-
proef reads files it does not own, so it must never write them:
fmtrefuses.hurlin directory discovery, and a foreign corpus stays byte-untouched.
Charter
PRD §3’s hurl non-goal is amended in the same change (PRD “Amendment (2026-08-11)”): it forbids generating Gherkin/macros/prose from hurl, not hurl text being an input source. Nothing here generates anything — features and macros stay hand-authored, and a fragment is inert until a macro names it. ADR-0016 stays declined on the untouched reasoning. The amendment records that OPEN-FINDINGS M3 asked for this re-examination to arrive with a measured port cost, and that it has not.
This is not M2 (mechanical equivalence between a hurl corpus and its proef port). The integration test that runs one fragment both ways proves the file is dual-runnable; it does not compare two suites’ results, and M2 stays open.
Alternatives considered
Region-claiming annotations (one comment claims every entry until the next) — fewer
comments, but a claimed region is a fixed sequence, so the reuse requirement kills it;
and pairing YAML modifiers positionally with claimed entries is exactly the coupling
emit.rs’s MapEntry.step comment already warns against (“explicit, never positional”).
One file per macro — trivial rule, but it dictates the layout of a corpus we do not
own. Metadata in the annotation (# @proef.step retry=10x300ms) — maximum locality,
at the price of a second configuration language with its own parser, schema and
finite-retry lint, inside somebody else’s file. Replacing inline entirely — rejected:
the 844-line corpus is evidence the paste path is sufficient for real work, and
${docstring} splicing has no binding equivalent. Implicit binding from config keys
— shortest packs, but a name would bind from a file the pack never mentions, and a
foreign corpus’s names rarely match anyway.
Prior art
The shape is well-established: -- name: magic comments naming queries inside valid
.sql files (yesql, HugSQL, aiosql, sqlc); # @name naming requests in JetBrains’
.http client, which can import and run them by name across files; and the OpenAPI
Initiative’s Arazzo specification, a separate declarative workflow document whose
steps reference operations by operationId defined elsewhere — the same dependency
direction chosen here, where the definition file knows nothing about its consumers.
Reuse across hurl files is an acknowledged, unresolved gap upstream
(Orange-OpenSource/hurl #317, #4574), so nothing here conflicts with a shipped hurl
feature. The known failure mode of magic comments — invisible coupling — is answered by
the name-only rule and by reading the annotation off hurl’s own AST, where
comment-to-entry attachment is already modelled.
Amendment — 2026-08-12 (the scanner reports unannotated entries, by line)
FragmentScanner returns ScannedFile { fragments, unannotated } rather than
Vec<ScannedFragment>. The named half is unchanged; the addition is the 1-based start
line of every entry carrying no # @proef annotation.
The original contract dropped those entries at scan time, on the argument — still
correct — that nothing downstream can use one, and that a corpus proef did not write is
expected to be mostly unannotated, so building a ScannedFragment for each would be the
bulk of a scan for nobody’s benefit. That reasoning covers building fragments. It does
not cover counting, and the difference showed up in the field: a 97-entry corpus port
found that missing an annotation on one entry produces a green dry-run and a silently
absent test, with no signal at scan, bind, or run time — because the entry that would
prove it was never built. Neither could any command state how many entries a corpus
held, so there was no denominator against which the gap could be noticed.
A line number costs a push and is all a listing can point at, there being no name to
print. The performance argument is therefore preserved intact: nothing extra is
constructed, and proef fragments consumes what the scan already had to walk past.
Unannotated is not an error. proef fragments --check fails on annotated fragments
no scenario runs; failing on unannotated entries requires --require-annotated. During
a port “unannotated” means not done yet; in steady state it means deliberately not
exposed, which is the premise that lets this ADR promise that pointing at a corpus you
did not write costs nothing. Gating every adopter on the porting reading would have
contradicted it, so the porting team asks for that check explicitly.
StepKindSpec also gained options, an engine-contributed recogniser mapping a raw
option key to what the core’s ADR-0007 budget rules should make of it. The fragment half
of that rule already crossed the seam (ScannedFragment::declared_options) while the
inline half matched "retry-interval:" as a literal inside proef-core — one rule at
two altitudes, and a second engine would have had its fragments linted and its inline
blocks not. Option spellings now live only in the engine that owns them. Option
baking (lower.rs) still writes hurl syntax directly; the emitter is hurl-shaped by
ADR-0010 and is a separate question this amendment does not address.
Amendment — 2026-09-05 (a fragment’s file assets are its own, and staging is what delivers them)
“The same bytes run under stock hurl and under proef” was stated for the entry’s
text. It was never true for the files that text reads. hurl resolves file,…; — a
request body, a multipart part, a file-valued assert, and [Options] output: — against
the directory of the file that wrote the reference (--file-root, defaulting to the
.hurl file’s own parent). proef resolved every such reference against the feature,
so a fragment’s asset, sitting where its own author put it, was unreachable: the same
file passed under stock hurl and failed under proef as exit 2, blaming the author for
a path that was correct.
Nothing worked around it. Moving the asset beside the feature breaks the standalone run
this ADR exists to guarantee; a reaching ../ path is refused by hurl’s sandbox; and
the advice that refusal prints — check –file-root option — names a flag proef does not
expose and no proef.toml key supplies.
Per-source resolution, delivered by staging. Each asset is copied from beside the
source that referenced it — the feature for an inline hurl: block, the fragment for a
ref: — into that scenario’s own asset root, which is then the engine’s context dir.
Two roots at once was the obvious alternative and is not available: hurl exposes one
context dir per run of entries and no per-entry override (its 42 OptionKind variants
contain no file-root), while a single batch may mix both body forms — measured, not
assumed: an inline step and a ref: step in one macro lower to one batch. Honouring two
roots would therefore mean splitting batches on the authoring layout, which trades away
the “batch maximally” rule TECH-SPEC §5 derives from run_entries building its client
per call. Staging keeps batching intact, and copying fixtures into the build output is
the standard answer to exactly this problem.
Three further consequences, each a correction rather than a cost:
- The root is per scenario. Flat staging was keyed by the asset’s bare name, so two
features that each kept a
data.jsonstaged to one file — last writer wins, silently — and the loser’s artifact replayed against the other’s bytes.artifact_slugalready refuses that trade for the.hurltext; the files it reads now match it. Two sources claiming one name within a scenario, which no per-scenario root can separate, is refused (proef::run::asset_unstageable). - Staging is load-bearing, so its failure is fatal. A missing asset used to be skipped in silence because the copy only fed the record. It now feeds the run, and a file that did not arrive is a request reading nothing, not an incomplete record.
- ADR-0010’s promise widens. “Artifacts are the executed input” held for text while
assets were read from the suite. The artifact directory is now the whole executed
input, and an artifact that reads a file says so in its replay line
(
--file-root assets/<slug>). One that reads none is byte-identical to before, which is why exactly one snapshot in the corpus moved.
The sandbox also narrows: the context dir is a directory holding only what proef staged, rather than the suite tree it used to be (TECH-SPEC §13).
ADR-0019 — Reserved tags and the authored skip
Status: Accepted · Date: 2026-08-24
Emerged from the Robot Framework capability audit (OPEN-FINDINGS, “RF wave 2”); every design fact below was verified against the tree or reproduced empirically before acceptance.
Context
proef had no way to park a scenario. A test that must not run — mid-migration,
a known-broken dependency, a seasonal flow — could only be deleted or dodged
with --tags, and both are invisible: nothing in any report says “this exists
and was deliberately not run”. Robot Framework’s SKIP model (its 4.0 headline
design, replacing criticality) is the industry convergence point: skip is a
first-class visible status with a reason that survives into every report.
Mechanically, proef already had a scenario-level Skipped status — but it
arose only from cancellation, its reason existed nowhere, and two consumers
had baked “Skipped means never-ran” into their logic (--rerun re-queues
Skipped-on-cancelled; diff reads Failed→Skipped as fixed — and
--fail-on-regression certified it). An authored skip that ignored those two
would have shipped a laundering bug, not a feature.
Decision
- A reserved tag namespace, recognized at the CLI edge.
@quarantineand@skipare the reserved tags; recognition lives in exactly one place (front::reserved), and core never reads tags — the front computes an instruction (ScenarioSpec.skip, likeexclusivebefore it) per ADR-0014’s split. Reserved tags in[run] setup/teardownfeatures have no effect: phases never pass throughbuild_specs, and skipping your whole setup deliberately is spelled by deleting the config key. - The spelling is
@skipor@skip:<reason-token>. The gherkin grammar acceptsskip:migration-pendingas one tag (verified empirically through the real pipeline). The recorded reason is the pasteable tag spelling itself —"@skip"/"@skip:migration-pending"— the same philosophy as the fragment field’sfile.hurl#name. - Authored reasons start with
@; mechanical reasons never do. That is the contract--rerunkeys on:Skipped ∧ cancelled ∧ reason not authoredre-queues as never-ran; an authored skip never re-queues. Pre-field records (no reason) read as mechanical, which they were. - No tag-list normalization. An earlier draft injected a canonical
skipatom beside@skip:xso--tags "not @skip"excluded both. Tag globs shipped first, andnot @skip*says the same thing without proef ever rewriting an authored tag list. Authored tags stay exactly authored. - A skipped scenario is selected, counted, and reasoned in every sink:
console (
∅ … — @skip:x), JUnit (<skipped message>), TAP (# SKIP @skip:x), the record (ScenarioFinished.reason, additive, schema stays 1), the HTML report,explain,flows --format json("skip"), and the harness (libtest’s ignored flag).--tagsremains the unselection mechanism — the two semantics stay distinct, as in RF. - All-selected-scenarios-skipped exits 0. Exit 2 is for faulty input; the empty-selection refusal exists for the typo’d filter whose silent green run nobody sees. An all-skipped run is neither silent (every surface prints the totals and reasons) nor accidental (each skip is authored, versioned, and visible in review). RF and pytest agree; pytest reserves its special code for empty collection, which is exactly the case that stays exit 2 here.
diffgives skip transitions their own bucket. Into-Skipped is neither fixed nor regressed (now skipped (was failing/passing)); out-of-Skipped has no meaningful baseline and takes theaddedshape.- A quarantined test-failure reaches JUnit as skipped-with-message. The
exit code already said “non-gating”; the XML said
<failure>, so Jenkins marked UNSTABLE and every dashboard contradicted the verdict. RF converts the status for the same reason. User/System faults stay failures — quarantine is for flaky tests, not broken input. --dry-runstill validates skipped scenarios. Skip is an execution-time decision, not a validation waiver — a broken-but-skipped scenario still fails--dry-run, deliberately.
Consequences
-
Library-breaking (clean break, no shims):
ScenarioSpec.skip,ScenarioOutcome.reason,Event::ScenarioFinished.reason,ScenarioRun.reason,write_junit/write_ci_reportsgain the non-gating list. Wire-additive;EVENT_SCHEMA_VERSIONstays 1. -
The sink wrappers that rebuild scenario events field-by-field (
stamp_scenario_timing,phase_sink) must thread every new field — the exhaustive constructions turn forgetting into a compile error, and the e2e test pins the stamped stream. -
A related stance this ADR writes down because the audit found it held but unwritten: control flow lives in packs (
when:conditional skip at step level,optional:soft-fail, finiteretry:) — prose stays declarative; there is no scenario-level IF/WHILE/TRY and none is planned. -
2026-09-06: a tag within a short edit distance of a reserved one (
@quarantined,@skipped,@Skip) stays an ordinary, inert tag — but it now warns (tags::reserved_tag_typo) with the spelling it likely meant, since a scenario its author believed quarantined would otherwise gate the build in silence. Short reserved words get only a case-fold or a suffix match (ship/slip/stepare one edit fromskip); the longquarantineaffords a distance-2 backstop.
ADR-0020 — Run metadata is explicit-injection-only
Status: Accepted · Date: 2026-08-24
The RF-audit wave-2 companion to ADR-0019; codifies the boundary R12-1 drew and the JUnit provenance decisions applied.
Context
A proef record could not say which commit, build, or environment produced it
— run_started carried schema and run_id alone. Robot Framework’s
--metadata name:value fills this in reports, and CI post-mortems genuinely
need it: diff across records from different commits or environments has no
context for what changed.
The hazard is on the other side. R12-1 removed harvested machine identity
from every artifact because absolute paths broke the two-checkouts
byte-equality ADR-0010 guarantees, and the JUnit sink deliberately omits
timestamp/hostname for the same reason. Metadata must not reopen that
door.
Decision
- The axis is harvested vs. handed-over, not automatic vs. manual.
proef never reads git, the hostname, wall-clock provenance, or CI
environment variables (
GITHUB_SHA,CI_COMMIT_SHA, …). If the user wants the SHA recorded, their shell harvests it:--meta commit=$(git rev-parse HEAD). What the user explicitly hands over, proef records verbatim. - One precedence chain, three scopes:
[meta]<[env.<name>.meta]<--meta k=v— the same base < env < flags shape asjobsand[url]/[vars]. A duplicate key among the flags is exit 2 (loud over last-wins); a flag overriding a config key is the designed use. There is noPROEF_META_*— the values that motivate env vars are already in the shell where the flag is typed. - The active
--envprofile name is recorded automatically, as its own field (run_started.env). It is user-chosen input to the invocation, not an observed machine fact — and without it the record is uninterpretable: the same suite deep-merges different[url]/[vars]per profile, sodiffwarns loudly on a cross-env comparison. shuffledrides the same head: with the permutation seeded byrun_id, the bool plus the id reproduces an order exactly (deferred out of--shuffle’s own change sorun_startedmoved once, not twice).- Metadata reaches the record,
explain,diff, the HTML report, the GitHub summary and the--format jsonbody — and nothing else. Never artifacts (.hurlbytes stay identical across checkouts and commits — ADR-0010, R12-2); not TAP (no slot a consumer reads); not JUnit<properties>(GitLab ignores them, Jenkins reads them only behind a non-default opt-in — same named-consumer method as R3-6, additive later if a consumer asks); not the console (the record andexplainown it). - Everything passes the sink-boundary mask — keys and values both: a secret-bearing URL pasted into either position must not survive into the record or the body. The known limit stands recorded: a token proef was never told is a secret matches no needle, the same standing as any CLI argument.
Amendment (2026-09-07) — a computed input fingerprint is not harvested metadata
proef flaky’s equivalence-class fingerprint (the inputs.json sidecar)
prompted the obvious question: does §1 forbid it? It does not, and the
boundary is worth stating so the next reader does not re-litigate it.
§1 forbids proef from harvesting an environment fact — reading git state,
the hostname, or CI variables and putting them in the record. The input
fingerprint reads none of those. It is a hash of proef’s own inputs — the
feature sources, the loaded macros and fragments, the resolved config scope —
the same category as the artifact slug (emit::artifact_slug) or the shard
hash: a derived identifier over data proef already holds, not a fact lifted
from the surrounding machine. Derived identifiers have never been in scope
here; ADR-0020 governs [meta]/--meta metadata, which this is not.
The git-commit case remains exactly as §1 requires: a user who wants
commit-based grouping hands the commit over (--meta commit=$(git rev-parse HEAD), §1’s own worked example) and proef flaky --by commit groups on it.
proef never runs git itself. So the survey’s “git tree SHA equivalence class”
splits cleanly along this ADR’s own axis — a computed fingerprint proef may
derive, plus a commit the user may hand over — and needs no new decision.
The sidecar is a derived aid like timings.json, not a second record
(ADR-0008): the JSONL event stream remains the only record format, and the
event schema is untouched (no run_started field was added).
Consequences
run_startedgainsenv,metadata,shuffled— additive, skip-serialized when unset,EVENT_SCHEMA_VERSIONstays 1 (ADR-0008 erratum extended). The empty case is byte-identical to every existing record.- Library-breaking:
ProjectConfig/EnvProfilegainmeta,exec::executetakes the merged map,RunRecord::opentakes the head trio; clean break per policy. - proef stays sans-IO in core: the CLI merges and injects; core never reads an environment.
ADR-0021 — Run discovery and rotation are separate questions
Status: Accepted · Date: 2026-09-02
Context
--run-id accepts any single path component, and TROUBLESHOOTING.md
demonstrates --run-id pr. A run named that way writes a complete, valid
record. It is then invisible to proef explain, diff, flaky, report and
--rerun whenever they resolve the latest run, because every one of them
enumerates through record::all_runs, which admits only 36-character uuid
names.
That predicate is not an oversight. rotate_runs consumes the same function
and says so:
record::all_runsis the one answer to “what is a run record here” — sorted, uuid-named directories only. Rotation adds a single further exclusion (the in-flight run), not a second enumeration rule.
and is_run_id states the reason:
rotation deletes the oldest run-shaped directories, so breadth here is a deletion hazard when the runs dir points somewhere shared.
Both are right. runs-dir may be ., so a broad predicate would let rotation
delete directories proef never created. That risk is real and the narrow
answer is the correct one — for rotation.
The defect is that one predicate serves two questions whose risks point in opposite directions:
| Question | Unsafe when | Consequence |
|---|---|---|
| May I delete this directory? | too broad | destroys user data |
| Is this a run I can show you? | too narrow | hides a real record |
Sharing the predicate meant the deletion-safety choice silently became a
visibility choice. CONFIG.md documented the rotation consequence of a custom
id (“this will never delete it”) and not the discovery one, so the surprising
half was the undocumented half.
Decision
Split the predicate along the risk, not along the file.
-
Rotation keeps the uuid rule (
is_rotatable, formerlyis_run_id) — uuid-named directories only. Unchanged in behaviour, for the reason it was written; renamed for the question it answers. A custom-id record is still never deleted by[run] keep-runs, and that stays documented. -
Discovery admits any directory containing an
events.jsonl. A directory holding proef’s own record file is a proef record; the test cannot mistaketarget/ornode_modules/for one, and — decisively — it authorises no deletion. Reading a directory that turns out not to be a record fails as a parse error naming the file, which is already how a corrupt record behaves. -
Ordering stops relying on the name — but keeps relying on the uuid.
all_runsdocumented that “uuid-v7 names sort chronologically, so lexical order is time order”, which aprdirectory breaks. The fix takes the timestamp rather than the spelling: a uuid-v7 name carries 48 bits of unix milliseconds — the moment proef minted it — which is precisely why the lexical sort worked. Runs order by that where it exists, and by directory mtime where it does not (a custom--run-idcarries no time). Both are wall-clock unix time, so the two sources compare directly.Not read from the record, though the first draft of this ADR said it would be. The head event carries
event/run_id/schemaand no timestamp at all — per-event times are injected observability onscenario_started(ADR-0015) — so “the record’s ownrun_startedtimestamp” does not exist. Reading it returnedNonefor every real record and silently ordered everything by mtime; the test that covered it passed only because its fixtures fabricated a field no record has. An ordering that depended on the suite having run at least one scenario would not be an ordering anyway.
Consequences
proef explain,diff,flaky,reportand--rerunfind custom-id runs. “Latest” means latest in time rather than latest in the alphabet, which is what every one of those commands already claimed to mean.- Rotation’s blast radius is unchanged — the one property that could have made this dangerous.
- Two predicates now exist where the code deliberately had one. That is the
cost, and it is why this is an ADR rather than a patch: the earlier
single-predicate statement was a considered position, and superseding it
needs to be on the record. The two are named for their questions
(
is_rotatable/holds_a_record) so a future reader cannot reach for the wrong one by picking the shorter name. - Ordering costs a
statper custom-id run directory and nothing at all for a uuid-named one, whose time is read straight out of its name. Bounded by[run] keep-runs(200 by default) and paid only by commands that resolve “latest”.
Alternatives considered
- Leave it, document it. The status quo before this ADR. Rejected: the
invisibility surprises exactly the user who chose a memorable id so they
could find the run again, and
--rerunsilently operating on a different run than the one just produced is the worst shape of it. - Make
--run-idreject non-uuid names. Honest, and it would remove the trap — but it also removes the feature’s point. CI archiving.proef-runs/pr-1234/by a known path is the use case--run-idexists for. - Rotate custom-id directories too, and keep one predicate. Rejected
outright: it makes
runs-dir = "."a data-loss configuration, which is the hazardis_run_idwas written to prevent.
Erratum — 2026-09-06
The JUnit report’s own uuid still had the trap this ADR removed elsewhere: a
non-uuid --run-id parsed to the nil uuid, so every --run-id ci run reported
00000000-… and collided in any consumer keyed on it. A custom id now derives a
stable UUIDv5 from its bytes; a uuid id passes through verbatim.
proef — Improvement Plan
Status: complete. Round 1 (§1–§11) and Round 2 (§12) are both shipped — every
batch (N2·N6a·N9·N8·N7·N4·N6b·N3), with N1 deferred per ADR-0013, N5 rejected, and the
OpenAPI generator settled in ADR-0016. This file is the historical feature roadmap
and its file:line citations are from 2026-07/08; the live worklist is
OPEN-FINDINGS and the running ledger is CLAUDE.md’s Status block.
· Date: 2026-07-31, appended 2026-08-02, closed 2026-08-31 · Owner: Emre
Companion docs: PRD (scope + the binding non-goals, §3), adr/ (the
invariants every item must respect), TECH-SPEC (types/pipeline),
IMPLEMENTATION-PLAN (milestones + definition of done).
0. What this is (and isn’t)
A competitive analysis of proef against the BDD and API-testing field, converted into a roadmap of candidate improvements — each validated against the current architecture at file:line, and filtered strictly through proef’s permanent non-goals (PRD §3).
Nothing here is committed work; it is the durable output of the post-M5 review round. Effort grades (S/M/L) and sequencing are advisory. Feature numbers (#1…#15) are stable identifiers used across the tables. File:line citations are as of 2026-07-31 and will drift — treat them as “start reading here”, not addresses.
Every item is in scope by construction: none proposes a second engine, mocking, contract testing, load testing, a dashboard/server, OpenTelemetry, or importing hand-written hurl — all permanent non-goals (§3 below). The work is almost entirely surfacing data the sans-IO core already computes, not new engine capability.
1. Headline finding
proef’s engine is already ahead of its cohort; the gaps are in reporting surfaces and
authoring/maintenance DX, not in HTTP power. Because proef pins hurl 8.0.1, it
already ships hurl 8.0’s full assertion arsenal (RFC 9535 JSONPath with filter functions,
type predicates isUuid/isIsoDate/isString/isObject/isList, response-time
duration < ms, the filter chain split/count/toDate/base64Decode/daysAfterNow).
Pack authors can write Karate-grade assertions today, inside raw hurl: blocks — they
are just undocumented and under-surfaced. The roadmap is therefore mostly exposure and
tooling, which is cheap, rather than engine work, which is done.
2. Where proef already wins (positioning to defend)
| proef strength | Competitor weakness it beats |
|---|---|
| One canonical way, raw-hurl-only, no escape-hatch language | Karate’s most-cited flaw: a three-language model (Gherkin + DSL + embedded JS) the moment anything gets non-trivial |
| Deterministic sans-IO core | Karate/Tavern/Newman are non-deterministic; reproducibility is now a headline selling point |
Artifacts = executed bytes (hash-locked, git-diffable, replayable hurl --test) | Postman’s opaque JSON collections (unreviewable merges) — the reason Bruno is displacing it |
| Finite retries + budgets + watchdog (ADR-0007) | hurl itself has no cancellation and unbounded retries — proef fixes its own engine’s biggest gap |
| Property-tested secret masking + typed exit codes (ADR-0009) | Most tools treat masking loosely and lack a stable exit-code contract |
| Dev-maintained macro packs = an enforced “step dictionary” | The AI-authoring trend is groping toward exactly this; proef has it structurally |
| No cloud, no account, plain text | Postman’s 2026 pricing exodus is driving the whole git-native wave |
3. Scope guardrails — what this plan will NOT propose
Permanent non-goals (PRD §3) — never revisit under “competitor parity”: further engines (browser/gRPC/etc.), API mocking, contract testing (OpenAPI drift / Pact / Schemathesis), load testing, a desktop dashboard or server mode, OpenTelemetry export, dynamic plugin loading, importing hand-written hurl into Gherkin (artifacts flow outward only), static musl/Windows binaries.
Named anti-patterns (from the 2025–2026 trend research) to avoid: silent retries / green-on-attempt-2; retry ceilings > 3 or unbounded (proef’s finite cap is already correct); permanent quarantine without owner+expiry; treating masking as a security boundary; LLM self-healing / non-deterministic test mutation; config sprawl.
A Karate-style marker DSL (#uuid, ##optional) is explicitly rejected — it would be
a second assertion mechanism competing with raw hurl predicates. Achieve the same
readability by surfacing hurl’s native predicates (item #2), not by inventing a layer.
The overriding design gate is one-canonical-way (see §6): four items must replace or augment an existing mechanism, never add a parallel knob.
4. What code validation changed about the roadmap
Three deep code-validation passes (reporting, CLI/tags, authoring) reshaped the outside-in list. Six meta-findings:
-
“Surface, don’t build” is confirmed at file:line. The sans-IO core already computes and records the data behind most items:
attempts/duration_ms/detailonStepFinished(proef-core/src/event.rs), the winningmacro_nameper bound step (bind.rs:20), deterministic seeded fakes, and hurl’s owncurl_cmd(already returned byEntryResult, currently discarded in engine-hurlsession.rs). -
Two items are already ~90% built — validation downgraded them.
- #14 seeded fakes: fakes are already deterministic (hand-rolled SplitMix64, no
randcrate,fake.rs) and already seeded by the injectedrun_id, which is already recorded inRunStarted. The feature collapses to “lettestpinrun_idthe wayartifacts --run-idalready can.” - #9 stub-gen: the “did you mean” fuzzy suggestion already exists (
bind.rs:214, levenshtein); only the paste-ready stub template is missing.
- #14 seeded fakes: fakes are already deterministic (hand-rolled SplitMix64, no
-
One item’s premise is broken — #13 impacted-only re-run. There is no stored content hash anywhere; ADR-0010 is enforced as a byte-identity
assert_eq!on outputs (crates/proef-cli/tests/execute.rs:188), not a reusable digest. The only honest impact fingerprint — the emitted.hurl— is deliberately not stable run-to-run (run_idlives in runtime globals). #13 is gated behind a determinism prerequisite. -
Two plumbing gaps gate a cluster. Scenario tags stop at the CLI edge — they never reach the runner (
ScenarioSpec/ScenarioOutcomeinrunner.rscarry no tags) or the event stream — so #15 (quarantine) and part of #8 (rerun) need new plumbing. And #4 (boolean tags) is hard-blocked byvalue_delimiter=','on--tags(main.rs:57); it is a contract-changing replace, not an add. -
The recurring architectural rule is one-canonical-way. Four items (#4, #9, #14, #15) each risk a second mechanism; each must replace or augment an existing one.
-
The exit-code contract (ADR-0009) is a live wire for #15. Quarantine changes which scenarios feed
RunSummary::exit_code(fine — the fold stays pure in core) but must extend the pinned assert_cmd tests and never let a quarantined system fault mask exit 3.
5. The validated roadmap (master table)
Status is what exists in the tree, re-verified against main on 2026-09-08 —
--help for flags, source for the rest. Verdict is the 2026-07-31 judgement of
whether the idea fits the architecture; it never meant “done”, and reading it that
way is why this table looked like a backlog when, by 2026-08-10, 13 of its 16 items
had shipped (14 today; one partial, one gated).
Status: shipped · partial · open · gated (premise rejected). Verdict legend: ✅ FITS · ⚠️ NEEDS-ADAPTATION · 🚫 premise broken. Effort: S ≤ ~1 day · M ~days · L ~weeks.
| # | Item | Status | Verdict | Lives in | Effort | Architectural truth (as of 2026-07-31) |
|---|---|---|---|---|---|---|
| 2 | Assertion cookbook (surface hurl-8.0 predicates/filters) | shipped | ✅ docs-only | docs/ | S | Lowering copies all but ${…} verbatim (resolve.rs:163); load runs the real hurl_core parser (engine-hurl/src/lib.rs:95). Predicates already work. Caveat: grammar validation, not JSONPath semantics. |
| 1 | GitHub ::error file=,line=,title= annotations | shipped | ✅ | proef-cli ci_reports.rs | S | file+line+detail already flow to write_github_summary (ci_reports.rs:98). Gate stdout vs --output json; percent-encode multiline detail. Line-only (no byte-span at runtime). |
| 5 | --curl export per request | partial | ✅ | engine-hurl session.rs | S | hurl’s EntryResult.curl_cmd is already returned, just discarded (session.rs:346). Must redact (holds resolved secrets, ADR-0005). Fold into the existing reproduce: block (exec.rs:230). Partial as of 2026-08-10: the curl line is surfaced on a failing step (exec.rs), which covers the debugging case; a per-request export flag for passing steps is still open. |
| 3a | “passed on attempt N” badge (JUnit/summary) | shipped | ✅ | proef-cli ci_reports.rs | S | attempts:u32 already on StepFinished (event.rs:72) + StepOutcome; JUnit ignores it today (ci_reports.rs:43). |
| 9 | Stub-gen for unbound steps | shipped | ⚠️ | proef-core bind.rs | S | Augment the existing did-you-mean help (bind.rs:89), zero-match arm only — not a new command. Derive {param} from quoted tokens (matcher already sheds quotes, matcher.rs:85). |
| 10 | SARIF export of --dry-run diagnostics | shipped | ✅ | proef-cli new sarif.rs | S–M | Diag (diag.rs:56) → SARIF result ~1:1: code→ruleId, byte span→region.byteOffset. Pre-populate rules[] from the closed diagnostic-code set. A parallel serializer to render.rs. |
| 14 | --seed (reproducible fakes) | shipped (as run-id) | ⚠️ | proef-cli main.rs/exec.rs | S | Thread into front::run’s existing run_id param (artifacts already exposes --run-id, main.rs:101). Caveats: arbitrary seed breaks JUnit’s UUID parse (ci_reports.rs:22); occurrence is per-scenario, not per-run — identical ${fake:X} at the same position in two different scenarios still draws the same value (a known limitation; see OPEN-FINDINGS). Note: the per-step reset this row originally cited (Refs::default() on every lower() call) was fixed in 0.6.0 — the counter now threads through lower with a high-water mark. Only the cross-scenario half remains. Shipped as the one knob §7 demanded (no separate --seed): test --run-id pins the fakes, and --shuffle seeds its permutation from the same id. |
| 7 | Dead-macro / usage report | shipped | ✅ | proef-cli new macros --usage | S–M | BoundStep.macro_name (bind.rs:20) vs packs.macros. Count use:-only macros (pattern:None, pack/mod.rs:165) as reachable via the use: graph. Report the whole corpus, not a --tags subset. |
| 6 | Self-contained HTML report | shipped | ✅ | core render_html(&[Event]) + cli write | M | Post-hoc proef report <run-id> replaying events.jsonl like explain (explain.rs:12) — the command is the HTML report, so it took no --html flag; -o picks the file. Bodies live in artifacts/ — deep-link, don’t inline. Derived view, never a second record (ADR-0008). |
| 12 | proef diff between two runs | shipped | ⚠️ | proef-cli diff.rs | M | Identity (file,scenario) (why ADR-0008 added file, event.rs:86); key step diffs on text not line (lines shift on edit). attempts+duration_ms → free flakiness/perf-regression detector. Pre-file records replay file="". |
| 4 | Boolean tag expressions (@a and not @b) | shipped | ⚠️ (replace) | grammar in core, apply in cli front.rs | M | value_delimiter=',' (main.rs:57) actively breaks and/or/not; must replace the CSV/OR contract (front.rs:388), keep empty-match=exit-2 (front.rs:382). Grammar/evaluator is deterministic → proptest/fuzz-shaped, belongs in core. |
| 8 | Rerun-only-failures (--rerun) | shipped | ⚠️ | proef-cli, reuse explain replay | M | explain already reads the latest record + failed (file,name) (explain.rs:100, event.rs:83). Needs a multi-identity predicate (today’s scenario/scenario-file filters are single-valued, exec.rs:301,321). Factor a shared record::failed_scenarios. |
| 15 | @quarantine non-gating tag | shipped | ⚠️ (contract) | thread gating:bool core+cli | M | Tags must first reach ScenarioSpec/ScenarioOutcome (P1). exit_code() stays pure in core (runner.rs:89) and skips non-gating outcomes. Events still emit the scenario → not hidden. Extend the pinned assert_cmd tests; never mask a Fault::System (exit 3). |
| 3b | True <flakyFailure> with earlier-attempt detail | shipped | ⚠️ | schema + engine-hurl | M | Needs an additive attempt_details field (ADR-0008 additive-only) + engine-hurl collecting per-retry bodies before the final one. Bigger than 3a. |
| 11 | proef lsp (feature/pack language server) | shipped | ✅ | new proef-lsp crate | L | All diagnostic substrate is headless/sans-IO already (bind, pack::load, resolve Probe mode, matcher). New: a sync lsp-server (tokio ban forbids async), a byte-offset→token API (not exposed), and a partial-results wrapper (bind/load are all-or-nothing today). Karate notably lacks good IDE support → differentiator. |
| 13 | Impacted-only re-run (content-hash) | gated | 🚫 | — (gated) | L | No input hash exists; raw-input hashing is unsound (shared packs, use: nesting, config vars fan out). Honest fingerprint = per-scenario emitted .hurl, but it is not run-to-run stable (run_id in globals). Needs a determinism prerequisite first; silent-green risk. 2026-09-07: a suite-level input hash now exists (inputs.json, for flaky windows) — not per-scenario; still gated. |
6. Prerequisites that unlock clusters
- P1 — carry scenario tags + a
gatingflag past the CLI edge intoScenarioSpec/ScenarioOutcome(runner.rs:30,121) and, additively, the event stream. Done (RF waves): tags rideScenarioSpec/ScenarioOutcomeandscenario_finished.tags; gating became the reserved-tag instruction plus the non-gating list (ADR-0019) rather than a bool. #15 and #8 both shipped on top of it. - P2 — a shared
record::failed_scenarios(run_id)+ a multi-identity scenario predicate. Reused byexplain,--rerun(#8), andproef diff(#12). - P3 — a deterministic emitted-
.hurlfingerprint (stable run-to-run despiterun_id). Prerequisite for #13; do not attempt #13 without it. 2026-09-07:proef_core::fingerprintnow hashes a run’s whole input set (feature sources, loaded macros and fragments, the resolved[url]/[vars]scope) intoinputs.json—proef flaky’s equivalence class. Suite-level and deliberately coarse (any edit ends a window), so it is not the per-scenario,run_id-independent fingerprint #13 needs; P3 stands.
7. The one-canonical-way watch-list
Each of these must fold into an existing mechanism, never ship beside it:
| Item | Must replace / augment (not duplicate) |
|---|---|
| #4 boolean tags | Replace the CSV/OR --tags semantics — no second tag syntax |
| #9 stub-gen | Augment the existing did-you-mean Diag.help — no separate proef stub command |
#14 --seed | Fold into the existing run_id determinism knob — no parallel seed unless it replaces run_id-keyed fakes |
| #15 quarantine | Exactly one non-gating tag name; must not spawn a second “skip” concept |
#5 --curl | Attach to the single reproduce: mechanism, not a parallel debug path |
8. Recommended sequencing
- Batch A — free / small, all FITS, all reuse existing data.
#2 cookbook (docs) → #1 annotations → #5
--curl→ #3a attempt badge → #9 stub. - Batch B — small, high-leverage. #10 SARIF · #7 dead-macro · #14
--seed. - Prereqs → Batch C — medium, now unblocked. Build P1+P2, then #6 HTML · #12 diff · #8 rerun · #15 quarantine · #4 boolean tags · #3b flaky-detail.
- Batch D — strategic. #11 LSP. And #13 only after committing to P3.
9. Per-item detail & competitor provenance
Each entry: what it borrows from whom → the validated architectural note. Numbers cross- reference §5.
#2 Assertion cookbook — from Karate’s fuzzy markers + hurl’s own docs. Confirmed a
docs task: resolve() leaves everything but ${…} byte-for-byte (resolve.rs:163-208,
test runtime_tier_passes_through), and pack load validates the full grammar via
hurl_core::parser::parse_hurl_file (engine-hurl/src/lib.rs:95), re-checked on the
emitted artifact (front.rs:159). Extension: ship a tests/features/ reference feature
exercising each predicate so the cookbook is snapshot-locked against hurl upgrades (the
canary catches drift).
#1 GitHub annotations — from the 2025–2026 CI-reporting shift (annotations displace
log-diving). The failures loop already prints `{file}:{line}` — {detail}
(ci_reports.rs:98). Emit ::error workflow commands as a sibling; title = scenario +
step text. Risk: stdout is owned by --output json (exec.rs:114) — gate it.
#5 --curl export — from hurl’s loved --curl; Bruno/Postman “copy as curl”. hurl
hands us EntryResult.curl_cmd already (session iterates result.entries at
session.rs:346 but reads only captures/errors/duration). Must pass Redactions
(session.rs:396) before any sink — the curl line contains resolved secrets. Cannot be
derived pre-execution (needs runtime {{…}}).
#3a “passed on attempt N” — from the flaky-test-honesty consensus (never hide a
retry). attempts is first-class (event.rs:72, step.rs:149) and already printed on
the console (report.rs:220); JUnit simply drops it. Count-based badge is S.
#9 Stub-gen — from Cucumber/Behave snippet generation. The matcher already computes
the nearest macro via closest_pattern/levenshtein (bind.rs:214, matcher.rs:248);
add a match:+hurl: | skeleton to the help text for the zero-match arm only (an
ambiguous step, bind.rs:99, must not get a stub).
#10 SARIF — from SARIF’s rise for static/validation findings inline in PRs. Dry-run
diags are a structured Vec<Diag> before miette (diag.rs:123, front.rs:64). Diag
maps ~1:1 to a SARIF result; the closed code set (one per tests/errors/ dir) pre-fills
rules[]. cli-edge serializer, no core change.
#14 --seed — from seeded-faker reproducibility. Fakes already deterministic
(fake.rs:12 SplitMix64/FNV, seeded fnv1a(run_id) ^ …), seed already recorded in
RunStarted{run_id} (event.rs:26). Design fork: alias run_id (zero core change, but
must stay uuid-parseable for JUnit) vs a dedicated recorded seed field (cleaner, but
a second knob — resolve per one-canonical-way). Known limit: the occurrence counter is
scoped per scenario (lower.rs:69-86, Refs::fakes, threaded through resolve()’s
fakes: &mut usize parameter, not reset per call) → cross-step uniqueness within a
scenario now holds, but cross-scenario uniqueness still does not: two different scenarios
each resolving ${fake:X} at the same position in their own step order get the same value.
Document before advertising “unique fakes” — it means per-scenario, not per-run or
per-entity.
#7 Dead-macro report — from Cucumber’s usage formatter marking UNUSED. Binding
records BoundStep.macro_name (bind.rs:192); iterate
front.features[].scenarios[].bound.steps[].macro_name vs packs.macros.keys()
(pack/mod.rs:116). use:-only macros (pattern:None) need reachability via the use:
graph (pack/validate.rs:628) to avoid false “unused”. Report the whole corpus.
#6 HTML report — from Cucumber/Karate/hurl HTML reports; the industry’s convergence on
the Cucumber-Messages/JSONL stream. Core render_html(&[Event]) -> String, cli writes;
best as post-hoc over events.jsonl (explain.rs already replays it) so historical runs
render. Events are pre-redacted at the sink (report.rs:124).
#12 proef diff — from Allure history / test-observability-without-OTel. Identity is
(file, scenario) (report.rs:147 ScenarioKey; ADR-0008 added file for exactly this).
Key step diffs on text, not the volatile line. attempts+duration_ms make it a
flakiness/perf-regression detector. run_id is uuid-v7 → chronology recoverable.
#4 Boolean tags — from Cucumber tag expressions (and/or/not/()). Single filter fn
tag_selected (front.rs:388), three callers. value_delimiter=',' (main.rs:57) blocks
the operator syntax → drop it, take one expression string, replace the CSV contract.
Grammar/evaluator → core (deterministic, fuzz-shaped). Preserve empty-match=exit-2.
#8 Rerun-only-failures — from Cucumber’s rerun formatter (@rerun.txt). explain
already discovers + replays the latest record and extracts failed identities
(explain.rs:63). Add a multi-identity predicate reusing build_specs (exec.rs:289).
Empty failure set → reuse no_scenarios_matched (exit 2), never silent-pass.
#15 Quarantine — from the flaky-quarantine-with-owner+expiry consensus. Thread
gating:bool from CLI (which sees scenario.lowered.tags, exec.rs:318) into
ScenarioSpec/ScenarioOutcome; exit_code() (runner.rs:89) skips non-gating outcomes
and stays pure in core. Events unchanged → scenario still reported. Extend the
cli.rs/execute.rs exit-code assertions; a quarantined Fault::System still exits 3.
#3b Flaky-failure detail — from JUnit <flakyFailure> / Allure retries. Needs an
additive attempt_details on StepFinished (ADR-0008 additive-only) and engine-hurl
collecting per-retry messages (hurl retry is per-entry internal — verify the adapter isn’t
already discarding earlier bodies).
#11 proef lsp — from Cucumber’s language server (unbound-step diagnostics,
go-to-def, completion); a gap Karate never closed. Reuses feature::parse, bind,
pack::load, resolve Probe mode, matcher — all headless, all with stable codes +
byte-offset spans that already map to editor ranges. New work: sync lsp-server (tokio
banned), byte→token API, and a “collect diags, don’t early-return” wrapper (bind/load are
all-or-nothing today).
#13 Impacted-only re-run — from selective/affected-test re-run. Premise broken: no
reusable input hash (ADR-0010 is a byte-identity assert_eq! on outputs,
execute.rs:188), and the honest fingerprint (emitted .hurl) is not run-to-run stable
because runtime globals include run_id (execute.rs:304). Gated on P3; a hash miss must
never skip a scenario that would fail (needs --force/first-run fallback).
10. Karate feature ledger — considered / adopted / rejected
Karate (github.com/karatelabs/karate) was the closest competitor and the deepest research stream. This ledger makes the “considered → decision” trail explicit, so each Karate idea is an intentional call rather than an omission.
Adopted — drove a plan item or the positioning:
| Karate feature | proef outcome |
|---|---|
Inline fuzzy markers (match response == { id: '#uuid', age: '#number' }) | #2 assertion cookbook — the same readability via hurl 8.0’s native predicates (isUuid, isIsoDate, …) surfaced in docs, not a new marker DSL. |
| Weak/immature IDE support (Karate’s own gap) | #11 LSP — reframed as a differentiator proef can win, since Karate never closed it. |
| HTML report with a timeline view | #6 self-contained HTML report (with Cucumber’s and hurl’s). |
| Three-language cognitive load (Gherkin + DSL + embedded JS) | proef’s headline positioning (§2): one-canonical-way, raw-hurl-only, sans-IO — the inverse of Karate’s most-cited flaw. |
call / callonce cross-feature reuse | Already covered by proef’s use: / with: macro composition (ADR-0004) — no new work. |
Rejected — with the reason (so it stays rejected):
| Karate feature | Why not |
|---|---|
A marker DSL (#uuid, ##optional, #? _ > 0) | A second assertion mechanism competing with raw hurl predicates — violates one-canonical-way (§3). Readability comes from #2 instead. |
| Embedded JavaScript escape hatch | Conflicts with the sans-IO deterministic core and one-canonical-way — it is the thing proef exists to avoid. |
Soft assertions (configure continueOnStepFailure) | Conflicts with proef’s deliberate stop-at-first-failed-step model (the ∅ cascade); a failed step’s downstream is intentionally not run. |
Service mocking · karate-gatling perf · UI automation | Permanent non-goals (PRD §3 and §3 above). |
Deferred — genuine candidates, not yet planned:
| Karate feature | Note |
|---|---|
Dynamic data-driven Examples (rows from a read('data.json') array) | Table-driven coverage from an external data file. Plausible, but in scope-tension with product-neutrality and the sans-IO/determinism line (an external read at lower time). Revisit if a real need appears. |
match each / schema-as-a-value reuse | Achievable today via the #2 cookbook’s hurl predicates and reusable expect: macros — no new engine feature needed. |
11. Sources (competitive research, 2026-07-31)
- Karate — match keyword / fuzzy markers / reuse: https://docs.karatelabs.io/assertions/match-keyword/, https://docs.karatelabs.io/reusability/calling-features/
- Cucumber-JS formatters (usage / rerun / snippets / html): https://github.com/cucumber/cucumber-js/blob/main/docs/formatters.md
- Reqnroll HTML report + Cucumber Messages; SpecFlow EOL: https://reqnroll.net/news/2025/06/roadmap-update-html-report/, https://reqnroll.net/news/2025/01/specflow-end-of-life-has-been-announced/
- Hurl 8.0 (RFC 9535 JSONPath,
--curl, TAP, secrets redaction) + release history: https://hurl.dev/blog/2026/04/27/announcing-hurl-8.0.0.html, https://hurl.dev/docs/filters.html - Bruno vs Postman (git-native trajectory): https://www.usebruno.com/compare/bruno-vs-postman
- GitHub Actions workflow commands (annotations + job summaries): https://docs.github.com/en/actions/reference/workflows-and-actions/workflow-commands
- Flaky-test quarantine/hardening; GitHub Actions OIDC/masking limits; Allure history: https://pie.inc/blog/flaky-tests-cicd/, https://www.stepsecurity.io/blog/github-actions-security-best-practices, https://allurereport.org/docs/how-it-works-history-files/
- Cucumber Language Server: https://github.com/cucumber/language-server
12. Round 2 — post-execution competitive re-review (2026-08-02)
Round 1 (§1–§11) is largely shipped (across the releases that followed). Round 2 re-ran the Karate + Cucumber + adjacent-landscape survey against the post-execution codebase, then put every surviving candidate through a four-stream deep code-validation pass (matcher/binding · reporting/events · lifecycle/i18n/snapshot · scope-boundary), mirroring Round 1’s method. Identifiers N1–N9 are stable and doc-local — this registry is their only sanctioned home; they never appear in code comments (per the no-task-ids-in-source rule). File:line citations validated 2026-08-02 and will drift.
12.1 What the re-survey confirmed is already shipped (positioning to defend)
The external agents, blind to the just-landed work, flagged many “gaps” that Round 1 already closed — recording them so the omission-vs-decision trail stays explicit:
| Re-flagged “gap” | Already shipped as |
|---|---|
| Cucumber snippet/stub suggestion on unbound step | #9 stub-gen (unbound-step diagnostic prints a paste-ready macro) |
| Rerun-only-failures | #8 --rerun (keyed on (file, name)) |
| Soft-fail / allow-failure tag | #15 @quarantine (non-gating, still reported) |
| Boolean tag expressions | #4 proef_core::tags (fuzzed grammar) |
“retry until assert passes” (Karate retry until) | macro retry: → hurl [Options] retry (retries until asserts pass or budget ends) |
| Data-table → step arguments | bind.rs merges | key | value | rows into macro args |
| One trial per Examples row | outline expansion → ScenarioDef per row → one harness Trial |
| Explicit skipped/pending status | Status::Skipped (post-failure steps) + Warned (optional) |
| Whole-run JSONL event record; git-native plain text; JUnit/SARIF/HTML | ADR-0008 event spine; .feature+YAML+proef.toml; the reporter family |
Convergent-evolution note (validates the architecture, nothing to adopt): Karate v2’s
karate-events.jsonl (2025) is proef’s ADR-0008 event stream re-invented; Bruno’s plain-text
git-native rise is proef’s text model; both confirm the design is industry-aligned.
12.2 The validated Round-2 roadmap (master table)
Verdict legend as §5 (✅ FITS · ⚠️ NEEDS-ADAPTATION · 🚫 rejected/premise-broken). Effort: S ≤ ~1 day · M ~days · L ~weeks.
| # | Item | Verdict | Lives in | Effort | Architectural truth (validated 2026-08-02) |
|---|---|---|---|---|---|
| N2 | Run-level SLA thresholds (p95/max(duration) gate) | ✅ | proef.toml [sla] + cli exec.rs | S–M | Strongest — zero schema change. duration_ms already on StepFinished (event.rs:74) and in the record; a pure CLI fold. Config as [sla] (env-overridable, ADR-0012), not a flag. Breach = TestFailure (exit 1) folded before exec.rs:361; malformed table = exit 2; no new exit code. Must be opt-in by presence of [sla] — absent, behaviour is byte-identical, so pinned exit-0 tests + reference snapshot are untouched. Distinct from hurl per-request duration < (aggregate vs per-entry) → keep SLA aggregate-only, one home. |
| N6a | HTML per-scenario timing waterfall | ✅ | core html.rs | S | Zero schema change. Step start-offset = cumulative sum of prior duration_ms in the scenario, width = own duration_ms; new render in html.rs:176. Cannot show cross-worker occupancy (no clock/worker id) — intra-scenario only. Ship this first. |
| N8 | i18n # language: — verify, fix, test, keep the claim | ⚠️ | core feature.rs + tests | S | Claim asserted twice (PRD.md:104, TECH-SPEC.md:110) but unverified. gherkin-0.16 honours the header transparently; proef strips no keywords. One English-only bug: feature.rs:167 detects outlines via keyword.contains("Outline"/"Template") — false under any dialect (fr Plan du scénario, de Szenariogrundriss). Blast radius small (only a no-Examples malformed outline degrades). Fix: detect via !examples.is_empty(); add a localized fixture + byte-span test. Keep the docs claim — fix, don’t retract. |
| N9 | Curated expect: shape-macro library (expectUuid, expectIsoDate, expectNonEmptyList…) | ✅ | helpers/*.yaml + docs | S | New — surfaced by validation. Augments the existing expect:/MergedAsserts mechanism (step.rs:60, emitted emit.rs:192); zero engine/core change; product-neutral (generic shapes only). Same lever as the §12.5-B ergonomic uplift. Narrows the deep-equality ergonomic gap — not the semantic one (§12.4). |
| N7 | Near-duplicate macro lint (extend proef macros) | ⚠️ | core sim-fn + cli commands.rs | S–M | Absent; reuse literal_skeleton (matcher.rs:236) + levenshtein (matcher.rs:260). Extend the shipped dead-macro report (commands.rs:270), not a load pass (those are hard errors). Tight heuristic — skeleton-equal-modulo-captures — or it false-positives on the shipped corpus (boardShows* family; activateChannel “…and ready”). Advisory JSON field beside unused (commands.rs:306), exit 0, never a gate. Drop the conjunction + “organize-by-domain” sub-lints (false-positive on shipped prose; no machine model of “domain”). |
| N4 | TAP reporter | ⚠️ | cli new tap.rs via --output tap | M | Valid, but the “surface hurl’s native TAP” rationale is wrong — proef calls run_entries in-process, never shells out; TAP must derive from the event spine (scenario = test point), like every reporter. Live Reporter (report.rs:120) → inherits sink redaction. Plan count from exec.rs:194. @quarantine → # TODO needs the non_gating set injected (it’s computed at exec.rs:313, not in the stream). One surface: --output tap (reuses the stdout-ownership machinery), never also a proef tap replay. |
| N1 | Typed parameter types in the matcher ({int}/{uuid}/custom, bind-time) | ⚠️ | matcher/bind in core | M | Genuinely absent (captures are untyped strings, matcher.rs:221; params: Vec<String>). Must be declaration-site, not inline {name:type}: 3 of 4 arg sources aren’t captures (data-table, defaults, with:, and use:-only macros have no pattern), and inline typing is invisible to proef schema. One-canonical forces a single params spelling (params: {q: uuid}, bare = any) via custom Deserialize → breaking pack migration → needs an ADR. Model on the fake::GENERATORS typed registry. Two-tier caveat: skip validation when the raw arg contains ${/{{ (resolves later) → a best-effort literal-args lint, not a type system. Diagnostic proef::bind::param_type_mismatch at bind.rs:193 (+ defaults at validate.rs:57, with: at validate.rs:353). |
| N3 | Suite-level setup/teardown (once-before / once-after) | ⚠️ | proef.toml [run] + cli exec.rs | M | Real gap (only per-scenario Background; teardown is engine-internal session.finish()). Premise correction: tags never reach the core runner — ScenarioSpec/ScenarioOutcome carry no tags/gating (runner.rs:30,131); quarantine is a CLI-edge non_gating set (exec.rs:313). So use a proef.toml [run] setup/teardown construct, not a tag (a tag would entangle with --tags/--rerun/flows/dedup + undefined ordering). Orchestrate in execute() around runner::run (exec.rs:220); state crosses only via saveAs: global, which must merge before the parallel pool snapshots the store (runner.rs:439). Explicit failure short-circuit required — an assert-failed setup is fault:None→exit 1 and would not abort the pool (cascading failures on un-seeded state); refuse to launch and surface the fault. Excluded from build_specs so it never double-runs; --dry-run unaffected (never calls runner::run). |
| N5 | Golden response snapshots | 🚫 | — | L | Rejected. Response bodies exist but are discarded (HurlResult…calls[].response.body; session reads only captures/errors). It is a second assertion mechanism competing with hurl body-asserts + expect:, and whole-body regression is already proef diff’s job. No normalization machinery exists (sink redaction masks only known injected values, report.rs:29). Secret-leak risk: backend-minted tokens/PII would be committed unredacted — against ADR-0005 (session.rs:346 already refuses to persist a capture equal to a secret). Do not build. If ever needed: diff-time over run records, never committed goldens. |
12.3 Verdict-change ledger (validation overturned the first sketch)
The deep pass is on the record because it changed conclusions — the point of validating:
| Item | First sketch | After validation | Why |
|---|---|---|---|
| N5 golden snapshots | ⚠️ candidate | 🚫 rejected | Duplicates hurl asserts + diff; no normalization; leaks backend-minted secrets |
| N4 TAP rationale | “surface hurl’s native TAP” | corrected: derive from the event spine | proef never shells out; hurl --report-tap is unreachable + per-file, not per-scenario |
| N3 selector | @setup/@teardown tag | proef.toml [run] construct | tags never reach the core runner; a tag overloads “filter” with “phase” and races the pool |
| N6 timeline | one “timeline” item | split N6a (zero-schema, now) / N6b (injected timestamp_ms, later) | true cross-worker occupancy needs an injected clock/worker id |
| N1 typed params | “small matcher tweak” | M + ADR + breaking migration | declaration-site forced by schema coherence; single spelling forced by one-canonical; literal-args-only forced by two-tier vars |
12.4 The named architectural ceiling (accepted, not a defect)
Karate’s match response == { id:'#uuid', items:'#[]' } — order-insensitive whole-body
deep-equality with type-holes, exhaustive-key checking, and one readable structural diff —
cannot be assembled under hurl-only (asserts are path-at-a-time). Reusable expect:
macros over hurl jsonpath cover per-path type/value/shape, collection membership
(contains), cardinality (count), optional keys (exists/not exists), and RFC-9535
filtered queries (AUTHORING.md:100) — the ergonomic gap, narrowed further by N9.
The semantic gap (single order-insensitive whole-body diff + exhaustiveness) stays open by
design. The Round-1 marker-DSL rejection stands (§3, ledger §10 — a second assertion
mechanism). This is a deliberate ceiling of the hurl-only bet, stated honestly, not
engineered away.
12.5 Non-goal-adjacent — explicit scope decisions (keep excluded absent an ADR)
The two biggest capabilities a market reviewer would name are on/over the PRD §3 line. The governing boundary the validation extracted: CLI-edge IO that injects values into the sans-IO core is sanctioned (ADR-0012); IO that re-shapes the corpus or acts as a recurring oracle is contract testing (out).
- A — OpenAPI → scenario generator (
proef generate). Verdict: needs an ADR; default = deferred/out-of-scope. Strict generate-then-freeze clears sans-IO/determinism (ADR-0012 precedent) and echoes #9 stub-gen at suite granularity — but the bright line is “the spec may be a one-shot seed; it may never become a recurring oracle.” It sits one--checkflag from OpenAPI-drift (§3non-goal), introduces a new inward generation direction, and pressures one-canonical-way on regeneration (a second maintenance path). Only an ADR that bans the oracle/drift mode and accepts the OpenAPI dependency can green-light even the narrow scaffolder. → Now settled in ADR-0016 (Proposed): the oracle/drift mode is permanently rejected; the narrow one-shot scaffolder is deferred (output-quality + dependency cost) but buildable later under the bright line. - B — JSON-Schema conformance assert. Verdict: shape/type conformance is
already-achievable today via
expect:+ hurl type predicates (the #2 cookbook — zero new features); the ergonomic uplift is N9 (curated shape macros). Full external.schema.jsonwhole-body validation is out — hurl has nojsonschemapredicate, and using the API’s canonical schema as oracle is drift-detection.
12.6 Prerequisites & one-canonical-way watch-list (round 2)
- P4 — injected per-event
timestamp_ms/worker(Option,skip_serializing_if), stamped by a CLI sink-wrapper on the worker thread (therun_idinjection pattern), leftNoneby the sans-IO core; kept offRunStarted(exact-bytes pinevent.rs:157). Note: old records still parse (additive holds), but the new-run reference snapshot changes and needs a new insta filter + deliberate review. Unlocks N6b. - ADR needed: N1 (params-shape migration + literal-args-only semantics); Tier-3-A OpenAPI (bans the oracle mode).
One-canonical watch-list — each must fold into an existing mechanism, never ship beside it:
| Item | Must replace / augment (not duplicate) |
|---|---|
| N1 typed params | One params spelling (name→type map); no second inline {name:type} form |
| N3 setup/teardown | Exactly one [run] setup + one teardown; not a tag, not a second Background |
| N4 TAP | One surface (--output tap); no parallel proef tap replay |
| N9 shape macros | Augment the existing expect: mechanism; never a schema/marker DSL |
| N2 SLA | Aggregate run/scenario budget only; per-request latency stays hurl duration < |
12.7 Recommended sequencing (round 2)
- Batch E — free / small / zero-schema — SHIPPED 2026-08-02: N2 SLA (opt-in) ·
N6a waterfall · N9
expect:library · N8 i18n verify+harden. - Batch F — small–medium — SHIPPED 2026-08-02: N7 near-duplicate lint ·
N4 TAP (
--output tap). - Batch G — medium, design/ADR call first — mostly SHIPPED 2026-08-02:
N6b full timeline (ADR-0015, P4 delivered) · N3 setup/teardown (ADR-0014,
[run] setup/teardown). N1 typed params — deferred (ADR-0013): the research-grounded call given proef’s deferred-heavy, string-ish corpus. - Blocked pending an ADR: Tier-3-A OpenAPI generator. Rejected: N5 golden snapshots.
12.8 Sources (round 2)
Competitor sources unchanged from §11 (Karate/Cucumber/Hurl/Bruno/Schemathesis/Pact/k6/ Playwright). Round-2 findings are code-internal — every verdict is anchored to a file:line validated 2026-08-02, not to an external claim.
proef — open findings
This is the worklist. Every open defect and gap lives here, whichever review found it. Each entry is self-contained: the evidence, the reasoning, and — where something was declined — why.
Companion: IMPROVEMENT-PLAN is the feature roadmap (own numbering, 14 of 16 shipped) and stays separate because five ADRs cite it by section number. CHANGELOG records what shipped, per release.
Provenance. Three reviews fed this list, each validated claim-by-claim against the tree and then retired into it:
| Review | Scope | Contributed |
|---|---|---|
| v0.5.3 external (2026-08-06) | 40 claims → 38 confirmed, 1 partial, 1 already fixed | the A/B/P/Q items below |
| first-run UX (0.5.3, engineer’s first 30 min) | F1–F4 | R1–R2 |
| non-technical UX (0.8.0, PRD §4 P1 calibration) | N1–N5 | R3 |
| round-7 pre-merge review of PR #13 | 4 defects + residue | §2.1–§2.4 below; three shipped in #31 |
| corpus-port report (0.8.0, a real 844-line hurl suite ported) | 12 items → 3 shipped (#41, #43), 2 premises corrected | M/E/D items below |
The review documents themselves were removed once their open items landed here; their
full text, transcripts and citations are in git history (git log --diff-filter=D -- docs/FIRST-RUN-UX-REVIEW.md docs/NON-TECHNICAL-UX-REVIEW.md). The shipped/open split
was re-checked against main on 2026-09-11 (after 0.19.0, which closed
H3, H4 and H5).
Read the citations as “start reading here”, not as addresses. They were accurate on
2026-08-06 and files have moved since; locate symbols with rg, not line numbers.
Noted after the 0.18.0 release (2026-09-10)
The exit-130 interrupt test asserts a race nothing holds open (open)
a_second_interrupt_hard_exits_with_130 (crates/proef-cli/tests/execute.rs)
failed once on gates (ubuntu-latest) and then passed on a re-run of the same
commit with no change: run 34341778587, attempt 1 red, attempt 2 green, both
at a802bfe. The diff under test was documentation only, so it cannot have
been a regression. Filed because a flake that is only ever re-run is a flake
nobody is counting — and because the test is young, added by #168 as the first
assertion anywhere on exit 130.
Verified from the failed attempt, not inferred. The panic carries the child’s stderr, and the interrupt notice is in it:
stderr:
interrupt — cancelling after current batches (a second interrupt hard-exits)
left: Some(1)
right: Some(130)
So the sequencing the test is built around worked — the first signal landed and
the handler announced itself — and the process still exited 1, the graceful
cancelled code, rather than 130. Two further facts bound what can have
happened. nextest timed the whole test at 57 ms; the scenario’s only
request is GET /slow, which the fixture answers after a deliberate
sleep(5s) (proef-fixture/src/lib.rs — the one documented exception to
that server’s own “never sleep-raced” rule). A 57 ms test never waited on that
sleep. And the banner the test synchronizes on, running N scenario(s), is
written at exec.rs:661 — before runner::run is called at exec.rs:680.
The most-supported reading, to be confirmed by a Linux reproduction rather
than assumed: the banner proves the run started, not that a batch is in
flight. Cancellation is cooperative at batch boundaries (ADR-0007), so a first
signal that wins the race against dispatch has nothing to wait for — the pool
starts already-cancelled, the scenarios record as skipped, the record closes
and the process exits, all inside the time it takes the test to spawn an
external kill(1) for the second signal. The window the test needs is the
5-second sleep; on that run the window never opened.
Why this is not just one red run. TESTING-STRATEGY §5 already names the rule — assert normalized event order, never raw interleaving, and treat wall time only as a generous upper bound — but it says parallel tests, so a signal-delivery test sits outside its letter while squarely inside its intent.
The fix shape, not applied here. The assertion is worth keeping (nothing else pins exit 130), so the answer is to make the window deterministic rather than to weaken it: the test needs a synchronization point proving the request reached the fixture, not that the run began, so the 5-second sleep is genuinely in flight when the first signal arrives. That is a fixture and test change on the one path that exercises the second-signal escape hatch; it wants its own change, and a Linux reproduction first — macOS has not reproduced it, and this is exactly the trap P5’s atomic-save half is held away from (“do not chase it on a Mac — that is how it gets fixed by coincidence”).
Ingested — the 0.18 survey (2026-09-06), validated then implemented
A check-the-world round over the CI-consumer surfaces — delivery failures,
signals, staging, the ADR-0007 budget family, the redaction boundary, the
machine sinks, and the flaky predicates — executed as waves A–F (#168–#175)
and then swept by a /simplify pass (#176, #178–#179). Recorded like every
external round so its verdicts are not re-derived.
Premises that did not survive validation — do not re-raise as filed:
- “The flaky predicates are missing.” The hardest one,
broken≠flaky, already existed inflaky.rs, with transition-counting,latent, and the quarantine lifecycle. Four gaps remained (a sample floor, hysteresis, an outage guard, an input equivalence class) — all pure folds over the retained history, no new state, advisory by design. Designed first (#174, closed unmerged once approved; its decisions: the fingerprint is the default key,new→insufficient-datais a MINOR break, and ADR-0020 takes a clarification rather than a new ADR), then shipped as #175. - “Group flakiness by git tree SHA.” Collides with ADR-0020 (proef never
harvests git state). Split along the ADR’s own axis: a proef-computed
input fingerprint (
inputs.json— feature sources + loaded macros/fragments- the resolved
[url]/[vars]scope; a sidecar, so ADR-0008’s schema stays frozen where the design had proposed arun_startedfield) plus a handed-over commit via--meta commit=…and--by commit.
- the resolved
- “Five sinks leak secrets.” They bypassed the masker for identity fields
only; secrets lower to
{{name}}and the engine pre-redacts details, so no live leak was found. An unenforced boundary, not a leak — closed per sink (#171), then made structural (Redactions::apply_outcome, #178). - “Adopt cargo-auditable, attestations, machete.” All already in place
(
release.yml,just gates); the genuine gap was on-demand coverage (#173).
Shipped: #168 (the record’s own write failure and the GitHub summary’s
reach exit 3; SIGTERM/SIGHUP graceful; a second signal exits 130 without
printing; a UUIDv5 JUnit identity for a custom --run-id) · #169 (staging
beside the file the parser read — the feature-side twin of H5, updated in
place below; --sarif lines from the carried source; the symlink and
case-insensitive edges; a 120-byte slug cap) · #170 (ADR-0007 amendment:
max-time:/retry-interval: capped, a four-hour batch ceiling,
timeout-ms = 0 refused) · #171 (every sink masks identities; lsp --env;
the CTRF key set pinned; three tests de-flaked) · #172 (warned/cancelled
in --format json, JUnit and CTRF; tags::reserved_tag_typo) · #173
(just cover, the ratchet policy) · #175 (the flaky guards).
Deferred, with dispositions:
- Restoring the Given/When/Then keyword and the
Rulename into the reporters — schema-additive, but it moves the pinned event snapshot: its own review, unscheduled. - Removing the two fake-generator aliases — a breaking change; bundle it with the next MINOR that already breaks.
secret list --format jsonandmacros --check— surfaces the survey wanted and nothing yet needs; build on a request.- A
cargo-mutantsCI job, anllvm-cov+ coverage-service job, and the immutable-releases repository setting — cannot be validated without triggering CI, and their cadence and cost are a maintainer’s decision. The coverage job, when it lands, must be a ratchet (TESTING-STRATEGY §3).
Noted while simplifying, not filed: the quarantine-match closure is
spelled three times (tap, ctrf, ci_reports) — one helper would do; and
apply_outcome clones an outcome even when the needle set is empty, a
Cow/is_empty short-circuit away from free on a secret-free run.
Ingested — the hurl-coverage audit (2026-09-05), validated claim-by-claim
The question: can every hurl test case now be wrapped in Gherkin? Answered
by enumerating hurl 8.0.1’s own surface from its AST — 8 section kinds, 7 body
byte kinds (5 multiline variants incl. GraphQL), 42 options — and checking each
against both body forms. Coverage is near-total by construction: the fragment
scanner implements hurl’s own Visitor, so it has no per-construct enumeration
to fall out of date, and only 4 of the 42 options are constrained at all (the
ADR-0007 budget rules on retry/repeat/delay/retry-interval).
Shipped: the two defects the audit found — a file,…; body in a ref:
fragment was unresolvable, and staged assets collided across scenarios. See
the ADR-0018 amendment of the same date.
A premise of the audit’s own first pass that did not survive validation.
Path-valued options were reported as sharing the file-body defect. They do not:
runner/options.rs never consults context_dir, so cacert, client-cert,
client-key and netrc-file reach curl as raw CWD-relative strings under both
runners, identically. [Options] output: is context-dir mediated
(runner/output.rs), via a later path than the option table — which is what
made the first reading look right.
Open — the two remaining gaps are by design, and stay that way
H1. One ref: names one entry. A .hurl file is usually one test case
spanning several chained entries; wrapping it means annotating each entry and
writing one ref: step per entry. Deliberate (ADR-0018 fixes the annotation at
one entry, permanently) and not silent: proef fragments prints
UNANNOTATED — not referenceable per entry with its line, and
--require-annotated exits 1. Declined rather than open: a multi-entry ref:
would have to decide where the run ends, which is the orchestration ADR-0018
keeps in YAML.
H2. A cross-entry [Options] variable: does not carry into a fragment.
hurl’s variable: assigns into one shared set that persists forward, so a
corpus file whose first entry declares variable: term=ok and whose second
reads {{term}} runs standalone but is refused by proef at --dry-run
(proef::lower::unbound_placeholder). Correct as it stands: a fragment is
independently runnable by definition, so its inputs must be satisfiable
without a neighbour having run first. The refusal is early, names the variable,
and offers the fix that preserves standalone runnability — give the fragment
its own [Options] variable:. Worth a porting note in AUTHORING if adopters
hit it; not worth weakening the check.
H3. --dry-run does not notice a missing file,…; asset. Verified: a
suite whose asset has been deleted still reports dry-run OK. A missing file
is statically knowable and --dry-run is the gate CI runs before standing an
environment up, so catching it there is the right end state. Deliberately not
done in the same change as the staging fix, for two reasons worth writing
down. dry_run has its own path and never calls build_specs, so the check
would be second code walking artifacts for assets — which “one way to do one
thing” says should instead be one shared checker both paths call. And the
reference corpus is run from temp working directories with settings passed by
environment (TESTING-STRATEGY), so a new filesystem requirement at validation
time needs its own regression pass over those tests before it can be trusted.
Until then the run-time failure is early (before the request is sent), names
the file and the directory it was sought in, and cannot be reached silently.
Closed 2026-09-11, on both of the conditions this entry set. The check is
one function with two callers, not a second walker: stage_assets split into
assets::resolve_assets — every refusal that is statically knowable, and no
destination touched — plus the copy, and --dry-run calls the first half.
Staging and validation cannot disagree about whether a suite’s assets resolve,
because they are the same code.
And the regression pass this entry asked for came back clean without needing
anything: all 712 tests pass, including the reference-corpus suites that run
from temp working directories with settings passed by environment. The reason
is the other H-item — since the feature-side twin of H5 landed, staging
resolves against LoadedFeature::read_from, the path the parser actually read,
so a new filesystem requirement at validation time does not inherit a
cwd-dependency. The concern was correct when it was written and had been
retired by a change filed under a different number.
H4. The file,…; scan in proef-core is hurl grammar the grammar guard
cannot see. emit::file_refs_in finds asset references by scanning for the
literal "file," and a closing ;. That is engine syntax living in core, and
source_guards.rs::hurl_grammar_in_core_is_the_closed_set_the_adr_names
does not catch it: engine_grammar_kind classifies fences, HTTP,
[Section] headers, method lines and key: value options, and a body
reference matches none of those — so the literal is neither on the sanctioned
list nor detected as missing from it. Pre-existing, not introduced by the
staging change (git show confirms the scan body is byte-identical to the
former file_references), which is why it was not fixed alongside it. Two
ways out, both real work: widen engine_grammar_kind so the set is closed
over the shapes ADR-0002 names rather than the shapes the guard happens to
classify — the same correction the method-line arm already records — or move
the scan behind the seam, where proef-engine-hurl’s Visitor already reads
filenames from hurl’s own AST (fragment.rs, visit_filename). The second is
the ADR-0002 answer; it needs a StepKindSpec entry beside validate,
fragments and options, and hurl_core supplies the hooks for it already
(visit_file for Bytes::File, visit_filename_param/visit_filename_value
for multipart parts — hurl_core-8.0.1/src/ast/visit.rs).
Closed 2026-09-10 — by the second way, and the first way as well. The scan
is now StepKindSpec::assets, a fourth engine hook beside validate,
fragments and options; emit() takes &[StepKindSpec] to reach it and
FrontEnd carries kinds beside the kind_to_engine table registry already
documents as a pair that must not be re-derived apart. The guard was widened
too, rather than left blind because nothing currently trips it: a bare
lowercase keyword followed by a comma (file,, hex,, base64,) is now
classified, and planting the literal back in emit.rs fails with
body "file," in emit.rs. Two things this entry predicted came true on
contact. The engine hooks are exactly the two named above — and the tempting
third, visit_filename, is the wrong one: hurl routes the [Options] file
paths through it, and output: names a file the run writes, so staging it
would demand a source that cannot exist. And the AST reading fixed the defect
this entry recorded as a consequence: file, inside a JSON body is no longer
an asset. What this entry did not anticipate is that ADR-0002’s amendment had
miscounted — it says thirteen literals, and this was the fourteenth, missing
for precisely the reason that amendment had already written down about the
method line.
The second consequence recorded below stands unchanged: collect_assets still
inspects only StepPayload::HurlEntries, never Structured. Recognition is
now the engine’s, but which payload variants carry assets at all is still
core’s assumption.
Two consequences of the text scan worth recording with it. It cannot tell a
real file,…; body from the same six characters inside a JSON or text
assertion body. And collect_assets only inspects StepPayload::HurlEntries,
never StepPayload::Structured — the variant reserved for a future non-hurl
engine — so the root (assets/<slug>/, per scenario) generalizes while the
recognition of what belongs in it does not. ADR-0002’s acceptance test
(“adding an engine leaves proef-core diff-empty”) is what would catch that,
and the seam above is what would satisfy it.
H5 — updated 2026-09-07. The 0.18 survey found and reproduced the
feature-side twin of this finding, worse than the fragment side it
records: the feature’s staging root was parent_dir(portable name) resolved
against the cwd, so a typed-absolute or config-written suite path run
from any subdirectory failed staging with exit 2 (a name’s anchor — project
root, or as-typed — is not recoverable from the string). Closed by exactly
the fix this entry prescribes, applied to the feature side: the resolved
discovery path travels beside the name (LoadedFeature::read_from) and
staging is a lookup, not a re-parse. The fragment side below still resolves
by name-join (correct while both are seeded from config.root(), per the
original analysis) and this entry stays open for it.
H5. A fragment’s directory is re-derived from its display name, inverting
SourceNaming without its canonicalize fallback. assets.rs::AssetRoots:: source_dir turns a recorded file.hurl#name back into a directory by
splitting the qualifier and joining against the project root. But that name is
produced once, at what the codebase calls the naming boundary
(front::read_corpus → naming.name(&path)), and SourceNaming::relative is
more than a strip: it falls back to comparing canonical forms precisely
because a lexical-only version already shipped a bug (a suite reached through
a symlink — macOS /tmp → /private/tmp — silently failed to match, R11-9).
The inverse here has no such fallback. The two agree today because both are
seeded from config.root() and discovery walks from that same root, so only
the lexical case is exercised; nothing enforces that they stay inverses, and
AssetRoots’ unit tests hand-build the struct rather than going through a
real SourceNaming. The deeper fix is to carry the resolved source directory
through the data model — Fragment/ScannedFragment holding the real
PathBuf beside file: String, threaded onto AssetRef — so staging is a
lookup rather than a re-parse. Not done here because it is a data-model change
across three crates, and because the record must keep carrying the portable
name: the resolved path would have to travel beside it, never replace it.
Closed 2026-09-11 — and it was one crate, not three. The prescription above
aimed the change at Fragment/ScannedFragment/AssetRef, which would have
put host paths into proef-core. The feature side had already answered this
differently and better: FeatureFile.path (core) carries the portable name and
LoadedFeature::read_from (CLI) carries the IO path beside it. Doing the
fragment side the same way keeps core untouched and makes the twins symmetric —
front::CorpusDirs records the directory each fragment file was read from, at
the naming boundary where both the name and the path are in hand, and
FrontEnd carries it beside kinds. AssetRoots::source_dir is a lookup;
there is no inverse left to drift.
Two things fell out. AssetRoots loses its project field and build_specs
its project_root argument — with nothing recomputed, the project root was
staging’s business only as the join’s left-hand side. And a fragment the corpus
never read is now a named error rather than a directory guessed from its name;
it is unreachable from a loaded suite, which is exactly why the old code’s
silent guess would never have been noticed.
Both new tests were checked against the old resolution and fail under it. The regression test is deliberately a case the join gets wrong rather than a symlink reproduction, because this entry is right that the two resolutions agree on every path a suite takes today: the defect was that nothing held them together, not that they had already come apart.
Ingested — validation round 19 (2026-09-02), validated claim-by-claim
An external round against v0.15.0+v0.16.0 (66 commits). Every finding was reproduced against the tree before being acted on, and the round’s own correction of two earlier rounds (the Rust pin) is accepted — see below.
Shipped: the P1 and all eight P2s. --rerun on a truncated record (a
silent green over a suite that never ran); the artifact slug collision
(ADR-0010, silent overwrite); diff‘s phantom “now skipped (was passing)”;
the tab exempted from the control-character guard; the unreachable Warned
scenario status and its four dead consumers; rerun composition (headline vs
page, and a non-transitive overlay); --shard-weights’ zero pileup; the two
ADR-0020 §5 metadata consumers that never received any; and the ADR-0002
grammar guard’s blind shapes — which, once taught method lines, surfaced
exactly the token the report predicted.
Two P2 sub-claims declined, with reasons:
-
Header lines in the grammar guard. The report names method and header lines as undetectable. Method lines were taught and found a real token. Header lines were not: no instance exists in core today, and the only workable heuristic (a Capitalized key with a colon) fires on ordinary diagnostic prose. Trigger to revisit: the first header literal that appears in core — at which point it should be pinned by hand rather than by pattern.
-
production_texttruncation was latent, not active. The report calls it “already the shape ofhtml.rsandpack/validate.rs”. Checked: both do carry a second#[cfg(test)] mod, but neither has production code after one, so nothing was actually unscanned. Fixed anyway (the scan now excises every test module) because it was one edit away from real.
P3s — shipped
.cargo/audit.toml’s stale quick-xml ignores (it claimed to mirror
deny.toml, which had deliberately removed them — the lockfile is on the
patched 0.41.0 line, so the nightly job was suppressing for no reason, and
would have silenced any new advisory against that line); explain dropping a
step’s authored name: while step_label’s own doc enumerates explain
among its six readers; the HTML “Slowest” section counting [run] phases into
“% of run time” while the tag table on the same page excludes them (ADR-0014);
the toolchain policy stated correctly in RELEASING.md/CLAUDE.md but not in
the normative spec that rust-toolchain.toml cites as its authority; and five
stale --output json spellings in documents describing current behaviour,
now guarded — narrowly, by an allowlist of present-tense docs, because
CHANGELOG/RELEASING/this file quote the flag as it really was.
P3 — closed by ADR-0021
-
(closed 2026-09-02 — ADR-0021, the decision this entry asked for). Split along the risk rather than the file: rotation keeps the uuid predicate (--run-idrecords are invisible tolatest,flaky,diffand--rerunis_rotatable), discovery asks whether a directory holds anevents.jsonl(holds_a_record), and ordering follows the uuid’s own embedded timestamp, falling back to directory mtime for a custom id. Not the record’srun_started, which is what this entry and the ADR’s first draft both proposed: the head event carriesevent/run_id/schemaand no time at all, so there was nothing there to read. The analysis below stands as the reasoning; it is kept because the tradeoff it names is what the ADR decides, not because the item is open. -
--run-idrecords are invisible tolatest,flaky,diffand--rerun.record::all_runsfilters onfsutil::is_run_id, which requires a 36-character uuid, so a--run-id prdirectory (which TROUBLESHOOTING demonstrates) is never enumerated.Do not “just widen it”.
rotate_runsconsumes the same predicate and says so in its own words — “all_runsis the one answer to what is a run record here” — and its narrowness is what keeps rotation from deleting user content underruns-dir = ".". Broadening the shared predicate broadens deletion. The two uses have opposite risk profiles: discovery is unsafe when narrow, rotation is unsafe when broad.So the fix is to split them, which contradicts an explicit design statement and therefore wants a decision on the record. A safe discovery predicate exists (a directory containing
events.jsonlcannot be mistaken fortarget/and deletes nothing), but ordering does not come free:all_runsdocuments that uuid-v7 names sort chronologically, so lexical order is time order — aprdirectory breaks that, andlatestwould need mtime or the record’s ownrun_started. CONFIG.md documents the rotation consequence of custom ids; it does not document the invisibility. That gap is real either way.
Noted while reviewing ADR-0021 — recorded, not scheduled
-
One doc check is still in the binary-half’s file.
docs.rsstates its own charter — it holds the checks that need a built binary, because they ask clap rather than parsing help text — andxtask docs-checkstates the mirror rule for the checks that only read files. Three of the four tests left indocs.rsgenuinely need the binary;no_current_behaviour_doc_spells_a_format_as_an_output_pathreads files and nothing else, so it belongs indocs_check()besidecheck_examplesandcheck_links. Consequence, the same one that moved the changelog check: it never runs in the fast doc-only CI step. Not moved with that one because it depends oncollect_markdownand theDESCRIBES_TODAYallowlist, both local todocs.rs— porting them is a real change, not a relocation, and it earns its own. Closed 2026-09-10: ported toxtask docs-checkascheck_output_path_spelling, and cheaper than this entry expected —living_docs()already collects the ADRs, socollect_markdownwas deleted rather than ported and “which files are documentation” stays one answer. OnlyDESCRIBES_TODAYmoved. The shrink guard was tightened in the move: it had counted ADRs into the same total, sochecked >= DESCRIBES_TODAY.len()could be satisfied bydocs/adralone, masking the one failure it exists to catch. -
--rerunreads the base record’sevents.jsonltwice.exec.rscallsrecord::read_events(&dir)for the JUnit overlay, thenrecord::rerun_candidates(&dir), which callsread_record→read_eventson the same directory. Two full reads and two full deserializations of one file, bounded only by the 256 MiB record ceiling. Pre-dates ADR-0021 and is untouched by it. The fix is small and shaped like the rest of the module —rerun_candidatestakes&[Event]rather than a&Path, and the one caller passes the events it already has — but it is a signature change on a path--rerunalone exercises, so it wants its own change, not a ride on this one. Closed 2026-09-10, in exactly that shape.rerun_candidatesalso became infallible, which surfaced a second defect this entry had not seen: the caller’s first read swallowed its error with.ok()and the second rediscovered it a line later, so which call reported a read failure was an accident of ordering. One read now, one error path. -
Discovery now costs a second
statper custom-id run, and that population is the one nothing bounds.all_runsstats each directory once forholds_a_record;began_atthen reads a uuid-v7 name’s time out of the name itself (no syscall) but falls tostd::fs::metadatafor any other name. Since rotation deliberately never deletes custom-id directories,[run] keep-runsdoes not cap that set — so a CI job minting--run-idper build pays one extra stat per historical build on every command that resolves “latest”. Accepted, not a defect: the stat is what buys correct interleaving of custom-id and uuid runs in one time order, which is the point of the ADR. Recorded because it is the one cost here that grows unbounded, and a future reader measuring a slowflakyon a long-lived runs dir should find it named.
Corrections this round made to earlier ones (accepted)
Rounds 17 and 18 reported the 1.97.1 pin as “overdue”. It was not: R18-2
changed the policy to latest stable adopted at its x.y.1 point release,
and 1.98.1 does not exist yet. The round is right that the remaining defect is
documentary, and right about where — the correction had reached
RELEASING.md and CLAUDE.md but not TECH-SPEC §15, which
rust-toolchain.toml names as its authority. Fixed in all four places.
Ingested — the 2026-09-02 survey (internal), validated then implemented
A deliberate check-the-world round: repo state against upstream releases, standards movement, and the open list itself. Recorded like every external round so its verdicts are not re-derived.
Validated as needing nothing — do not re-raise without new evidence:
- The hurl pin is current. 8.0.1 is the latest upstream stable (2026-04-28); the canary covers the next one.
- The Rust pin is correct per the written policy. 1.98.0 landed
2026-08-20; no 1.98.1 exists yet, and policy adopts at
x.y.1— a calendar item (~mid-September 2026), not a drift. notify9.0 is still a release candidate (rc.4, 2026-05);=8.2.0stands.- Release engineering already ships the modern supply-chain story —
Sigstore attestations (
attest-build-provenance@v4),.sha256sidecars, Homebrew tap, binstall metadata, SHA-pinned actions gated by pinned zizmor. The survey’s own candidate (“add attestations”) died against the tree. - Competitor movement is OpenAPI-generative testing (Schemathesis et al.) — a different product shape (generated negative tests vs. declared business scenarios); no charter-fit gap. The hurl-fidelity niche is uncontested.
Shipped from the survey (this series): [http] cookie-store = false
(hurl 8.0’s env-shaped option; the one [http] key with no per-entry
spelling at all); --ctrf (CTRF report off the JUnit fold, quarantine
parity per ADR-0019, real retryAttempts); the mid-run console write
failure latch (the deferred v0.6–v0.8 item, to its own written design);
emit::feature_stem/emit::artifact_slug closing Q6 structurally; Q2
re-verdicted closed (the #146 cache had already closed it).
Still open from the survey, dispositions unchanged: [source-links]
(build verdict of 2026-09-01, unscheduled); P13 (the CI llvm-cov job —
its local half, just cover, shipped 2026-09-07 in #173); the text-scan
honesty bundle (capture-name charset / ≤2-char methods / key_line_spans
flag — see the deferred list); P12 (measure first, alone, per the
complexity-guard lesson). Decision items untouched: E2’s split-invocation
remainder (trigger not fired), E3 (wants an ADR), R1 (wants its own spec).
The shipped-changelog duplicate headers (maintainer’s call) — closed
2026-09-10: no release carries a repeated kind heading any more, the
regrouping is recorded in CHANGELOG.md’s own preamble, and
xtask docs-check’s check_changelog_kinds fails if one returns, so the
call does not need making twice.
Shipped since validation
Kept here so the list reads as live rather than stale, and so a finding is not re-reported after it is fixed.
| ID | Finding | Shipped in |
|---|---|---|
| P10 | Abandoned-scenario events appended after RunFinished — worse than reported: past the run’s terminal event, not just the scenario’s | #15 |
| B11 | ${fake:*} collided across a scenario’s steps | #15 |
| P9 | .map.json gained phantom capture rows (fence-unaware scan, unrecognised custom methods) | #15 |
| B1 | Whitespace-only expect: produced an inverted sidecar span [9,8] | #15 |
| P2 | Non-UTF-8 PROEF_KEY/PROEF_ENV/PROEF_SECRET_<NAME> read as absent | #18 |
| P6 | Full disk: --output json exited 0 with truncated JSON | #18 |
| P7 | No stdout-side pipe-close test (both existing ones closed stderr) | #18 |
| P1 | Tee re-wrote the full slice on every write_all retry, duplicating run.log tail bytes | #18 |
| P8 | proef fmt rewrote CRLF → LF wholesale | #18 |
| Q3 | report -o outside the run dir shipped dead relative artifact hrefs | #18 |
| B8 | diff flagged a brand-new retried step as flaky | #18 |
| A3 | CLAUDE.md status stopped at post-M5 | #21 |
| N1 | First run reported system error with no explanation (NON-TECHNICAL) | #24 |
| N2 | proef macros printed identifiers, never the match: sentence | #24 |
| N3 | macros refused to list when any step failed to bind | #24 |
| N4 | unbound_step’s help led with the pack maintainer’s action | #24 |
| N5 | No document described the scenario author’s workflow | #24 |
| §8 | init announced four files and reported five | #24 |
| Q5 | Ctrl-C skipped teardown silently — cleanup never ran, nothing said so | #26 |
| Q4 | --dry-run validated neither [run] setup nor [run] teardown | #26 |
| R2 | doctor did not report a missing pack schema (FIRST-RUN F4b’s second half) | #29 |
| B7 | secret set --value put the secret in argv, and the error text steered to it | #29 |
| §2.2 | init destroyed an authored proef-pack.schema.json (round 7) | #30 |
| §2.3 | a mixed suite+phase failure lost the phase label exactly when it disambiguated | #31 |
| §2.4 | --rerun after a phase-only failure blamed filters never passed | #31 |
| §2.1 | pre-0.6.0 records reported the wrong verdict with confidence | #31 |
| — | diff counted a failing teardown as a test regression | #31 |
| P4 | fmt homogenized mixed-endings files beyond its hurl-blocks-only promise | #33 |
| P4 | proef --help described macros with pre-prose wording | #33 |
| P4 | WRITING-SCENARIOS’ two sample outputs drifted from the binary | #33 |
| — | init destroyed an authored proef-pack.schema.json (round-7 §2.2) | #30 |
| Q7 | fuzz_tag_expr compiled but was in neither fuzz loop | #30 |
| B3 | windows.yml built and tested without --locked | #30 |
| B13 | justfile gate list omitted public-api (and the fuzz gate) | #30 |
| B5 | explain/diff/report each inlined ProjectConfig::load() | #30 |
| A6 | TROUBLESHOOTING’s exit table omitted 130 | #30 |
| A4 | README’s ADR range and flag rows, and TECH-SPEC §10’s command surface, were stale | #30, #34 |
| A5 | TECH-SPEC’s publish claim and its run-dir inventory were stale | #30 |
| A1 | EDITORS.md claimed go-to-definition cannot land on a match: line | #34 |
| A7 | GETTING-STARTED’s copy of the scaffold comment had a word the scaffold does not | #34 |
| P11 | ADR-0015 described a worker on ScenarioFinished that is always None | #34 |
| B2 | a templated retry:/delay: under-counted the batch budget, abandoning healthy scenarios | #35 |
| B4 | --output json’s exit_code disagreed with the real exit after a JUnit failure | #35 |
| B6 | LSP completion snippets did not escape $/}/\ | #36 |
| B9 | GitHub annotation file= and job-summary table cells were unescaped | #36 |
| — | the --dry-run nudge echoed a command that was not the run validated (round-7) | #37 |
| P3 | --sarif emitted no startLine, so it annotated nothing | #37 |
| P5 | --watch did not retrigger on proef.toml | #37 |
| — | a run against untouched scaffold routes got no coaching (round-8 §5) | #38 |
| — | truncated-record fallback totals dropped Warned scenarios (round-7) | #39 |
| — | fmt rewrote any file handed to it, not just a pack | #40 |
| — | fmt trimmed the YAML skeleton, turning --check red outside its scope | #40 |
| C1 | negative-case authoring had no signposted catalogue form | #43 |
| C3 | expect: composition documented as a mechanism, never shown as the pattern | #43 |
| R9-1 | no proef fragments listing — neither way a fragment dies had a denominator | 0.11.0 |
| §2.1 | a bind: key nothing reads passed silently — the one authoring mistake with no signal | 0.11.0 |
| §2.2 | duplicate_fragment said “in both x and x” and offered a remedy that cannot work | 0.11.0 |
| §2.3 | unbound_placeholder named two of ADR-0018’s three supply routes | 0.11.0 |
| §3.1 | doctor did not know fragments exist — a path error surfaced as a name error | 0.11.0 |
| §3.2 | config discovery searches only up, undocumented; no way to name the file | 0.11.0 |
| §3.3 | init scaffolded only `hurl: | , so ref:` was invisible to the persona built for it |
| — | ADR-0007 value caps never crossed to fragments: retry: -1 validated clean | 0.11.0 |
Q7 is now closed (#30): fuzz_tag_expr is in both fuzz loops as well as the
compile gate.
Ingested — round 19 (2026-08-31), validated claim-by-claim
Five confirmed defects, all shipped; five checks that cleared; four external triggers re-tested. The round’s shape: the heavily-audited paths (scheduler, record gate, outline expansion, the shard×shuffle×rerun composition) were probed and found correctly defended, so the yield came from what the output surfaces contain rather than from what the core computes.
R19-1 — a step’s name: reached the artifact and nothing else (shipped)
A macro with more than one step turns one feature sentence into several engine
steps sharing a StepRef exactly. The emitter always wrote the authored
name: into the artifact’s entry comment; StepRef never carried it, so the
console, HTML report, JUnit, TAP, the job summary and explain printed the
same sentence once per step with only the status glyph between a warning and
the failure beside it. In a fresh reference run, 17 of 44 step identities
were duplicates, and the pinned event snapshot was encoding the defect —
three byte-identical step_finished for the cookie session is exercised.
Fixed by mirroring fragment (StepOutcome + step_finished), not by
extending StepRef: several engine steps share one StepRef, so the label
belongs to the engine step. Additive on the wire; schema stays 1. Retires two
untrue claims — AUTHORING.md’s “they anchor artifacts, events, and failure
output” and LoweredStep::label’s own “(events/console)”. Same class as
reproduce_hint in the R18 wave: computed all along, printed all along,
dropped by the record.
R19-2 — report -o wrote the machine into the shared file (shipped)
Absolute artifact hrefs, 12 per report, naming the author’s home directory —
in the one output built to be uploaded. 0.13.0 scrubbed machine identity from
the record (R12-1) and the record is clean; the HTML put it back. The
absolute path was deliberate and pinned by a test, but it resolves only on the
machine that produced it, which is exactly where -o output is not read. A
relative href strictly dominates. Windows CI then caught a second half the
local gate could not: the href was built with Path::display, and \ is
not a separator in a URL, so a Windows-generated report’s links were dead
either way — it is now built from components joined with /. The known
macOS-only-gate hazard, paid again.
R19-4 — one palette token failed WCAG AA, and every dark pill did (shipped)
--skip was the single token the dark block does not redefine: a grey chosen
against #0d1117 left carrying white text on white at 3.45:1. Writing the
guard rather than the fix found the larger one — .pill painted color:#fff
on status colours the dark palette tunes as text on a dark ground, so all four
dark pills sat between 2.52:1 and 3.45:1. The pill foreground is a token now.
Tests assert the ratio, not the hex, and that both palettes define the same
token set (the absence that caused it).
R19-5 — the report had one heading and no outline (shipped, narrower than filed)
Filed as “no headings at all”; the timeline already had an <h2> — the
first inventory ran against a record with no timing, so the timeline never
rendered. Corrected before implementing: only the tag table and the scenario
list lacked one. Both gained one, sharing the class the timeline already used.
R19-3 — the three post-run commands had no machine output (shipped)
explain, diff and doctor. A run directory carries no structured summary,
so anything analysing a run it did not launch had to fold events.jsonl
itself — the fold proef’s own two copies disagreed on three ways. Each object
mirrors its prose field for field; doctor had to start collecting its checks
before rendering them, so JSON is a second rendering rather than a second walk.
Cleared — checked, not defects (do not re-raise as omissions)
RecordGate’s(file, name)identity is safe:feature.rs’sdedup_namesguarantees uniqueness feature-wide and its doc names this consumer.- Duplicate step rows are not a counting bug — the record keys steps by
(text, occurrence ordinal). Only the surfaces were blind (R19-1). - Report keyboard focus is intact: no
:focusrules, but nooutline:noneeither, so native rings survive on button/anchor/summary. - The tag table is a real
<table>; it renders no rows only when a run carries no tags. - crates.io showing no
homepagefor 0.14.0 is publish lag — the field landed after that release was cut (verified by ancestry), and appears on the next publish. Confirmed 2026-09-09: publishing 0.18.0 carriedhomepage = https://emrecdr.github.io/proef/through, closing that half.documentationis still unset: for a binary crate that falls back to a docs.rs library page rather than the book, worth setting deliberately.
External triggers re-tested 2026-08-31 — three hold, one has since fired
- OpenTelemetry export stays a non-goal. OTel graduated CNCF (2026-05), so
the umbrella argument weakened, but the attributes that would carry a test
run —
test.case.name,test.case.result.status,test.suite.name,test.suite.run.status— are all still Development stability. PRD §3’s stated reason is current as written; only the re-check date moves. - CTRF was declined here and has since shipped. As re-tested on
2026-08-31 this read “still community-adoption phase; Microsoft’s test
platform has a discussion issue, not an implementation. Trigger unfired” —
accurate for its own date. The 2026-09-02 survey shipped it anyway as
--ctrf(#160), rendered off the same fold as JUnit. Corrected 2026-09-10; the two sibling statements of the same deferral, in the RF audit below and at R3-5, were stale with it. - Both sacred pins are correct. hurl 8.0.1 is the latest release
(2026-04-29) — no 8.1, no 9.0. Rust 1.97.1 is right under the written
policy: stable is 1.98.0 (2026-08-18) and
channel-rust-1.98.1.toml404s, so the point release the policy waits for does not exist yet. R18-2’s refutation survives contact with the calendar. - An MCP server is declined, with a named trigger. The largest ecosystem
shift since the last research round — Playwright, Cypress, BrowserStack,
Maestro and ReportPortal all ship one, and Claude Code / Cursor / Windsurf
consume them natively. They shipped MCP because their primary surface is a
GUI or a cloud API and an agent had no other way in. proef is CLI-first with
--format jsonand a pinned four-code exit contract: an agent already has a complete interface, and a second one is a second way to do one thing. The only real gap an agent hit was R19-3, now closed. Trigger: a concrete agent workflow that--format jsonplus exit codes cannot express. Recorded soproef lsp’s precedent is not read as an open door.
Noted, not filed
source_guards’ malformed-plural scan matches the literal (y) anywhere in
a non-comment source line, so any code with a single-character y parameter
false-positives (x.max(y) did). The guard’s intent is user-facing strings;
scanning all code is broader than that. Left alone — it is working as a guard
and tightening it to string literals is more risk than the trap is worth — but
the next author to trip it should know why.
Ingested — the deep improvement report (2026-08-25), validated claim-by-claim
A twelve-stream self-audit plus competitive/ecosystem research (five code audits, three UX audits, four research streams; ~125 findings), every load-bearing claim re-verified against the tree before acceptance and three proved empirically (measured stack-overflow abort and exponential backtracking in the tag glob; observed nondeterministic gherkin error ordering). The full report is the session artifact “proef — deep improvement report”; this section records the verdicts and what remains open.
Wave 1 — shipped (#112–#116)
- #112 — a comment on a section header no longer blinds any scan
(
[Options] # tuning+retry: -1dry-ran clean — ADR-0007’s named hole;[Captures] # idsdropped sidecar rows;[Asserts] # notedoubled a section); the delay cap learned hurl’shunit (delay: 5hvalidated clean at 5× the cap); pack-scopebind:resolves arg-free instead of in whichever macro ran first; the tag glob is the two-pointer match (oracle-property-tested — the metachar branch previously had zero generated coverage);multiline_bindrefuses\r/controls; theexpect:merge shares the emitter’s hardened response-line check. - #113 — a Ctrl-C in
--watch’s debounce window no longer launches one more full run; a delivered watcher error or rescan burst (queue overflow) retriggers instead of leaving the watch permanently deaf. Punctured and re-closed 0.12.0’s “staleness class closed for good” claim. - #114 — eleven silent-failure sites gained voices (store-poison save,
suite walker, doctor/fmt over unreadable trees, non-UTF-8 env values,
.map.json, LSP config,flakydegrade,docs-checkvacuous pass,Sinksseverity filter);fmtrecognizes every literal-block spelling. - #115 — a travelling record can no longer lie (
scenario_finished.file = ""key mismatch silently emptied every step map —flaky’sLatentverdict was unreachable anddiff --fail-on-regressioncertified green), crash (256 MiB read ceiling; saturating sums; saturatingSpan::len), or steer (rerun_of/--run-idsingle-component validation;[tag-links]URL percent-encoding + http(s)-only in both sinks). - #116 —
saveAs: globalrefuses a secret it can find (needle set, in core’sWorld, every engine covered) rather than one it can equal (engine-side, raw values only); the invariant is now property-tested as CLAUDE.md had claimed. The SLA gate applies the same@quarantinenon-gating list as the exit code.
Waves 2–5 — shipped (#118–#128), audited 2026-08-31
This section said these were open for a week after they landed. The list exists so a finding is not re-reported once it is fixed, and it failed at exactly that: a re-read sent one round toward rebuilding wave 2, and repeated two of its claims to a reader as open work. Corrected by checking the tree for each item rather than trusting the entry.
| Wave | Shipped in | Spot-checked by |
|---|---|---|
| 2 — CI-sink conformance | #118 | failure_detail_reaches_attribute_and_text_node_alike, illegal_bytes_and_ansi_never_reach_the_xml, an_oversized_summary_truncates_and_says_so, annotations_cap_at_ten_with_an_honest_notice, composed_identities_form_a_set, times_are_three_decimal_seconds |
| 3 — UX | #119–#122 | console is_terminal colour, clap_complete/clap_mangen in the archives, doctor’s project block, the report’s jump nav + data-f filter, the --format / -o split |
| 4 — diagnostics | #123–#125 | suggest_or_enumerate, code_description, proef::config::* codes, match_span in use |
| 5 — docs & distribution | #126–#128 | docs/INSTALL.md, .sha256 sidecars, the README comparison |
Two wave items did not ship, and one of them should not:
[[ATTACHMENT|path]]in a testcase’ssystem-out— declined, with a trigger. It is a Jenkins-plugin convention: GitLab and GitHub ignore it, so it buys a link for one vendor’s users who also installed the JUnit Attachments plugin. It would put a filesystem path inside an artifact built to travel — the class of defect R19-2 had just finished removing from the HTML report — and the reader’s need is already met twice over, by the reproduce-hintcurlin the failure content and by the HTML report’s own artifact deep-links. Trigger: a user on Jenkins reporting that neither reaches the artifact for them.llms-full.txt— still unshipped, and the entry that proposed it already records that the SEO case for it is empirically dead. Left as-is.
Corrections to this list’s own claims (all four were stale)
- “the exclusive-tags scheduler and
RecordGatehave no direct tests” — false.proef-core/tests/runner.rscarriesan_exclusive_scenario_never_shares_the_pool,back_to_back_exclusive_scenarios_each_get_the_pool_alone,cancelling_during_an_exclusive_drain_still_completes_the_run, and — for the gate —abandoned_scenario_emits_nothing_after_run_finished. - “
bake_entry_optionsdeserves a proptest” — it has one, inlower.rs. - “the
--format/-osplit is open” — shipped in #122. - “
match_spanis computed and unused” — it is used; the diagnostics wave wired it.
Verified against the tree (each entry says whether it is open or closed)
-
Closed (2026-09-02) by--shardbalances by hash while the timing data to balance by duration is already retained.--shard-weights, after validation found the obvious design silently wrong.shard_bucket(file, name, count)took identity only, so a 4-way split was balanced by count and not by time — and a CI matrix finishes when its slowest shard finishes. The weight now shipped is the sum of a scenario’s step durations (record::StepRun::duration_ms), which measures work rather than queue wait; the wall-clock span would have been the wrong number and the record reader does not retain it anyway.The hazard this entry existed to record. The natural implementation — “weight by the latest record in
runs-dir” — is silently incorrect for the only case sharding exists to serve. Each shard of a CI matrix runs on a separate machine with its own (usually empty)runs-dir, so every job would compute a different weight table and therefore a different assignment. Scenarios would run twice or not at all, and the suite would still report green. Nothing about that failure announces itself.Shipped shape: every run that reaches its suite writes a small
timings.jsoninto its run directory (an aborted setup has no suite to weigh, and writing its own scenarios would skew the next split with identities that never run), CI archives that one file, and each matrix job points--shard-weightsat the same copy — so the split is a pure function of (selected scenarios, that file). Weighted scenarios are placed longest-first; unweighted ones fall back to the frozen hash, and the two rules partition rather than compete, so a test added after the timings were captured still runs exactly once. A three-way matrix test asserts set equality both ways; mutating placement by one bucket drops two scenarios and the test names them.What it gives up is what hash mode was chosen for — a balanced split is not stable under insertion — which is why the flag is opt-in. See
CONFIG.md. -
Closed (2026-09-02), and the premise it was filed under was wrong.lower.rsthreads the same mutable trio through twelve functions.The original filing said “mechanical, no behaviour change: introduce a context struct and make them methods”. That would have broken the code. The closures (
resolve_in,resolve_pack_scope) takerefsandsinksas explicit parameters rather than capturing them, precisely so they remain callable while other state is mutably borrowed — and a method on&mut selfcannot be called whileselfis borrowed elsewhere. Threading was not an oversight; it was load-bearing, and validating that is what turned a rename into a design.Shipped: three bundles, each a type the code already implied —
Emit { out, refs, sinks }(the mutable outputs, always passed together),StepScope { step_ref, ctx, at }(what stays fixed for one authored step however deep expansion recurses), andFinishedfor the four values describing a completed step. The threading discipline is unchanged; only the arity is. Arity suppressions workspace-wide: 13 → 6,lower.rsat zero. -
Hurl grammar inClosed (2026-09-01) by an ADR-0002 amendment plus a guard — and this entry was wrong three times over. “~290 lines” countedproef-corevs ADR-0002’s diff-empty claim.#[cfg(test)]fixtures; the correction to “19 lines, all inlower.rs, four concerns” fixed the count and kept two errors. It is 19 lines across three files —lower.rs,emit.rsandpack/validate.rs— and the four “concerns” mostly are not concerns:is_method_line,is_section_header,is_response_lineandis_header_lineare already one canonicalpub(crate)set that all three files share.What the entry missed entirely is the finding: the seam already solved this once, on the reading side.
StepKindSpec::optionsexists, in its own words, as “the seam that keeps option spellings out ofproef-core” — added because matching"retry-interval:"as a literal meant “one rule lived at two altitudes.” It covers recognising options. The core still writesretry:,retry-interval:,delay:andvariable:as literals, so the same rule still lives at two altitudes, in the other direction. That, not the line count, is the actual asymmetry.Resolved as: the thirteen-token vocabulary is sanctioned and closed on the ADR record, pinned by
source_guards::hurl_grammar_in_core_is_the_closed_set_the_adr_names(which fails on growth and on an existing token spreading to another core module, and names both remedies in the failure). Moving the written half behind the seam is deferred with a named trigger — a second engine being scheduled — because until then it relocates seven literals that exactly one implementation will ever supply, at the cost of a public-API break.The meta-lesson, and the fifth instance of it this programme: a claim that lives only in prose decays, and decays in whichever direction makes the writer’s point. Every wrong version of this entry overstated the problem.
-
ReadingClosed (#150), and it was not the duplication it was filed as. Chasing it found that the 256 MiB record ceiling reached two of its four readers:events.jsonlis spread across seven files.explainandreporteach opened the file with a bareread_to_string, so neither had it —reporteven used the guarded reader for the base record two dozen lines below the raw read of the primary one. Both go throughrecord::read_eventsnow, and a source scan insource_guardsmakes the next reader use the same door. The folds that could disagree were already unified (record::parse_record,report::suite_totals), andexplain --format jsonhands consumers the canonical answer rather than inviting an eighth reader. -
captures_beforeis O(steps²) — deliberately left. It runs only when aref:/bind:consumes it and is bounded by scenario size, so threading a running set through lowering is churn against a bound that is not tight. -
Redaction runs inside the reporter mutex.Fixed (#149). The allocation cost had already been addressed (the miss path no longer allocates, and a clean field keeps itsArc); the structural point stood until the masking simply moved above thelock(). It reads the event and the needle set and writes neither, so it never needed the lock at all.
Analysed 2026-09-01 — none of the three was an ADR question
This section carried three items as charter questions needing a new or amended ADR. Checked against the tree, none of them needs one, and two had the wrong governing principle attached. Each entry below states the verdict and the evidence; the decision to act is a one-word answer, not a design exercise.
-
GitHub-summary permalinks to failing lines — build it; ADR-0020 is untouched. The entry assumed the commit must come from
GITHUB_SHA, which §1 forbids by name. It need not, on two counts. First, a link to the failing line already ships:github_annotations(ci_reports.rs:363, live atexec.rs:1444) emits::error file=…,line=…::per failing scenario, and--sarif(sarif.rs) carriesstartLine. GitHub resolves both against the commit the job checked out — proef never reads a SHA to make that work. The residual gap is narrow: the job-summary tables are inert text where the annotations are linked. Second, closing that gap needs only two mechanisms that are already accepted and already shipped —[tag-links](config.rs, glob → URL template with{tag}substituted, applied by the HTML tag table and the GitHub summary, documented as “base config only — a link is a project fact, not an environment one”), and ADR-0020 §1’s own worked example,--meta commit=$(git rev-parse HEAD). A[source-links]table over{file}/{line}/{commit}, with the commit handed over as metadata, sits inside both rules unchanged. §1 forbids proef harvesting the variable; it does not forbid a user handing it over — that is precisely the distinction the ADR was written to draw. It also serves GitLab, Bitbucket and self-hosted forges, which aGITHUB_SHAread never would. -
Chrome-trace export of the scheduling timeline — decline; the filed reason is false and the conclusion survives on a different one. “A second rendering of what the HTML timeline already shows” is wrong:
render_timeline(html.rs) draws one bar per scenario per worker lane, while a trace’s whole value is step-level nesting and zoom, which the report does not have. (Cited without a line number on purpose — the first version of this entry named one, and adding the neighbouringrender_slowestmoved it.) The adjacent gap that entry implied — that the page could not say which scenarios cost the most — is closed separately by the report’s ranked Slowest section; the trace question is unaffected, because that section ranks scenarios and a trace nests steps. The correct reason to decline is that the JSONL record already carries every step’s start and end (ADR-0015 injected timestamps), so a trace is a short transform of data proef publishes in full — and a second export format for already-published data is what one canonical mechanism forbids. No consumer has asked, which is the same CTRF/TAP-14 discipline applied above. Action: document the conversion recipe instead of building an exporter. Trigger: someone who has run the transform and hit something it cannot express. -
Templated report output paths — decline; the one-path rule was the wrong lens. That rule governs the resolution base (a path in
proef.tomlresolves against the config’s directory, a flag against the cwd); templating touches neither, so the two never conflicted. The governing principle is ADR-0020 §1’s axis again:-o "reports/$RUN_ID/index.html"is shell interpolation the caller already controls, and asking proef to interpolate it is asking proef to own a value that is already handed over. Per-run records also already have their mechanism —[run] runs-dirpluskeep-runsrotation (config.rs:101-111), deliberately project-wide so there is one record store and one policy — so a second per-run path scheme would be the duplicate, not the gap.
The pattern is now consistent enough to be worth stating: an item filed as
a charter question is usually an item whose governing principle was guessed.
The only-failed console sat here until #147 shipped it as --console failed,
a fourth mode on the existing flag — exactly what the entry asked for and no
ADR at all. Two more had already been decided. Check the tree and name the
actual principle before filing the next one.
Declined — do not re-raise
- In-run scenario
@retry(cucumber-rs’s headline feature, ranked first by one research stream): the CI-standards stream independently established retry-until-green as the anti-pattern (a 25%-failure bug passes 99.6% of the time under three retries) and proef’s detect-then-quarantine shape as the consensus architecture; per-stepretry:already covers polling. Document the stance in TESTING-STRATEGY instead — it reads as a gap until stated. Done 2026-09-10: stated in TESTING-STRATEGY §5, with the arithmetic, beside the determinism rules it belongs with. - CTRF — shipped 2026-09-02 as
--ctrf(#160); the “deferred trigger was checked and has not fired” verdict recorded here is superseded. It was right about the ecosystem — no CI platform ingests CTRF natively (GitLab/CircleCI are JUnit-only, GitHub has no format at all, Buildkite has its own JSON) — and that is why the entry is kept rather than deleted: the trigger genuinely never fired, and the format shipped for a different reason, that rendering it off the existing JUnit fold cost a renderer rather than a mechanism. Buildkite JSON remains the higher-yield target if a consumer ever does materialize (its span model maps 1:1 onto step outcomes). Corrected 2026-09-10. - TAP 14 (unratified branch, zero declared consumers) · Bruno-style
granular exit codes (ADR-0009 is a contract) · an OS-keychain secret
backend (second storage mechanism) · a user-level personal config file
(proef.toml is the one channel;
NO_COLORcovers terminal taste) · Karatematch withinsugar (hurl predicates cover it).
Environment note (machine-side, not repo-side)
Homebrew’s Rust (1.98.0) shadows rustup on this machine’s PATH
( Resolved 2026-09-10: the PATH is reordered —
/opt/homebrew/bin/cargo first), which breaks cargo +nightly and the
public-api gate and silently un-pins builds; Wave 1 gates were re-run
under the pinned 1.97.1 explicitly. Owner action: brew uninstall rust
or reorder PATH.~/.cargo/bin precedes /opt/homebrew/bin, and a login shell now resolves
both cargo and rustc to the pinned 1.97.1. The Homebrew formula is still
installed and harmless where it now sits; nothing needs uninstalling.
Ingested — Robot Framework capability audit (2026-08-24)
A deliberate mining of Robot Framework 7.x for transferable ideas, run as five extended-context investigations (one per adoption candidate, plus a counter-audit attacking the first-pass verdicts), every load-bearing claim reproduced against the tree before anything shipped.
RF wave 1 (shipped, #96–#101)
Detail cap at the engine boundary (RF’s 40-line rule) · tag-atom globs
(*/?, anchored; the silent-no-match became the intended selection) ·
flows feature descriptions (parsed, was dropped) · --shuffle seeded by
the run id (R3-9, one determinism knob) · reproduce_hint into the record
(the console knew more than explain did). The shard parity fix (#96) was
round 18’s, not RF’s, but shipped in the same wave.
RF wave 2 — the schema wave (all three shipped, #103–#105)
- RF-W2-skip (shipped — ADR-0019, #103) —
@skip/@skip:reason-token(prefix verified to parse as one tag);reasononScenarioFinished+ScenarioOutcome; one reserved-tag module (quarantine moves in); the two mapped collisions are the point of the work:--rerunre-queues Skipped-on-cancelled (authored reasons start with@, mechanical never do), anddiffbuckets Failed→Skipped as fixed — three-way bucketing required. All-skipped → exit 0 (ADR-0009 argument recorded); harness → libtest-mimic ignored. - RF-W2-tags (shipped — #104) —
tagsonScenarioFinishedonly (the cancel-skip path emits noStarted),exclusiveonScenarioStarted(closes R11-6), schema stays 1; HTML + GH-summary per-tag tables, suite-only per ADR-0014; the quarantinenon_gatinglist re-derivation collapses into the new one owner; D1 becomes its predicted recipe. NOT building RF’s tagstat combine/link/doc knobs. - RF-W2-meta (shipped — ADR-0020, #105) —
--meta k=v+[meta]/[env.<name>.meta]on the existing precedence chain;RunStarted.envauto-recorded (handed-over, not harvested — R12-1’s real axis); values through the one sink-boundary mask; JUnit<properties>deferred by the named-consumer method; never in artifacts; ADR codifying explicit-injection-only ships with it. Ashuffled: boolmarker rides the sameRunStartedchange (deferred out of #100 for one wire change instead of two).
Hazard both schema items must clear: stamp_scenario_timing and
phase_sink (exec.rs) rebuild scenario events field-by-field — a new field
compiles clean and is silently stripped from every stamped stream unless
threaded there, with an integration test per field.
RF wave 3 (shipped, #106–#108)
- Rerun merge-at-report (shipped — #106) — the record-composition half
of E2: a rerun’s JUnit carries the base’s not-re-run suite scenarios as
ordinary testcases, and
reportoverlays the base for one whole-suite page. Composition over records — the record files themselves never merge (ADR-0008); totals and the exit code stay the rerun’s own (ADR-0014). - Console modes (shipped — #107, extended #147) —
--console full|failed|dotted|quiet. The OSC-8 hyperlink half did not ship: printed paths stay plain until a terminal consumer asks (the same named-consumer method as JUnit<properties>). - Quarantine in JUnit (decision taken; shipped with ADR-0019, #103) —
a quarantined test-failure maps to
<skipped message="quarantined failure (non-gating): …">, so Jenkins, the dashboards and the exit code agree.
Deferred with named triggers
Report-size mechanism (first >1k-scenario record; failures-only render, one
mechanism not three knobs) · --runemptysuite (first CI consumer; settle
the early-error record first) · JUnit <properties> (a
Jenkins-keepProperties user). The --tagstatlink analogue left this list:
it shipped as [tag-links] (#108).
Rejected, and where the basis actually lives
--nostatusrc (ADR-0009 is a contract) · argfiles/ROBOT_OPTIONS
(proef.toml is the one channel) · pre-run modifiers, custom parsers,
listener API (PRD §3 non-goals + product identity) · GROUP · Set Test Message · --exitonerror (--max-fail is the one early stop) ·
robot:private (the macro listing exists to show the vocabulary).
Two stances the counter-audit showed are held but unwritten — control
flow lives in packs (when:/optional:/retry:), prose stays
declarative; and proef has no runtime extension surface, the record is the
observation API — both belong in AUTHORING or a short ADR when wave 2’s
ADR is written anyway.
Counter-audit corrections (for the record)
“--rerun ahead of RF” was half wrong — ahead on selection (cancelled-tail
union), missing the merge half entirely; see E2. “HTML report ahead of
log.html” — ahead on visualization, was behind on forensics (the
reproduce_hint gap, now closed by #101; request/response excerpts remain
a deliberate non-goal until asked). “@quarantine ≈ --skiponfailure” holds
only for exit-code CI (see wave 3). RF 7.4 added a Secret type — proef’s
redaction invariant predates it; banked as an ahead.
Ingested — round 18 (2026-08-24), validated claim-by-claim
R18-1 — the shard hash collapse, round two (confirmed — shipped)
Round 18 tested this registry’s R17-2.1 refutation instead of restating the
round-17 claim, and won the half that matters. The mechanism is arithmetic,
not statistics: FNV-1a’s multiplier is odd, so the accumulator’s low bit is
exactly the XOR-parity of the input bytes’ low bits; a scenario named after
its feature file — the commonest Gherkin convention — duplicates content
across the (file, name) identity, whose parity contributions cancel,
leaving a corpus-constant bit: N=2 → [20,0], odd buckets empty at N=4
(reproduced against the real shard_bucket, then pinned red in the balance
test before the fix). The R17 balance test could not see it by
construction — all three corpora held the file constant, the one condition
under which raw FNV behaves. Shipped: Murmur3 fmix64 finalizer on
shard_bucket (Breaking: every matrix re-deals), the mirrored corpus in
natural_corpora_spread_across_shards, and bounds recalibrated to what a
well-mixed hash yields (no empty shard at any N; the 3× skew bound at
N=2 only — a fair deal of 20 over 4 buckets legitimately produces
[2,7,5,6]). The reviewer’s own concession stands for the record: fmix64 is
mildly worse on constant-file corpora ([7,13] vs [10,10]), which is
randomness, not structure — no empties.
R18-2 — Rust pin “four days overdue” (refuted — and the policy is now written)
The pin follows the practiced policy — adopt a new stable at its x.y.1
point release, ~3–4 weeks after x.y.0 (1.98.1 expected mid-September) —
but the reviewer read CLAUDE.md’s “always latest stable Rust”, which said
otherwise. An unwritten policy that contradicts the written one is a docs
defect on our side: the policy now lives in CLAUDE.md and RELEASING.md, and
the pin bump lands on 1.98.1, as it always would have.
R18 closures
Eleven round-17 closures re-verified by the reviewer against cc75129 with
original repros; nothing reopened. The review singles out the machine-body
funnel and the flags-direction docs gate as the durable forms of their fixes.
Ingested — round 17 (2026-08-23), validated claim-by-claim
Two P1s filed; one confirmed both ways it can be read, one refuted by
measurement. Every confirmed item reproduced against b5b320a before any fix.
R17-2.1 — --shard hash collapse at power-of-two counts (refuted for constant-file corpora; corrected by R18-1)
The filed claim: FNV-1a’s unmixed low bits collapse the distribution at
N=2/4/8 (“100% of scenarios land in one shard”), fix with an fmix64
finalizer. Measured with a model calibrated against the frozen-literal test
(exact match on every pinned value), the claim inverts. Natural corpus
shapes — numbered scenarios, prose names, outline #N instances, camelCase,
verb templates, multi-file — are near-uniform under the current
fnv % count: [10,10], [11,9], [4,5,6,5] at their widths. The proposed
fmix64 is worse on the same corpora ([7,13] where FNV gives [10,10],
empty shards at N=8 that FNV does not produce): FNV’s parity-structured low
bit behaves like round-robin on templated names, which real suites are full
of. Collapse requires a degenerate corpus — every name an even-length run of
one character — which no suite exhibits. The filed measurement tables do not
reproduce from the calibrated function. What survives: no test asserted
balance — shipped as a distribution test over natural name shapes, so a
future hash change that does skew fails loudly.
Round-18 correction: the refutation above held only where its evidence
did — every corpus it measured kept the file path constant. Round 18 showed
the varying-file half was real (see R18-1): the low bit of raw FNV is byte
parity, and a scenario named after its feature file cancels to a
corpus-constant parity — [20,0] at N=2. “Degenerate corpus only” was this
registry’s error, not the reviewer’s.
R17-2.2 — bind: validation refused input the engine accepts (shipped)
Confirmed, both halves, plus a third the round missed:
{{newUuid}}in a bind value was refused as an unbound variable; it is a hurl function (ExprKind::Function), and stock hurl 8.0.1 runs the equivalent line (reproduced both directions). “What does this text read” is now the engine’s answer —FragmentSupport::template_reads, the same AST walk the fragment scanner uses — so the tree holds one answer, not two disagreeing ones.- A sibling literal bind sorting before the bound key is a real supplier (injected lines are written and evaluated in name order) and is now accepted; a later-sorting sibling stays refused, with the ordering named in the help.
- The round’s fix list missed the ordering half: injection landed at the
head of an author
[Options]section, so the fragment-supplies-it route the check accepts was assigned too late to be read at run time. Injection now lands at the section’s end; pinned bya_fragments_own_variable_evaluates_first.
R17-2.3 / 2.4 / 2.5 — machine output and phase reporting (shipped)
An empty shard wrote prose (plus a stray-space run) where a --output json/TAP body belongs; a setup abort wrote JUnit but zero machine-stdout
bytes; a failed teardown reached no report at all. Shipped as one
mechanism each way: emit_machine_body is called by every terminating
path (pool, empty shard, both setup aborts) with ADR-0014 suite-only totals
and the path’s own exit code — the note moved to stderr under machine
output — and a failed teardown’s outcomes ride into write_junit as their
own suite (#78’s rule made symmetric; a green phase stays out). Deliberate
scope as recorded then: the GitHub summary keeps pool-only totals. The
second audit pass showed the code does not hold to it — a setup abort passes
the setup summary as the primary, so setup failures render in the GitHub
summary while teardown failures do not. Queued: unify all three CI sinks on
“a phase appears when it fails” at the write_ci_reports boundary, with
totals staying suite-only everywhere (ADR-0014).
R17-2.6 — batch (shipped)
README omitted --shard/--max-fail (and the #73 gate was blind to the
flags direction) — closed with a reverse-flags gate whose measured burden was
exactly three flags; explain’s truncated-record fallback now filters
through is_suite() (the fourth consumer #72’s helper was built for); the
canary refuses a backport older than the pin by semver ordering, not
equality; identical warnings collapse to one with a repeat count (every
class, at the front-end aggregation — bind_shadows_capture was the
motivating fifty-warning wall); quick-xml rides at quick-junit 0.7’s
in-tree copy again, one generation in the lock. P4s (all shipped in #87): the pages workflow comment now states that
upstream/’s .patch files are served; docs/runbooks/ entered both
living_docs scanners; outline identity’s positional #N is documented in
AUTHORING with the column-placeholder remedy; the #79 comment stopped
claiming file:line survives in the failure detail.
Standards note
Rust 1.98.0 released 2026-08-20 (verified against the channel manifest). The
round calls the pin overdue; house policy waits 3–4 weeks after x.y.0 and
targets x.y.1 — the window opens ~2026-09-10.
Open — round-9 residue (ingested 2026-08-12)
The review’s P1/P2 and three P3s shipped in #48 and #50. What follows is what was verified and deliberately not built, so none of it depends on remembering.
R9-1 — proef fragments has no listing command (shipped)
flows lists scenarios and macros lists the vocabulary; nothing lists the
corpus. There is no way to ask which fragments exist, which are referenced, or
which .hurl entries carry no annotation — and an unannotated entry is dropped
at scan time by design, so the tool structurally cannot report what it never
built.
Raised by a consumer migration whose coverage gate (“every @proef name is
referenced, every entry is annotated”) had to become a script that repo owns.
Not built for 0.10.0 on purpose: new public surface, and the migration was
unblocked by correcting its own gate instead.
Shipped. A second migration report (ADOPTION-REQUEST.md, 97 entries)
supplied the field evidence this entry was waiting for and ranked it first of
seven. proef fragments now names both death modes apart, lists unannotated
entries by line, and gates CI with --check; --require-annotated is opt-in
because an unannotated entry is inert by design (ADR-0018), so “not done yet”
is a porting team’s reading of that signal and not every adopter’s.
R9-2 — fuzz coverage does not reach the fragment surfaces (shipped)
fuzz_pack_load runs with an empty corpus, so ref:/bind: clash logic never
executes under fuzzing; the annotation scanner’s entry-boundary arithmetic —
proef’s own code, not hurl’s — and bake_entry_options’ textual injection are
unfuzzed entirely. Split the fuzz input into pack and corpus halves, and consider
a fuzz_fragment_scan target (nightly, accepting the native-libs cost).
Shipped, and the prescription was half wrong — measurably. Splitting the
input into pack and corpus halves was tried first and did not work: a
byte-oriented target never resolved a single ref: in 1.45 million runs,
because reaching the rules means discovering valid YAML and a matching corpus
name simultaneously. Verified by probe (panic on a resolving ref:, run the
fuzzer, see whether it fires) rather than assumed from coverage numbers — which
is the same mistake this finding is about, one level up.
What shipped instead is fuzz_fragment_binding, structure-aware: it builds a
well-formed pack and corpus from the input and spends the budget on the name
space, so every run reaches the rules. The probe fires in seconds.
fuzz_pack_load stays byte-oriented and unchanged — parser totality is a real
job and the split would only have diluted it.
The fuzz_fragment_scan half was declined for a concrete reason, not on cost
alone: cargo dependencies are package-level, so adding proef-engine-hurl to
the fuzz crate compiles hurl for all five targets and drags native libraries into
a job that has none. Hurl’s scanner is instead property-tested in
proef-engine-hurl, where those libraries already are — pinning that every
reported line lies inside the file, that entries are accounted for exactly once
in order, and that no fragment’s text runs into the entry after it. The last
assertion was added after mutation testing: the first draft passed with the
boundary deliberately broken.
Still open from this entry: bake_entry_options’ textual injection is
unfuzzed. It is lower-time, not load-time, so it sits behind lowering rather
than pack::load and needs its own target.
R9-3 — no resource bounds on the corpus read (shipped)
No per-file or file-count cap: a multi-GB .hurl is read whole on every command
that loads packs. Pairs with the read-resilience work in #48, which made the read
survivable but not bounded.
Shipped, and worse than filed by one word: not “a multi-GB file” — a 279 MB
file cost 601 MB of resident memory on proef flows, a command that never
looks at a fragment, over a file carrying no # @proef annotation at all. The
doubling is read_to_string into a String and then Arc::from(&str), which
copies.
Bounded now at 8 MiB per file and 64 MiB per corpus, measured from the directory
entry so an oversized file is never allocated (601 MB → 15 MB on the same
input). Reported through the per-file diagnostic channel unreadable_file
already established — skipped, never fatal — and applied in proef lsp too,
where the corpus is held between requests rather than for the length of one
command. The laziness promise is intact: a corpus nothing ref:s still reports
nothing and exits 0, pinned by a test.
The Arc<str> copy itself was left alone. Removing it means changing
PackSource’s type across every reader, which is a wider change than a bound
and buys a constant factor on an input that is now capped anyway.
R9-4 — a bind that shadows a capture is silent (shipped)
hurl’s variable: assigns into one shared set, so a pack- or macro-scope bind:
re-assigning a name an earlier entry captured overrides it for every later entry,
with no diagnostic. A warning shaped like option_declared_twice fits — the
difference is that this one is only decidable where the capture set is known, at
lower time.
Shipped as proef::lower::bind_shadows_capture, a warning per the verdict
above — a fixed value over a live session is sometimes deliberate. Only a
literal bind warns: a secret bind skips the [Options] path entirely, so the
earlier capture’s assignment stands and there is nothing to warn about (pinned
by a unit test). En route it was validated that unread_bind_key already
narrows the surface to binds a fragment in scope reads — the live gap was
exactly the capture-shadow shape.
R9-5 — {{x}} inside a bind value is unvalidated at lower time (shipped)
It fails at run time instead of at --dry-run: loud, but late, and the late half
is what --dry-run exists to prevent.
Shipped as the same proef::lower::unbound_placeholder the fragment check
uses — one code for one defect class — naming both the placeholder and the bind
key, anchored on the feature step (pack-line anchoring from lower time is R1’s
recorded deferral). The accepted suppliers, each pinned: an earlier step’s
capture, the fragment’s own [Options] variable: (authored lines precede the
injected ones), and a secret in scope — a run-time {{secret}} reference never
puts the value in an artifact, unlike the ${secret:…} splice that
secret_in_composite_bind refuses.
R9-6 — provenance is cwd-dependent (shipped)
Run from a subdirectory and step_finished.fragment, explain’s via, JUnit and
the diagnostics carry an absolute machine path; the record-portability claim holds
only from the project root. Relativize against the config root rather than cwd —
the same boundary [run] fragments already resolves against.
Shipped as part of R12-1, which found the same defect reaching further than this entry describes — the safe case it names, running from the project root, had stopped being safe. The prescription here was the right one and is what landed: one anchor, the config directory, for every input kind.
R9-7 — smaller edges, verified and recorded
Artifacts written inside a fragments root poison the corpus with proef’s own
output (loud, but the remedies misdirect — skip files carrying the artifact
header, or document it); a step-scope bind: key the fragment never reads is
silently baked as a run-level variable: and can shadow a later capture, and an
unused ${secret:} bind silently widens the required-secret set (warnable at step
scope, where it is decidable); a # @proef annotation placed mid-entry is
silently ignored and the resulting unknown_ref does not hint at misplacement;
proef macros prints a corpus error twice on the degraded path; same-file
duplicate annotations read as “declared in both f.hurl and f.hurl”.
The double print is broader than filed (verified 2026-08-14 while adding the
corpus bound, which inherits it). It is not specific to macros: proef fragments does it too, and to any corpus diagnostic — unreadable_fragment_file
and the new oversized_fragment_file alike. The mechanism is that
commands::fragments renders corpus.diagnostics() itself and then loads the
suite, whose failure path renders the same diagnostics again. Both land on
stderr, so the count line reads 1 error(s) under two rendered copies. Left
here rather than folded into the bound: it is a rendering decision about which
of the two sites owns corpus diagnostics, not a property of any one diagnostic.
Open — round-10 residue (ingested 2026-08-12)
Found by a cleanup review over the fragments branch, after its own gates were green. All three are consequences of what that branch added; none is a defect in what shipped before it. Recorded rather than fixed in place because each is a behaviour change, and the branch was already carrying two correctness fixes.
R10-1 — --config is honoured by the runner and ignored by the editor (shipped)
--config <path> bypasses the upward search so a proef.toml beside the suite
becomes usable. proef lsp never sees it (it re-discovers via
ProjectConfig::load_from), and --watch watches the config found by its own
fresh upward search, not the one the run was given.
So in exactly the layout the flag exists for, proef test --config … runs
green while the editor gets no [run] fragments and reports every ref: as
unknown — diagnostics disagreeing with the runner, which is the drift that makes
an editor untrustworthy.
Shipped. ProjectConfig now keeps the file it was read from and derives
root from it, rather than storing the directory and leaving every consumer that
needed the file to search again. --watch watches the config the run resolved
through; proef lsp takes the flag and lets it outrank even the client-announced
workspace root, since a named file is not a guess to be improved on. The free
config::config_path() — the fresh upward search both bugs went through — is
gone, which is what stops the class recurring. proef lsp still starts when a
named config is missing (an editor offering less beats one that will not boot),
where the runner exits 2; the asymmetry is deliberate and documented.
R10-2 — proef fragments judges reachability over a smaller universe than the runner (shipped)
[run] setup / [run] teardown are not loaded, so a fragment used only by a
phase feature counts as never run and fails --check — a false CI failure in the
workflow --check was asked for, unless the phase feature happens to sit inside
the suite directory. exec::execute already threads one corpus through both
phase validations and both phase runs; the listing needs the same universe.
R10-3 — three predicates answer “is this a fragment file?”, and they disagree (shipped)
front::fragment_extensions (exact match, and its doc claims to be “the one
place that answers this”), pack::scan_fragments (exact), and the LSP’s own
is_fragment (case-insensitive). api.HURL therefore invalidates the editor’s
corpus but is never scanned by core or discovered by the CLI.
The shared home is proef_core::engine, beside StepKindSpec — it is pure logic
over the registry, so it is sans-IO-legal, and proef-lsp cannot reach
proef-cli’s copy. Worth pairing with the deeper question the LSP predicate
raises: membership in discover_fragments() is the real test, and an extension
match also claims emitted artifacts that happen to end in .hurl.
R11-1 — proef.toml resolved its paths against two different roots (shipped)
[run] fragments resolved against the config file’s directory and suite,
setup, teardown and runs-dir resolved against the working directory, so the
same relative spelling meant two directories depending on which key it sat under.
.proef-state.json and .proef-secrets.json were cwd-anchored too and appeared
in no inventory, making two shells in one project two Worlds and two secret
stores. One rule now: written paths resolve against the config, typed paths
against the working directory.
R11-2 — --watch retriggered on a config it then ignored (shipped)
The loop watched proef.toml and reran on an edit while the rerun used the
startup snapshot, so changing [url] base produced a rerun that called the old
host. Fixed by re-reading per rerun — and by moving the startup config out of
scope, which makes the stale value unreachable from the rerun closure and the
invariant a compile error rather than a habit. Which directories are watched is
still fixed at startup, so [run] fragments and [run] suite need a restart to
be watched. runs-dir was in that list until R11-8 showed it did not belong
there: it is not a watched root but an excluded one, and freezing it was the
bug rather than the limitation.
R11-3 — --config was honoured, swallowed, or ignored depending on the command (shipped)
doctor printed the error for a missing named file and then reported on
defaults, exit 0; fmt, init, schema and secret accepted a nonexistent
path silently. Three documents called the flag global to every subcommand. A
named-but-missing file is exit 2 everywhere now; doctor stays lenient about
discovery, which is a different claim.
R11-4 / R11-5 — [run] exclusive-tags did not validate itself (shipped)
--dry-run never parsed the expression, and a well-formed expression matching
nothing was silent — both defeat the reason the setting is a config expression
rather than a reserved tag name.
R11-6 — exclusivity is invisible in the run record (shipped — RF wave 2)
Event::ScenarioStarted carries no field saying a scenario ran exclusively, so a
post-mortem cannot tell a deliberate drain from a stall: the record shows
parallelism dropping to one and nothing explaining why. An additive field is
permitted by ADR-0008, and the reporters would need to decide whether to surface
it. Filed rather than built — it is a design question about what the record
should say, not a defect, and the run behaves correctly either way.
Closed (2026-08-24): scenario_started carries additive exclusive —
the very bool the scheduler read, never re-evaluated. Surfaced in the HTML
timeline title only; every other reporter deliberately ignores it.
R11-7 — the corpus-read rule is shared, its discovery is not (closed 2026-08-23 — discovery unified: one walker, one claims predicate; the surviving asymmetry is size measurement — fs::metadata vs text length — deliberate and documented, an unsaved buffer has no file to stat)
FragmentCorpus::unreadable_file now gives both readers one diagnostic, but the
CLI walks the fragment root with std::fs while the LSP reads through its
overlay provider. That difference is real — the editor must see unsaved buffers —
so the readers stay separate. What is worth watching is that “which files are in
the corpus” is still answered twice, and only the meaning of a failed read was
unified here.
R11-8 — a runs-dir edited mid---watch fed the loop its own output (shipped)
R11-2 made each rerun re-read the config, so records went to the new runs dir
while the watcher’s exclusion still named the one frozen at startup. Every
rerun’s artifacts/*.hurl, now under an unexcluded directory, requeued the next
run: 39 runs in 12 seconds, firing real traffic, from one edit. The third outing
for this class, so the fix removes the second answer rather than resynchronising
it — each rerun registers where it is about to write, before it writes, and the
exclusion is derived from the same config the run is. Deliberately not a
uuid-shaped exclusion: --run-id names a run directory that is not uuid-shaped.
R11-9 — a relative --config was never the file --watch matched (shipped)
The watcher compared the config by exact path while notify reports events under
the spelling the OS resolved them to, so --config proef.toml matched nothing and
config edits produced no rerun — silently, because feature edits kept firing and
the loop looked alive. Two questions had been conflated: where a path points
(answered once, lexically, when the flag is stored) and whether two paths are the
same file (answered by comparing canonical forms, since absolute is not enough —
macOS’s /var → /private/var aliasing and symlinks both survive it). The same
relative path had been costing proef lsp --config go-to-definition across the
whole corpus, because documents::name_to_url refuses a relative name.
R11-10 — doctor reported on defaults over a proef.toml that would not parse (shipped)
R11-3’s discovery arm became a silent unwrap_or_default, dropping the parse
error the previous code printed: a malformed config left doctor reporting on
invented defaults and printing “all checks passed”, exit 0. A project: row now,
so it reaches worst and the exit code CI reads. Leniency still means absent —
doctor must run outside a project — not broken.
Ingested — competitive research v2 (2026-08-16), validated claim-by-claim
An external research pass (prototyped against the built 0.12.0 binary) plus its round-14 companion review. Each actionable claim was re-reproduced here before anything was written down. Disposition:
S1 — an encoded reflection of a secret defeated redaction (shipped)
The one defect in the set, confirmed by live reproduction: a server
reflecting the bearer token base64-encoded put dG9r… (trivially decodable)
into an assert-failure detail; the raw needle never fired; the encoded
credential reached the console and events.jsonl. The raw-form invariant was
intact — this violated its intent. Shipped as derived needles inside
Redactions::new (see the changelog and the ADR-0005 amendment); property- and
mutation-tested, pinned end-to-end against a fixture introspection route.
Not covered, on purpose: hashed/split/re-encrypted reflections (not needle-
matchable), double encodings (an unbounded tower; echo endpoints produce one
level). The research doc’s companion ideas — a redaction-verifying scan over a
finished run record, GitHub ::add-mask:: for captured secret-typed values,
RF-style secret-typed macro arguments — are enhancements, not part of the
defect, and await triage.
Corrections to the research set, so they are not re-litigated
- S4 (Trusted Publishing plan) rests on a false premise: it plans a first publish with a classic token, but all four crates have been live on crates.io since 0.5.1 (0.12.0 current). Trusted Publishing can be configured directly against the existing crates; the token sequence is unnecessary.
- S2’s exposure check is right and already satisfied:
Cargo.lockcarriescurl-sys 0.4.90+curl-8.21.0, past the June-2026 CVE batch. The detection blind spot (RUSTSEC carries no advisories for*-sys-bundled C libraries) is real; the proposed libcurl-version print in release artifacts awaits triage with the rest. - R3-16/R3-17 (browser and Android engines) are foreclosed, not deferred: proef is API-testing-with-hurl only — a standing decision, not a gap the research reopens. The M6 line in CLAUDE.md is architectural readiness, with nothing scheduled. The seam-hygiene half of R3-15 stands on its own merits and awaits triage like the rest of the registry.
- The round-14 review audited
214a39d(a pre-amend commit never pushed; what merged isc3ac752, differing by one deliberately-removed proptest seed), counted 464 tests where 462 exist, and credited #63 with the LSP corpus-holding change that shipped earlier — recorded here because review counts have now drifted by +2 for three consecutive rounds.
The R3 registry — triaged 2026-08-17
Triaged as a set against the PRD, the ADRs, and current industry practice, with each seam re-validated against the tree first. The v1 research document was confirmed absent (only v2 exists on disk), so items defined only there are one-line summaries with no spec — that fact drives several verdicts below.
Built:
- R3-1
--max-fail N(shipped with this triage). The convention is universal — Playwright--max-failures, pytest--maxfail, nextest--max-fail— with one shared semantics: stop after N failures, un-run tests report as not-run rather than passed. proef’s seams made it a CLI-only change: a sink wrapper counts suite-scenario failures (thephasefield keeps setup/teardown out of the count) and cancels the run token, which is the tested Ctrl-C drain path — in-flight batches finish, the rest record as skipped, teardown still runs on its own token, and the record is a complete cancelled run. That last part is free correctness:diff --fail-on-regressionalready refuses to certify a cancelled run, which is exactly right for a deliberately-partial one. - R3-4
difftakes a record path — shipped earlier (#65), with the research’s--baselineflag spelling declined as a second name for the same positional.
Build next (validated, in order):
- R3-2 a flakiness verdict — (shipped as
proef flaky). The 2026 pipeline is detect → quarantine → resolve, and proef already owned the middle step (@quarantineruns-but-does-not-gate);flakyis the missing detect, a fold over the recordsruns-diralready retains, so the history window is[run] keep-runsand no new state exists. Transition-counting separates flaky from broken (a mutation test proved the test suite could not initially tell that apart from a naive fail-rate — the F,F,P,P case now pins it), per-step attempt counts surface the pass-only-on-retry latent class, and a cancellation-skipped row is not evidence. No--checkgate, deliberately — its siblingfragmentshas one, but a flakiness verdict is advisory by nature and@quarantineowns the gating decision; the asymmetry is a choice, not an omission, and the thresholds become contract (and move toproef.toml) only if a gating mode ever exists. - R3-3 sharding, hash-mode only — (shipped as
--shard I/N). The measured stability argument held end to end: the mutation test swapped index-slicing back in and the insertion case (prepend, which shifts every position) caught it — the append case did not, which is itself the finding’s point. The assignment is frozen by literal-pinned tests; changing the hash is a breaking change to every sharded matrix. Filter→shard order pinned; an empty shard of a non-empty selection exits 0 with a note. - R3-6 JUnit attributes — (shipped, from the fresh spec the triage
required). The spec was written from what the two consumers actually parse,
at source level: GitLab’s docs enumerate testcase
classname/name/file/timeplus suite and roottime— and explicitly ignore the count attributes andtimestamp; Jenkins’SuiteResult.javareads suitename/package/id/time/timestampand caseclassname, and never readshostname. What shipped, and why:- Identity became
classname+name— Jenkins keys test history on the pair, GitLab’s MR widget diffs head against base by it, and the old singlenameembeddedfile:line, so an edit above a scenario re-identified every test below it (a fleet of “new” tests on both tools).classnamecarries the feature file,namethe scenario alone — unique per file by construction (outline instances are#N-disambiguated). Breaking for anything keyed on the old names. fileon the testcase (GitLab source linking),timeon suite and root (both consumers), and the suiteskippedcount spelledskipped(quick-junit 0.5 → 0.7; 0.5 wrotedisabled, which neither consumer reads).timestampandhostnamedeliberately absent — GitLab ignores both, Jenkins substitutes its own build clock and never readshostname, and naming the machine would undo R12-1. Additive later if a consumer asks.
- Identity became
Deferred, with the trigger named:
- R3-5 CTRF output — shipped as
--ctrf(#160, 2026-09-02). The deferral read “a seventh format needs a consumer, not a trend”, against the six proef already emits (JUnit, TAP, JSONL, a GH summary, SARIF, HTML). What it had not weighed is that the seventh shares the JUnit fold, so it cost a renderer rather than a mechanism, and ADR-0019 quarantine parity plus realretryAttemptscame with the fold. Corrected 2026-09-10 — it sat under a “deferred, trigger named” heading for the eight days after it shipped. R3-9, four bullets below in this same list, was annotated the moment it shipped — that is the convention this entry missed. - R3-18 generated pack documentation — when pack-vocabulary discovery becomes a reported adoption pain; the LSP currently serves that need interactively.
- R3-15 pre-M6 seam refactors — when a second engine is actually scheduled (M6 has nothing scheduled; the snapshot corpus already provides the golden artifact-diff prerequisite).
- R3-7
--affected-by, R3-10 fake variants — defined only in the absent v1 document; need the source or a fresh spec before any verdict. - R3-9 seeded shuffle — shipped as
--shuffle(RF-audit wave 1). The old pointer here was dangling: IMPROVEMENT-PLAN #14 is the fakes seed and never mentioned order. The shipped form honors #14’s actual rule anyway — the permutation is seeded by the run id, no parallel seed.
Declined — do not re-raise (moved to the standing section’s rules):
- OTel trace export (R3-11) and Cucumber Messages (R3-12). ADR-0008: the JSONL event stream is the record, no second record format. Both are re-encodings of the record for ecosystems that can convert from JSONL outside proef; building them in creates permanent format-tracking obligations against moving upstream schemas.
- Browser/Android engines (R3-16/R3-17) — foreclosed by the standing hurl-only decision, not deferred.
- S4’s first-publish token sequence — false premise; the crates have been live since 0.5.1. The worthwhile residue (crates.io Trusted Publishing for the existing crates, then the token-delete) is an owner-side dashboard action, recommended to the user rather than something the repo can do.
Open — adoption report on 0.12.0 (ingested 2026-08-14)
From a suite that ported to ref: at scale — 15 hurl files, 112 fragments, 21
scenarios — and ran 0.12.0 as an installed release. Three items, each reproduced
here against the tree before being written down. Two shipped in the same change;
the third is recorded because the report’s diagnosis was wrong even though its
observation was right, and that distinction is the finding.
R12-1 — provenance named the machine that produced the record (shipped)
[run] suite resolves against the config directory (R11-1), so a path-less
proef test handed the front end an absolute path and every emitter printed
it: the .hurl # source: header, .map.json’s feature.file, every
step_finished event, the console, and pack diagnostics. Two checkouts of one
suite stopped producing equal artifacts, which is exactly the property ADR-0010
exists to guarantee.
Worse than R9-6 filed it. R9-6 says the portability claim “holds only from the project root”; this reproduces from the project root with the config in it. R11-1 was the right fix — one resolution rule — but resolution produces absolute paths, and nothing was named at the other end.
Shipped, and R9-6 with it. front::SourceNaming is the one naming boundary:
resolve against the project, then name against the project again. A relative path
is left exactly as it arrived (machine-independent already, and the caller’s own
spelling, which their terminal can open); an absolute one is spelled relative to
the config directory when it lies inside it. This also replaced the fragment
corpus’s cwd-relative strip, which was a second anchor for the same question —
the drift R9-6 predicted. The four ways to name one suite (derived, typed, typed
absolute, from a subdirectory) now emit one artifact byte-for-byte, pinned by
crates/proef-cli/tests/provenance.rs.
Two limits, deliberate: a corpus genuinely outside the project keeps its absolute
name, because no project-relative one exists; and DiskSourceProvider
(proef lsp) still yields absolute names, because it keys document identity on
them.
R12-2 — the run-record ceiling was a constant no project could reach (shipped)
Retention was const RUN_RETENTION = 200 with only runs-dir configurable, and
artifacts are byte-identical across runs of an unchanged suite — so a suite
re-run on every save accumulated identical bytes for a day before anything
signalled a ceiling existed. [run] keep-runs makes the policy expressible; 0
keeps none but the run in flight.
The report’s inference that artifacts should therefore not be stored is wrong,
and it said so itself: an old record’s artifacts are what that run executed,
and once the corpus changes proef artifacts no longer reproduces them. Bound
the cost, do not drop the evidence.
Not closed by this, and not reported: rotation only ever deletes directories
named by a generated run id, so --run-id <name> records sit outside the
budget entirely. A CI minting a fresh id per build accumulates without bound.
Guessing at user-named directories is the worse failure — runs-dir may be .
— so this stays, documented in CONFIG.md rather than fixed.
R12-3 — a [run] setup test failure is invisible to JUnit (shipped)
Reproduced: a setup feature whose assertion fails exits 2 with
summary: 0 passed · 0 failed · 0 skipped, and --output junit writes an empty
report, because the abort precedes the reporter. A CI reading JUnit sees nothing
at all.
The exit code is not the defect. ADR-0014 decided it explicitly — a setup failure maps to a user (2) or system (3) fault, never a test failure, “the same distinction Playwright draws between a clear setup error and a cryptic test failure”. Changing it needs a superseding ADR, not a bug fix.
Three of the report’s supporting claims do not survive checking, recorded so they are not re-litigated:
- “teardown already has a distinct code; setup collapses both into one” —
false. A teardown assertion failure exits 3, not 1: both phases map a test
failure onto a non-test code (
phase_failed(…, UserError)/…, SystemError). Neither distinguishes, by design. - “appears in nothing
explain/diffconsume” — false forexplain, which printsfailed (setup — excluded from the totals above)with the assertion detail and the artifact reference; the events are in the record withphase: setup. - “previously raised, still open” — no entry in this file matches it.
So the open item is narrow: the phase reporters run only for the pool. Worth fixing at the reporter, not the exit code.
Shipped at exactly that boundary: the CI-report block (JUnit, GitHub job
summary, PR annotations) is one function both enders call, so a setup abort now
writes the reports from the setup phase’s own summary — one testcase, failed,
suite named by the setup feature file. On main the gap was worse than filed:
no JUnit file was written at all (the finding said “empty”). Exit codes are
untouched, per ADR-0014. Nothing is fabricated for the pool that never ran —
the test pins that too.
Open — round-7 residue (ingested 2026-08-10)
The round-7 pre-merge review of PR #13 never entered any worklist; a round-8
revalidation re-reproduced its findings against v0.8.0. §2.2, §2.3, §2.4 and the
diff item shipped in #30/#31. What remains, carried on that report’s evidence
rather than re-reproduced here:
-
The early-error record — reproduced 2026-08-11, needs a decision.
proef test --tags <nothing-matches>prints the error and then asummary: 0 passed · 0 failed · 0 skippedline, and the record it leaves isrun_started+run_finished 0/0/0— byte-indistinguishable from a clean run of an empty suite. A post-mortem reader cannot tell “errored before dispatch” from “ran nothing successfully”.The fix is a design call, not a patch. Suppressing the tail on this path would leave the record incomplete, which the tooling already banners correctly — but
RunRecordemits its tail structurally, onDrop, precisely so no return path has to remember it, and adding an exception reintroduces the fragility that design removed. Opening the record later is blocked by setup, whose scenario events need it. The third option is an additive event carrying the early error (ADR-0008 permits it) — the most honest and the most work.
Open — residue of the two UX reviews
Verified against main on 2026-08-10. Everything else those reviews raised has shipped
(first-run: F1, F3, F4a and F2’s did-you-mean in 0.6.0 · non-technical: N1–N5 and the
init count in #24).
R1 — missing_config_var’s span points at the sentence, not the pack line
The diagnostic reports at the feature step that used the variable, e.g.
suite/case.feature:3:5, rather than the pack line where ${url:bse} actually appears —
so the reader goes hunting. The did-you-mean half shipped in 0.6.0; this half did not,
deliberately.
Why it was deferred, in full — this is the whole reasoning, do not re-derive it:
ResolveError carries no position, and resolve() is documented “pure and total”.
The comparable diagnostic that does land on a pack line (pack::invalid_hurl) gets
its position from hurl’s own parser reporting a line/column, which feeds
locate::payload_line_span(…, rel_line); nothing computes a rel_line for a resolve
failure. Supplying one means threading an offset out of a deliberately position-free
pure function and carrying pack identity to the diagnostic site. That is a design
change, not a fix — it wants its own spec.
Two sibling extensions were declined at the same time: resolve::missing_env must
not suggest from the injected environment snapshot (it would surface unrelated
environment variable names in diagnostics, against the secret-masking posture), and
resolve::unknown_namespace already enumerates all seven valid namespaces. Sibling
codes share a shape, not a candidate set.
R3 — the scaffold default is the dev fixture’s port (declined 2026-08-11)
init.rs writes base = "${env:PROEF_BASE_URL:-http://127.0.0.1:8787}", which is
proef’s own dev fixture port — so to someone who installed a binary and has no
fixture, the value looks configured and is not. The proposal was an obvious
placeholder (https://api.example.com) to cover prevention, since a failing run
already covers recovery.
Declined, with the reasoning recorded rather than a silent skip. Recovery is now
covered on both halves: an unreachable target and untouched routes each get their
own note (#28, #38). The remaining benefit is that the config file would read as
obviously unfilled. Against that, init.rs’s module doc states the scaffold
deliberately mirrors what GETTING-STARTED teaches — so changing the literal changes
the tutorial too, and the tutorial’s “run it against xtask fixture with no
PROEF_BASE_URL” flow stops working. That flow is a real onboarding asset for
contributors. Trading a working tutorial for a more obviously-fake string is not worth
it once the failure itself explains both halves.
Revisit if first-run drop-off is ever measured rather than reasoned about.
Decided against — do not re-raise
Recorded as decisions, so they are not rediscovered as fresh ideas.
- Re-classify the unconfigured-scaffold failure from exit 3 to exit 2. Not a
CLI-edge change: the verdict is set in
proef-engine-hurl(classify_error’s_ => Infraarm),Fault::System(String)carries no kind to match on, and the exit derives inproef-core(RunSummary::exit_code_excluding). Both routes — string- matching the engine’s opaque message, or adding a structured kind to core’s public surface — cost more than the value, which is vocabulary. The note delivers that, and fires on the exit-1 placeholder-route path a re-classification would have missed. - Degrade
proef flowsthe waymacrosdegrades.flowspromises every scenario; a list silently omitting the feature that failed to parse is a wrong answer, not a degraded one.macrosdegrades safely only because pack loading precedes binding and does not depend on it. - Ship
proef-fixturein the binary so the scaffold’s first run passes. Needs a new ADR (it is dev-only today), enlarges the binary and the security posture of a test runner with a listening server — and R3 plus #24’s note remove the need. - A GUI, web UI, or “no-terminal” mode. PRD §3 forecloses dashboard/server mode. The P1 gap was always about vocabulary and error text, never a second interface.
- Importing or round-tripping hand-written hurl, and anything OpenAPI-shaped as a recurring oracle. PRD §3 and ADR-0016 permanent non-goals.
Open — correctness
Q2 was the remaining Tier 1 branch (Q5 and Q4 shipped in #26); it closed with the #146 analysis cache — see below.
Q2 — the walk still happens twice per request (closed 2026-09-02)
Shipped in #27: the walk skips target/, node_modules/, vendor/ and
dot-directories, is depth-bounded, and no longer aborts the whole discovery on
one unreadable subdirectory (which analyze.rs swallowed into a silently empty
analysis). Shipped in #32: the server adopts the workspace root the client
announces — workspaceFolders, else rootUri, else the previous
config-then-cwd resolution — so an editor launched outside the project no longer
analyses the wrong tree.
Closed 2026-09-02 — by the #146 analysis cache, which this entry predated.
The premise (“on every completion/definition/references request”) is no longer
true: every request handler reads one cached Analysis through the single
read path (server.rs — “the debounced diagnostics publisher and every
on-demand feature go through here; they share one recompute per edit rather
than one each”), edits mark the suite dirty behind a debounce, and the
fragment corpus is held across recomputes (“called when a fragment file
changes, never per request” — analysis.rs). The invalidation hook this entry
said SourceProvider lacked turned out not to be needed: the whole analysis
is invalidated on any edit, which at this suite scale (tens of small files,
milliseconds per recompute) beats maintaining an incremental index — the
module doc says so in as many words. The two walks inside one recompute
remain, and are now a per-edit cost too small to file.
P5 — watch: the atomic-save half (remainder)
Shipped in #37: --watch now also watches proef.toml, matched by exact path.
Closed by inspection — the inspection was invalidated by a later change, and
the bug shipped. The original argument was: the retrigger filter is an allowlist
of .feature/.yaml/.yml, and no run-record file (.jsonl, .log, .hurl,
.vars, .json, .xml, .html) matches it. ADR-0018 then added the engines’
fragment extensions to that allowlist — .hurl, named in this very paragraph as
the thing that could not match — while every run writes
.proef-runs/<id>/artifacts/*.hurl. A watched tree containing its own runs dir
fed itself: 49 runs in 15 seconds, firing real traffic in a tight loop.
Now closed by construction, not inspection. The retrigger filter excludes
generated trees by directory name, reusing discovery’s own skipped_dir, so
there is one rule with two consumers rather than a second list to drift; the
configured [run] runs-dir is passed in for the case where it is not a
dot-directory. watch::tests pins both halves — that an emitted artifact never
requeues, and that a fragment edit still does.
The lesson is the general one: a “closed by inspection” note records a conclusion whose premise nothing watches. This one even enumerated the fact that later became false. Prefer a test that would fail when the premise changes.
Still open: “a single watched file dies after an atomic save”. It did not
reproduce on macOS/FSEvents; notify’s own docs say it is real but
platform-dependent and worst on inotify. Do not chase it on a Mac — that is how it
gets “fixed” by coincidence. It needs a Linux reproduction first.
Open — adoption and execution model (ingested 2026-08-11)
Source: a report written while porting a real 844-line hurl corpus onto proef —
field evidence rather than inspection, which is why it found a different class
from the review rounds. Every claim below was re-checked against main before
filing; where the report was wrong, the correction is recorded with the item.
Already closed from it: the docstring-placeholder documentation gap (#41). Two of its claims did not survive checking, and both are noted in place (M1, D2).
The through-line. These are adoption, not correctness. The first-run path is finished and the correctness series closed its bug class; the next constraint is whether a team with an existing hurl suite can move onto proef and demonstrate they lost nothing. M1 and M2 are that story. E1 is the first wall a real suite hits afterwards.
F1 — proef.toml now has two path-resolution rules (closed — duplicate of shipped R11-1)
[run] fragments resolves relative to the config file’s directory (ADR-0018);
suite, setup, teardown and runs-dir stay relative to the working
directory. The reasoning that produced the new rule — the config is found by
walking up, so a path in a config three levels above must mean “relative to the
project” — applies verbatim to all five keys, and setup/teardown/runs-dir
are consulted on every run rather than only when a path was omitted.
Cost: one file with two semantics and no marker distinguishing them. A user with
setup and fragments in the same proef.toml gets one working from a
subdirectory and one not, and every future path key re-litigates the choice
against four precedents for the older rule.
Not fixed here on purpose. Changing the four existing keys is a behaviour
change for every project that already relies on cwd-relative resolution, which
is out of scope for the change that introduced the fifth. The fix is a single
ProjectConfig::resolve_path used by every path accessor, shipped deliberately
with a changelog note — recorded so it is a decision rather than an oversight.
ADR-0018 (named hurl fragments) lands into this section — read it against these
items before assuming what it closes. It lets a pack ref: a named entry in a real
.hurl file, so a corpus file is annotated once instead of transcribed, and stays
runnable under stock hurl. Item by item:
- M1 is not closed and must not be built concurrently — both touch
fmtdiscovery. ADR-0018 requires the opposite of M1 at one entry point (directory discovery must never sweep.hurlinto the pack formatter) while leaving M1’s actual ask untouched (an explicitly named.hurlmay be canonicalized). Sequence them, either order, never at once. - M2 is not closed. ADR-0018’s integration test runs one fragment both ways, which proves a file is dual-runnable; it does not compare two suites’ result sets.
- M3 is unanswered and now overtaken: the charter re-examination M3 asked for has happened (PRD §3 amendment) without the measurement it asked it to rest on. The amendment argues from the non-goal’s own rationale instead, and says so. Measuring the port cost is still worth doing — it now informs priority rather than permission.
Closed 2026-08-23 (premise false — the “not fixed here” above HAS since been
fixed, as shipped R11-1). The exact fix this entry prescribed exists as
ProjectConfig::resolve (config.rs:328-338): every path-valued key routes
through it, its doc comment narrates this entry’s story, and CONFIG.md
documents the one rule. This entry and R11-1 were the same finding filed twice.
M1 — fmt cannot canonicalize a standalone .hurl (closed — foreclosed by ADR-0018)
The report had this backwards and it is worth recording why. It claimed fmt
refuses a file outside a pack, and proposed teaching it to accept .hurl as a
small plumbing change. fmt in fact accepted any file and rewrote it — two
defects fixed in #40, which now makes it refuse .hurl correctly, since
applying YAML block-location logic to hurl syntax would be nonsense.
So the item survives but changes shape: making it real means teaching fmt to
recognize a hurl file and run the block canonicaliser over the whole thing, with
no hurl: key to locate. That is a feature, not a flag.
Why it still ranks first. It is what converts M2 from clerical to mechanical, and it is the cheapest unlock for the most valuable capability.
Closed 2026-08-23, without building it. This entry predates ADR-0018, which
was accepted with the opposite principle: proef reads files it does not own,
so it must never write them — fmt refuses fragment files (ADR-0018,
“proef never writes”; carried as a hard constraint in CLAUDE.md). Building M1
would diverge from an accepted ADR without a superseding one. And the goal M1
served no longer needs it: it existed to make M2 mechanical — canonicalize both
corpora, diff the text — but fragments removed the transcription M2 was
guarding, so there is no ported copy whose equivalence needs proving. The file
the backend team owns is what proef runs, pinned per-file by the both-runners
test. Reopening this requires a superseding ADR, not a feature request.
M2 — no mechanical equivalence check between a hurl corpus and its proef port (deferred — trigger named below)
Verified when filed (diff now also accepts record dirs and .jsonl paths — R3-4/#65 — but still reads no hurl report); no
path reads a hurl --report-json, which the pinned hurl 8.0.1 does emit.
Why it matters. The safe way to adopt proef is to run both suites until the new one is trusted. During that window nothing proves the two assert the same things, so the equivalence gate degrades to a hand-maintained mapping table reviewed once by a human — and that table is what a team’s decision to delete their old suite rests on.
Scope. Not the hurl-import non-goal in disguise (PRD.md:42). Import means
reading .hurl and generating Gherkin. This compares two result sets,
which is diff’s existing job with one more input format. The non-goal
forecloses a direction of data flow, not the ability to check your own work.
Deferred 2026-08-23. The urgency rested on transcription drift — a port
that could silently assert less than its original. ADR-0018 removed the
transcription: a migrating team annotates the corpus it already has, and the
same bytes run under stock hurl and under proef (fragments.rs pins it
per-file against the fixture). What remains defensible is a results diff for
the trust-building window when both runners run in CI side by side —
diff’s job with hurl --report-json as one more input. Trigger: the
first concrete migration that runs both runners and asks to compare outcomes
mechanically. Building a seventh input format ahead of a consumer is the same
mistake the CTRF deferral records.
M3 — the port cost has never been measured (closed — overtaken by ADR-0018)
PRD.md:42 makes hurl import a permanent non-goal, and that rests on
persona P3’s “pastes between corpus and packs” (PRD.md:57) being cheap —
which nobody has measured. A 14-file, 844-line port is the first real datum
available. Recording the hours settles a recurring argument in one direction or
the other: cheap vindicates the non-goal with evidence instead of assertion,
expensive earns the charter a re-examination with numbers rather than opinion.
Closed 2026-08-23. The re-examination this measurement was meant to trigger happened: ADR-0018 narrowed the non-goal to generation and rewrote P3’s job from “pastes between corpus and packs” to “annotates once” — the exact charter change M3 said the numbers should decide. The two field data points stand recorded (an 844-line/14-file corpus ported by raw paste at 100% coverage; a 97-entry corpus that chose annotation and stopped the paste port deliberately), and no third answer would change a decision that has already been made and shipped.
E1 — no intra-run serialization primitive (report B1)
Verified. TECH-SPEC.md:313 — scenario ordering is “preserved for artifact
naming, not execution order.” No serial tag or config key exists anywhere in
core, cli, CONFIG.md or AUTHORING.md.
Why it matters. Real suites contain scenarios that mutate global state — the
reporting corpus has two, one needing an empty database for absolute items[N]
assertions and one installing a workflow definition governing everything created
afterwards. Neither can run in a parallel pool, and proef offers no way to say
so; the workaround is several CLI invocations driven by tag discipline in a
Makefile.
Charter fit. Scheduling, not a new engine or execution mode — the
orchestrator already decides what runs when, and [run] setup/teardown prove
the surrounding concept is in charter. Those cover before and after the pool
and nothing inside it.
Options. A reserved @serial tag, or [run] serial-tags = [...]. The config
form is more explicit and keeps runner semantics out of the feature files — and
E4 is an argument for it.
Shipped as [run] exclusive-tags, a tag expression rather than a list —
the same language --tags takes, so group membership is answered exactly as
selection is. Two corrections to this entry, both from checking before building:
- The filing describes one axis; the mature shape has two.
cargo-nextestseparates a group concurrency limit (max-threads, which bounds members against each other and leaves the rest of the pool running) from per-test weight (threads-required, which is what buys global exclusivity — they redefined it in 2024 precisely so limits “are never exceeded”, enabling mutual exclusion against all tests). Only the second is what was missing here, so only that shipped; a group table can be added later without breaking this key. - Of the two motivating scenarios, only the first is a serialization problem.
“Installs a workflow definition governing everything created afterwards” is
ordering, which
[run] setupalready provides — a feature run once before the pool exists. Recorded so an ordering primitive is not built on the assumption that it was needed.
E2 — N invocations produce N run records, with no merge (report B2; consequence of E1)
Verified. Each run writes its own .proef-runs/<run-id>/ (TECH-SPEC.md:299).
E1’s workaround therefore yields N records, N JUnit files, N HTML reports, and
pass/fail aggregation pushed onto the caller’s shell, while explain/diff
operate per-run so a post-mortem reader must know which to open. Recorded as a
consequence, not an independent item — solve E1 and this largely evaporates;
solving it alone (a proef merge) treats the symptom.
Largely closed by E1 shipping: a suite whose isolation needs are expressed
as exclusive-tags runs in one invocation, so it produces one record, one JUnit
file, one report and one exit code. Kept open rather than closed outright
because a suite may still split invocations for reasons E1 does not address
(different environments, different --tags in separate CI jobs), and nothing
merges those.
Shipped (2026-08-25) — the rerun half: run_started.rerun_of names the
base; the rerun’s JUnit carries the base’s not-re-run scenarios
(reconstructed from its record, exit code and totals untouched), and
report overlays the base into a whole-suite page with a merged-view
banner, degrading loudly when rotation ate the base. What remains of E2 is
the original split-invocation case (different --tags in separate CI
jobs), still open on its trigger.
Widened by the RF audit (2026-08-24): the class includes --rerun’s own
CI story, which this entry never named — a rerun writes a new record whose
JUnit/report contain only the re-run subset, so “the one JUnit at the end”
of the standard retry workflow describes 3 scenarios of a 300-scenario
suite. RF’s answer is rebot --merge. The proef shape, when built: overlay
a rerun record onto its base at report/JUnit emission — composition over
records, never a merged record file (ADR-0008); an additive
RunStarted.rerun_of field would make records self-describing for it.
E3 — no per-scenario state reset hook (report B3)
Verified. [run] setup/teardown are whole-suite only, run once around the
pool (CONFIG.md:63-64, 120-141).
Any suite against a real database wants before-each; today isolation is
convention (title prefixes so scenarios do not see each other’s rows) and
convention has no guardrail. proef knows nothing about databases, so “reset the
DB” cannot be a proef feature — but framed as a feature file run before each
scenario it is the same primitive as setup at a different scope, which is
engine-agnostic by construction. The cost is real: it multiplies run time by
scenario count and interacts with parallelism. This needs an ADR against
ADR-0014, not a patch, and it may well be declined — deliberately rather than
never asked.
E4 — nothing enforces tag-group discipline (report B4; record, do not build)
If E1 ships as a tag convention, a scenario added six months later lands untagged in the parallel pool and breaks isolation intermittently — the worst failure mode, because it reads as flakiness. A lint would have to guess which endpoints are global, which proef cannot know. Its value is as a marker: this is the follow-on cost of the tag form of E1, and therefore an argument for the config form.
D1 — no first-class requirement traceability
Verified. flows --format json prints one object per scenario
(main.rs:128-137), which with tags like @FRD-3.1-create gets most of the way.
Almost certainly a documented recipe rather than a feature — proef should not
learn what a requirement is — but the recipe does not exist, so every team
reinvents it and the capability is not advertised for this use.
D2 — report generation across N runs (premise partly corrected)
The report overstated this. It claimed a Makefile must capture the run id
because proef needs proef report <run-id>; in fact run_id is optional and
defaults to the latest run (main.rs:199-201), so the ordinary single-run case
needs nothing captured.
What survives is the compounding with E2: with N invocations, “the latest” is one of N. Minor on its own, and listed because report-generation friction is felt by every CI integration rather than by one team.
Positive evidence — recorded so it is not undone
- The raw-hurl paste path covered 100% of a real corpus. All 844 lines used
only
[Asserts](75) and[Captures](15) — no[Options],[Query],[FormParams]or[Cookies]— with seven ordinary predicates (==,exists,not exists,matches,count ==,>=,isString), every one passing through untouched. The strongest evidence yet for ADR-0004, and the kind of claim that gets doubted later. proef macrosprinting sentences (#29) is load-bearing. The porting plan gated its prerequisite phase on it, purely to author 14 files of new prose.--rerun(main.rs:120-122, re-run only the last run’s failures) fits conversion iteration exactly.
Suggested order (historical — every item now resolved or parked)
The order was M1 → M2 (adoption becomes provable) → E1 (dissolves E2). E1
shipped as [run] exclusive-tags; M1 closed against ADR-0018; M2 is deferred
on a named trigger; M3 closed as overtaken. The two documentation items, C1
and C3, shipped in #43. E3, E4, D1 and D2 remain record-only — none blocks
anyone today.
Closed — docs drift (2026-08-11)
Every item in this section shipped; the table above records which PR each landed in. Two did not reproduce when re-checked, and are recorded here rather than dropped, so the next reader does not spend the same time on them:
- A2 —
CONFIG.mdwas said to claim[env.<name>.run]overrides any section. It carries no such claim today: its precedence text namesjobsspecifically, which is whatRunOverrideactually allows. - B12 — the CHANGELOG’s 0.5.2 entry was said to lack a line about the directory-valued-phase hard error. It has one, first bullet under Fixed.
One half of A5 was deliberately not acted on: TECH-SPEC §11’s run-dir inventory
lists the files a run generates, and the [run] setup/teardown features are inputs
named by config, not run-dir output. The reviewer called this half “defensible-but-
interpretive” and it is; report.html, which the inventory genuinely omitted, was added.
Open — maintainability and CI
| ID | Finding |
|---|---|
| B10 | The canary would chase a hurl prerelease (no semver filter) — shipped: the index parse (latest_stable_in_index) skips - versions, unit-pinned; build metadata needs no rule, crates.io refuses versions differing only by +meta |
| P12 | The matcher re-tokenizes per (step, pattern) pair on every bind (performance) |
| P13 | no_guarded_secret_ever_enters_the_global_store); no CI workflow runs llvm-cov — the local half shipped 2026-09-07 (#173): just cover/cover-html/cover-lcov; the CI job stays a maintainer’s cadence/cost call and must be a ratchet, never a threshold (TESTING-STRATEGY §3) |
| Q1 | EngineLowering was a review’s name, never a symbol) — what survives: no registered engine claims a structured kind, so the path runs only under test fixtures |
| Q6 | html.rs re-derives the emitter slug; four file_stem() sitesemit::feature_stem and emit::artifact_slug are now the one definition of each, called by the emitter’s own caller, the dispatcher’s spec naming, the report’s anchors/artifact links, and the editor analysis. The other premise had gone stale the other way: ScenarioOutcome.artifact_slug has carried the emitter’s naming to runtime consumers since round 19, so “the schema carries no slug” no longer forced anyone to re-derive) |
Open — deferred during the v0.6.0–v0.8.0 correctness series
Found while fixing the above; each was validated and consciously left out of scope.
-
proef-harnessPROEF_BIN/PROEF_HARNESS_SUITE— fixed in #19, but the same reader is now duplicated inproef-cliandproef-harness. Justified today (a binary crate cannot be depended on; these are the only twoenv::varcallers in the tree). Tripwire: at a third caller, promote it to a shared crate. -
Capture-name charset is narrower than hurl’s grammar, so an out-of-charset name is silently omitted fromClosed 2026-09-11: aligned with.map.json.hurl_core’skey_string_text— anychar::is_alphanumeric(Unicode, not ASCII) plus_ - . [ ] @ $. Souser.id,items[0],@type,total$andprécisall parse as captures in hurl and were all absent from the sidecar. A leading[stays refused because hurl refuses it too;{/}stay out because a templated name has no statically knowable text. -
A
#comment inside a[Captures]run — fixed in #16;the one/two-letter-method gap it exposed remains (. Closed 2026-09-11, and the measurement found a second error in the opposite direction: the predicate also allowedis_method_linerequires three characters, hurl’s grammar does not)-, whichhurl_core’smethod(read_while(is_ascii_alphabetic), non-empty, uppercase) does not. Too narrow on length and too wide on charset, each masking the other, which is how both survived from 0.1.0. The failure is a phantom row, not only a missing one: with the run left open across a short method, a header of the next entry reaches.map.jsonas a capture nobody wrote — the first version of the regression test missed exactly this, because a response line closed the run anyway and it passed against the defect. -
Closed 2026-09-11, one step past the prescription. A flag can be ignored; the scan is instead private behind akey_line_spans’ flow-style undercount is guarded by convention, not types. Two callers guard it independently; a third would have to remember. Cheap hardening: have the primitive return a reliability flag.KeyLinesvalue whose only accessors arepaired_with(parsed)— the spans, and only when the counts agree — andsole()for a key that occurs at most once. There is no path to a positional list that does not state the count it expects, so the third caller has nothing to remember. Both existing guards became the call itself, andspans_reliableis gone. -
Cross-scenario
${fake:*}coincidence — two scenarios can still draw the same value. Documented as a known limitation in AUTHORING/CHANGELOG/TECH-SPEC. -
No corpus tier for engineered robustness fixtures.
tests/has zero custom-method entries and zero fenced blocks, so that bug class is pinned only by unit tests on private functions. -
fmt’s tie-break (equal CRLF/LF → LF) now applies only to the trailing newline of a file that lacked one — per-line endings are preserved (#33). Lone-\rfiles are still unhandled: the splitter keys on\n, so a classic-Mac file is one long line. -
normalize_packkeeps the skeleton verbatim by construction at eachpush, not by the algorithm’s shape. The “hurl blocks only” promise has broken three times (#18 line endings, #33 mixed endings, #40 trailing whitespace), each caught by an example pinning that one instance. #44 added properties — skeleton-only text round-trips byte-for-byte, and formatting is a fixed point — so a fourth over-reach now fails CI instead of shipping. The structural version would locate each block’s byte span and splice the canonicalized body back into the original text, making “bytes outside a span are never visited” a property of the shape. Not worth the rewrite for a small textual formatter; revisit if a fourth normalization rule is ever added to that loop. -
The stdout latch’s single-reader test isolation is safe under the mandated nextest (one process per test) but is a convention, not an enforced invariant.
-
A disk filling mid-run still truncates the human console report without reaching the exit code(closed: the console latch shipped in the 2026-09-02 series (#160), to its own written design; the record’s own writer got the same latch in #168, and both reach exit 3 throughescalate_environment_failures). -
Absent-secret fallthrough (“an unsetClosed 2026-09-11:PROEF_SECRET_<NAME>still reads the store”) is load-bearing and pinned only by an integration test, not a unit test.resolve_allcarries unit tests for the fallthrough, for the override winning over a stored value, and for the neither-source error naming both remedies — each checked against a mutation that breaks it. ThePROEF_KEYoverride supplies the key, so nothing touches a key file.Recorded because it cost a rewrite: the override test first claimed to prove the
from_store.is_empty()early return by using a corrupt store, and deleting that return left the test green.load_store’s error reaches the caller only through names that needed the store, and a fully env-supplied run has none — so the early return is an IO saving, not an observable behaviour, and the test’s stated mechanism was not the one making it pass. Kept as a separate test that says so.
Versioning & release procedure
This document is the versioning policy and the release runbook. The README carries a summary; this file wins on detail.
Versioning policy
Scheme: SemVer 2.0.0. Pre-1.0 semantics, applied strictly:
- MINOR (0.X.0) — any breaking change, or a coherent feature series/milestone.
- PATCH (0.x.Y) — fixes and purely additive changes that break nothing below.
What counts as breaking (these are the public contracts, per the ADRs):
| Surface | Breaking examples | Non-breaking examples |
|---|---|---|
| CLI + exit codes (ADR-0009) | removing/renaming a flag; changing an exit-code meaning | new flag; new subcommand |
| Pack schema (ADR-0004) | removing a key; changing key semantics | new optional key |
| Event wire schema (ADR-0008) | removing/renaming a field or variant; changing schema semantics | new variant; new field with a default (additive-only rule) |
| Canonical artifact format (ADR-0010) | any change to emitted bytes (snapshot-locked) | — (changes are inherently breaking; bump minor) |
| Engine seam (ADR-0002) | changing EngineFactory/EngineSession/StepBatch/ScenarioCtx shapes | new defaulted trait method |
| Config file | removing/renaming a proef.toml key | new optional key |
1.0.0 is declared when the pack schema, CLI grammar, event schema, and exit codes are stable enough to promise MAJOR-only breakage. Until then, downstream consumers should pin minor versions.
Single source of truth: [workspace.package] version in the root Cargo.toml.
Every crate inherits it (version.workspace = true); the workspace releases as one
set, always. Never version a crate individually.
Orthogonal versions, not to confuse with the crate version:
- The event schema version is the
schemafield inrun_started(EVENT_SCHEMA_VERSION). It only moves on a semantic break of the stream — additive variants/fields do not bump it. - The hurl pins (
=8.0.1) never move as a side effect of a release. Upgrades go exclusively through the canary + runbook (IMPLEMENTATION-PLAN §7, ADR-0003). - Toolchain policy: the pin tracks latest stable Rust but adopts a new
minor only at its
x.y.1point release, ~3-4 weeks afterx.y.0(tools and third-party crates track latest immediately; exact pins like hurl outrank everything). A reviewer reading “latest stable” as “bump on release day” prompted writing this down (R18-2). - MSRV is the toolchain pinned in
rust-toolchain.toml; it may rise in any MINOR release pre-1.0 and is not a separate contract yet.
Tags: annotated vX.Y.Z on main, linear history. Cadence: release when a
milestone or a coherent series lands — not on a calendar.
CHANGELOG rules
## [Unreleased]always exists at the top; every landed change adds a line there in the same commit series that lands it.- On release,
Unreleasedcontent moves under## [X.Y.Z] - YYYY-MM-DD(with a short parenthetical theme) and a fresh emptyUnreleasedis left behind.
Release runbook
From a clean, green main (all gates local + CI).
mainis protected — the release commit goes through a pull request, and the tag is pushed only after it merges. Do notgit push origin main, and do not tag before the merge.git push --follow-tagsis not atomic: git pushes refs independently, so a protected-branch rejection stops the branch while the tag still lands — and a tag is exactly whatrelease.ymltriggers on. That combination starts a release build from a commit that is not onmain. It happened cutting 0.10.0; the run was cancelled and the tag deleted before anything published, but the recovery is avoidable and this ordering avoids it.
# 1. On a release branch, cut the changelog: move [Unreleased] → [X.Y.Z] - date
# with a short parenthetical theme, and leave a fresh empty [Unreleased].
# (There is no link-reference section at the bottom of CHANGELOG.md — nothing
# to update there.)
git switch -c release/vX.Y.Z
# 2. Bump the version in the root Cargo.toml — BOTH places:
# [workspace.package] version = "X.Y.Z" (the crates' own version)
# [workspace.dependencies] proef-core / proef-engine-hurl / proef-lsp
# version = "X.Y.Z"
# (the inter-crate pins — belt-and-suspenders for independent crates.io
# publish; a stale pin no longer satisfies the bumped version and fails
# resolution, so these move in lockstep with the line above).
cargo build --workspace # refreshes Cargo.lock versions
# fuzz/ is a separate workspace with its own committed lock, and the gates job
# checks it with --locked: refresh it too or that gate goes red on the release
# commit.
cargo check --manifest-path fuzz/Cargo.toml --all-targets
# 3. Full gates — the same set CI runs, so a green local pass predicts a green PR:
cargo nextest run && cargo test --doc
cargo clippy --all-targets --all-features -- -D warnings && cargo fmt --all --check
RUSTDOCFLAGS="-D warnings" cargo doc --no-deps --all-features --workspace
cargo deny check && cargo audit && cargo machete
cargo run -p xtask -- docs-check # the gate a docs-touching release commit trips
zizmor .github/workflows/
# 4. Commit and open the release PR (no tag yet):
git commit -am "release: vX.Y.Z"
git push -u origin release/vX.Y.Z
gh pr create --base main --title "release: vX.Y.Z"
# 5. After CI is green and the PR is MERGED, tag the *merged* commit and push
# only the tag. The squash merge creates a new commit, so tagging the branch
# would leave the tag off `main`'s history.
git switch main && git pull --ff-only
git describe --tags --exact-match HEAD 2>/dev/null && echo "already tagged — stop"
git tag -a vX.Y.Z -m "proef X.Y.Z"
git push origin vX.Y.Z # this, and only this, starts release.yml
If the release commit was made on main locally before branching, git pull --ff-only refuses afterwards: the squash merge superseded it. Confirm the merged
commit carries the version bump, check git diff --quiet HEAD origin/main, then
git reset --hard origin/main.
The tag push triggers .github/workflows/release.yml, which:
- builds release binaries for five targets (macOS arm64/x86_64, Linux
arm64/x86_64-gnu, Windows x86_64-msvc — the Windows zip bundles the vcpkg
DLLs; macOS links the SDK’s system libxml2 and vendors OpenSSL, so shipped
binaries need no Homebrew), via
cargo auditable(binaries stay scannable) with no cache restore (cache poisoning must not reach published artifacts), attesting SLSA build provenance per artifact (the repo is public, so this runs unconditionally); - publishes the GitHub Release with the version’s CHANGELOG section and all
five archives (asset names must stay in sync with the
binstallmetadata in theproefpackage manifest); - regenerates
Formula/proef.rbin theemrecdr/homebrew-proeftap (deploy-key auth via theHOMEBREW_TAP_DEPLOY_KEYrepo secret) — only when the tag is newer than the version the tap already carries. That step is gated on nothing but “a tag was pushed” and rewrites the formula whole, so a tag pushed late or out of order would downgrade everybrew upgrade; it now skips green instead, leaving the tap alone while the release still publishes. The formula installs the binary, its man page and the bash/zsh/fish completions; from 0.16.0 until 0.18.0 it installed only the binary, so Homebrew users silently got neither the man page nor completion while every other channel did. The render step now checks the archive for each file the formula claims to install, because nothing else connects the two.
A tag runs the workflow as it existed at the tagged commit, not as it exists
on main — the ordinary push-event rule, and the one that decides what
backfilling a missing tag actually does. Patching release.yml therefore
protects future tags only: a tag cut on a commit older than a fix runs the
pipeline without it. This is not theoretical here. Backfilling v0.15.0
(commit dated 2026-08-25) would have run that commit’s unguarded tap job and
walked the published formula from 0.17.0 back to 0.15.0, defeating the
forward-only guard added in 0.18 — which lives on later commits and could not
apply. The backfill was done with gh workflow disable release.yml around the
push for exactly that reason, then the Release created by hand. Disable the
workflow before pushing any tag whose commit predates a release-pipeline fix.
workflow_dispatch runs build+attest only — a full matrix smoke without
publishing. crates.io publication remains a deliberate manual cargo publish
per crate in dependency order (core → engine-hurl → lsp → proef — proef-lsp
before proef, which depends on it non-optionally) and is not automated.
Manual because it is the one step nothing undoes: a published version can be yanked, never replaced or re-uploaded. So publish from the tag, not from a working tree that merely resembles it:
git describe --tags --exact-match HEAD # must print vX.Y.Z
git status --short # must be empty
cargo publish -p proef-core --dry-run --locked
cargo publish -p proef-core --locked
cargo publish -p proef-engine-hurl --locked
cargo publish -p proef-lsp --locked
cargo publish -p proef --locked
--locked throughout, so what ships is what the committed lockfile resolves.
Only these four go: [workspace.package] publish = false is the default and each
publishable crate overrides it, so proef-fixture, proef-harness and xtask
are excluded by construction rather than by remembering to skip them. Each
command waits for the registry before returning, which is what makes the next
one resolvable.
The registry does not carry every tag. 0.15.0, 0.16.0 and 0.17.0 were tagged and
released on GitHub but never published, so crates.io goes 0.14.0 → 0.18.0
(published 2026-09-09, from the tag, all four crates). Cargo resolves version
requirements rather than sequences, so the gap costs a consumer nothing — it is
recorded here so that a reader comparing git tag against the registry does not
read it as a failed upload.
History
v0.1.0— initial release (fresh history baseline, 2026-07-29)v0.2.0— deep-review correctness blockers (duplicate-request, body corruption, delay budget), panic containment, the output contract, and the author guides — breaking:--output jsonstream split, empty selections exit 2, event schema grew additivelyv0.2.1— review P0 (header grammar, pipe, filters, name dedup, secrets perms) + failure UX (hurl expected/actual, true error-line anchoring)v0.3.0— data-safety blockers (asset copy, run rotation, zero-entry false green), Then-step visibility with exact attribution, UserInput taxonomy (user mistakes exit 2), option caps + repeat budget, atomic locked stores — breaking: proef-core API pruned,when:skips on literal false, zero-entry packs fail validationv0.3.1— secret hardening:secret rm,PROEF_KEYCI override, the saveAs-vs-secret promotion guard, doctor store/key health, corrupt-store recovery, warned-step reasons on the consolev0.4.0— external config & environments (proef.toml[url]/[vars]/[env.<name>],${url:}/${vars:},--env/PROEF_ENV, ADR-0012), default suite path, and the competitive-review breadth pass — breaking: the pack root keytemplates:becamemacros:with no alias (ADR-0004 amendment)v0.5.0— theproef-lsplanguage server: diagnostics, completion, go-to-definition and references over the sans-IO core (ADR-0017)v0.5.1— LSP correctness: process-leak, malformed-request crash, broken-pack degradation, root-at-suite, overlay keying;use:/match:go-to-definitionv0.5.2— CLI correctness: diff step-collision, truncated-run gate, setup double-run, the first EPIPE guard, overflow hardening, bare-filename path resolution, exit-130 documentationv0.5.3— closed-pipe safety: every remaining raweprintln!inproef-clirouted through the EPIPE-safe guard (with a source-scanning drift test), andproef-lsp’s panic-recovery notice no longer kills the server it just rescuedv0.6.0— first-run UX & run-record correctness:proef init, a did-you-mean for unset config variables, a next-command nudge; onerun_started/run_finishedpair per record with suite-only totals, truncated-record banners inreport/explain, a real worker slot index — breaking: a scenario with no steps is now an errorv0.7.0— record & artifact integrity:run_finishedis the record’s last line again (a watchdog-abandoned scenario no longer appends past it),${fake:…}values no longer repeat across a scenario’s steps,.map.jsonstops listing captures that were never made and stops dropping real ones, and a whitespace-onlyexpect:is rejected instead of emitting an inverted span — breaking:proef_core::resolve::resolvetakes a caller-owned occurrence counter andResolution::fakesis gonev0.8.0— CLI output & exit integrity: an unreadablePROEF_KEY/PROEF_ENV/PROEF_SECRET_<NAME>is a loud user error instead of reading as unset, therun.logtee no longer duplicates bytes on a short write,proef fmtkeeps a file’s own line endings,report -owrites artifact links that resolve, anddiffstops inventing flakiness for a step with no baseline — breaking: a failed stdout write now exits 3 where it exited 0, and a malformed environment variable exits 2 where it was silently ignoredv0.9.0— tool-surface integrity & authoring guidance: values interpolated into LSP snippets, GitHub annotations and job-summary tables are escaped,proef fmtrefuses a file that is not a pack and stops trimming the YAML skeleton,--sarifcarriesstartLineso annotations land,--watchretriggers onproef.toml,--dry-run’s nudge echoes the run that was actually validated, a templatedretry:stops under-counting the batch budget,--output jsonreports the real exit, a truncated record counts its warned scenarios, a failing run says when the scaffold’s routes are still placeholders,macrosprints the sentence an author needs,proef lspadopts the client’s workspace root, and AUTHORING documents docstring placeholders and the validation-catalogue pattern — breaking:proef secret set --valuewas removed in favour of--stdin(a secret in argv is visible tops), andproef macros --output json’spatternfield changed from a boolean tostring|nullv0.10.0— named hurl fragments (ADR-0018): a step mayref:one# @proef <name>entry of a real.hurlfile, values supplied bybind:at pack/macro/step scope, so the same bytes run under stockhurland under proef;[run] fragmentsnames the scanned root, aref:step records the fragment it ran asfile.hurl#nameeverywhere a failure is reported, and the editor completesbind:keys and jumps fromref:to the annotation — breaking:pack::loadtakes a&FragmentCorpus,PackSet::fragmentsis anArc,LoweredScenario::secretsis a map, andLoweredStep/StepOutcome/Event::StepFinishedcarryfragmentv0.11.0— the adoption response: ADR-0007’s value caps reach fragment text (byte-identical[Options]exited 2 inline and 0 behind aref:, then ran),proef fragmentslists the corpus and names both ways a fragment dies with a--checkgate, abind:key nothing reads is refused with did-you-mean,doctorreports the corpus,initscaffolds both body forms,--confignames theproef.tomlto read, and[run] exclusive-tagsruns a scenario with the pool to itself — breaking:FragmentScannerreturnsScannedFile,AnalyzeCtxtakes the corpus rather than building one per call,StepKindSpeccarries anoptionsrecogniser, andScenarioSpeccarriesexclusivev0.11.1— the gaps 0.11.0 shipped with:--configreachesproef lspand--watch(it was honoured by the runner alone, so the editor reported everyref:as unknown in exactly the layout the flag exists for),proef fragmentscounts[run] setup/teardownusage instead of calling a phase-only fragment unreachable and failing--check, one predicate answers “is this a fragment file?” where three disagreed, and--junit/--sarif/report -ocreate the directories their paths name — asartifacts -oand the run directory already did, and as pytest, jest-junit, cargo-nextest and the embedded hurl all dov0.12.0— one path rule, and a watcher that stops lying: a path written inproef.tomlresolves against the config, a path typed on the command line against the working directory, with no exceptions — which finally inventoried.proef-state.json,.proef-secrets.jsonand the run records, all three cwd-anchored and unlisted.--watchrereads the config it retriggers on; stops feeding itself whenruns-dirchanges mid-loop (one edit produced 39 runs in 12 seconds against a live API); and matches a relatively-typed or symlinked--config, which also restoredproef lsp --configgo-to-definition across the fragment corpus.doctorfails aproef.tomlthat will not parse instead of reporting on invented defaults,--configis honoured or refused by every subcommand, and[run] exclusive-tagsvalidates itself in both paths — breaking: the secret store, the World and the run records move with the config rather than the shell, which reaches anyone who ran proef from a subdirectoryv0.13.0— a record that travels, and a secret that stays one: nothing proef records names the machine that produced it (one naming boundary, the dual of the path rule — breaking: artifact bytes change for path-less runs), and a secret reflected base64/hex/percent/JSON-escape-encoded is redacted like its raw form (live leak reproduced, then closed; ADR-0005 amended). The fragment corpus read is bounded (601 MB → 15 MB on the measured pathological input), fuzzing actually reaches the fragment rules (probe-verified),[run] keep-runsmakes retention expressible,difftakes a record path for the CI-baseline flow, the bundled libcurl gets a CVE floor no advisory scanner would catch, and a hung test is a five-minute failure instead of a five-day zombiev0.14.0— proef at CI scale:--max-fail Nstops a run honestly (the never-run tail records as skipped, the record is a cancelled rundiffrefuses to certify),--reruncontinues a cancelled run instead of a false green,proef flakyfolds the retained history into verdicts (flapping by transition-count, passes-only-on-retry, broken-not-flaky) completing the detect→quarantine→resolve loop the@quarantinetag already anchored, and--shard I/Npartitions a matrix by a frozen identity hash so adding a scenario never re-buckets the others — plus the reverse docs gate: every subcommand must be documented, enforced rather than noticedv0.15.0— validation rounds 17–18 + the Robot Framework capability audit:@skip/@skip:reasonand@quarantinevisible in every sink (ADR-0019), tag globs + per-tag report verdicts +[tag-links], explicit run metadata (--meta/[meta], ADR-0020), the rerun overlay (one JUnit and one report covering the whole suite),--console dotted|quiet,--shuffleseeded by the run id,reproduce_hintinto the record — breaking: quarantined failures reach JUnit as skipped-with-message,--shardre-deals (the hash gained fmix64), tag atoms glob, JUnit identity isclassname+name. Its tag was missing for two weeks (found 2026-09-09): the release commit landed 2026-08-25 butv0.15.0was never pushed, andrelease.ymlstarts on the tag alone — so the pipeline never ran and 0.15.0 had no GitHub Release, binaries or attestations. Backfilled 2026-09-09 with the workflow disabled for the push: the tag now points at the release commit and the Release carries the changelog section, marked not-latest, with no archives — the only release without them. Not repaired by simply pushing the tag; the runbook above says which workflow a tag actually runsv0.16.0— the surfaces tell the truth: an eight-wave improvement programme (#112–#142) plus the round that found what it missed (#143–#150). CI-sink conformance (JUnit detail into element content, an XML-1.0 control-character boundary, real limits on the GitHub summary and annotations), a triageable and linkable HTML report, console colour, shell completions and a man page in every archive, a project-awaredoctor,explain/diff/doctor --format json,--console failed,flaky --by,proef schema config, and the LSP wave — document symbols, hover, quick-fix code actions off a structuredDiag::fix, one analysis per edit rather than per keystroke, and a panic guard on both message-loop entry points. Pack validation became linear in the macro count (65× at 3200 macros) and the last unfuzzed parser gained a target — breaking:--outputsplit by meaning into--format(which format) and-o/--output(which path),World::set_globalreturns a#[must_use] bool,ConsoleReporter::newtakes acolorflagv0.17.0— the environment a suite runs in, and the guards that keep its claims true:[http]gained the keys that describe an environment rather than a request (TLSinsecure, proxy, mTLS cert/key,max-redirs,user-agent,cookie-store = false),--ctrfrenders the run off the same fold as JUnit,--shard-weightsbalances a matrix by measured duration from one sharedtimings.json, and the HTML report answers “what is slowest”. The hurl-coverage audit closed the two defects aref:fragment could not work around (#164–#166): afile,…;body resolves beside the file that wrote the reference, and two features’ same-named assets stop overwriting each other. A--run-idrecord is findable again (ADR-0021), a disk filling mid-run reaches the exit code, and the 23 diagnostic codes that had no test got one — breaking:emit::file_referencesbecameArtifact::assetscarrying each reference with the source that wrote it,emit::asset_rootis new,HttpDefaultsgained eight fields and lostCopy, and the canonical artifact format moved (an artifact that reads a file now names its--file-rootin the replay line)v0.18.0— the CI-consumer surfaces, run to exhaustion (#168–#179): output proef could not deliver never looks like success. A run-record write failure latches into exit 3 through one fold (escalate_environment_failures, beside the JUnit/CTRF and GitHub-summary failures), SIGTERM/SIGHUP take the graceful cancel so a CI job timeout leaves a complete record and its reports, and a custom--run-idno longer collapses the JUnit identity onto the nil uuid. Asset staging resolves beside the file the parser read wherever youcdfrom, with--sariflines from the carried source and the symlink and case-insensitive staging edges closed. The ADR-0007 budget family is closed over its inputs and bounded as a product (a four-hour batch ceiling,[http] timeout-ms = 0refused), everyRunSummarysink routes identities through the masker, andproef flakygained the 2026 statistical guards — a sample floor, hysteresis, an environment-outage guard, and an input-fingerprint equivalence class — breaking:proef_lsp::RootResolverreturns aResolvedRoot,timings::rendertakes a&Redactions, andflaky’snewverdict is renamedinsufficient-datawith its default sample floor rising from 2 to 10v0.19.0— the checks that could not see what they claimed to cover (#186–#191). The Homebrew formula installs the man page and the shell completions again — broken since 0.16.0 because the formula is a heredoc inrelease.ymland the archive is staged in another job, so nothing tied the two together; the render step now fails if the archive lacks a file the formula installs, and this tag is the first to carry it.--dry-runrefuses afile,…;asset that is not there, through staging’s own checker rather than a second walk, so the gate CI runs before standing an environment up is no longer blind to a defect that is entirely static. The asset scan moved behind the engine seam: it was a scan for the literal"file,"inproef-corethat the ADR-0002 guard structurally could not classify — so it was never sanctioned and never reported missing, and that ADR’s “thirteen literals” measurement is corrected to fourteen — while reading hurl’s own AST also stopsfile,inside a JSON body counting as an asset. A fragment’s assets stage from where its file was read rather than from where its recorded name points, the fragment-side twin of whatLoadedFeature::read_fromalready does.--rerunreads its base record once instead of twice, and the one doc check that only reads files moved into the half of the gate that only reads files — breaking:proef_core::emit::emittakes the registered step kinds andStepKindSpecgains anassetshook, replacingemit::file_refs_in
Contributing to proef
Small, focused PRs against main. The corpus in docs/ is the source of
truth — CLAUDE.md is the working summary, docs/TECH-SPEC.md and the ADRs
win on conflict.
Setup
The toolchain is pinned by rust-toolchain.toml (latest stable; rustup picks
it up automatically). One-time tools:
cargo install cargo-nextest cargo-deny cargo-audit cargo-insta just
# plus, for the full CI surface locally:
cargo install cargo-fuzz cargo-public-api cargo-machete cargo-llvm-cov # llvm-cov: `just cover`
Native build prerequisites (only proef-engine-hurl needs them):
Debian/Ubuntu apt install build-essential pkg-config libssl-dev libcurl4-openssl-dev libxml2-dev libclang-dev; macOS: Xcode CLT.
The gates (green before every commit)
cargo nextest run # all tests
cargo test --doc # doctests (nextest skips them)
cargo clippy --all-targets --all-features -- -D warnings
cargo fmt --all --check
RUSTDOCFLAGS="-D warnings" cargo doc --no-deps --all-features --workspace
cargo deny check
cargo run -p xtask -- docs-check # indexes ↔ reality
cargo run -p xtask -- public-api # proef-core API surface (nightly rustdoc)
cargo check --manifest-path fuzz/Cargo.toml --all-targets --locked # fuzz/ is its own workspace
cargo machete # unused dependencies
just gates runs the set above. CI additionally runs the #[ignore]d
complexity guard alone (just perf — TESTING-STRATEGY §7), zizmor, a
proef doctor smoke, a fuzz smoke, the hurl canary, a Windows gate, the
docs-site build, and — on a pull request — the changelog self-recording check;
cargo audit and the full fuzz run nightly.
Rules that are easy to trip over
- hurl pins are exact (
=8.0.1, built--locked). Never bump them in a PR — upgrades go through the canary + runbook (ADR-0003). - Snapshots are deliberate. Artifact bytes, sidecars, diagnostics, and
event streams are insta-locked;
cargo insta revieweach diff and be able to say why it changed. Never blind-accept. proef-coreAPI is snapshot-locked (crates/proef-core/public-api.txt). An intended surface change regenerates it:PROEF_PUBLIC_API_UPDATE=1 cargo run -p xtask -- public-api.- Core purity:
proef-coredoes no IO and reads no clocks/env/randomness — inject values instead. This keeps every snapshot deterministic. - New architectural decision → new ADR (
docs/adr/ADR-00NN-*.md, next number, same format) in the same PR. Diverging from an ADR without a superseding one is a bug. - New diagnostic → index it in
docs/DIAGNOSTICS.md, prefer a seeded case undertests/errors/<area>__<name>/(dry-running that corpus fails by design). - YAML is
serde_norway, datetime isjiff, and reqwest/async-trait/ a tokio runtime are banned (seeCLAUDE.mdfor the full list and why). - No raw print macros in
proef-cli. Usecrate::render::outln!for stdout andcrate::render::errln!for stderr.println!/eprintln!panic when the write fails, and a closed pipe (proef … | head) surfaces as EPIPE rather than a signal — so a raw macro aborts with 101, outside the typed 0/1/2/3 exit contract (ADR-0009). A source-scanning test enforces this.
Testing
docs/TESTING-STRATEGY.md is normative. In short: everything is device- and
network-free except the fixture integration suite
(cargo run -p xtask -- fixture runs the dev API server standalone). Assert
attempt counts and normalized event order, never wall-clock. just cover
measures line coverage on demand (cargo-llvm-cov; advisory, never a
threshold — TESTING-STRATEGY §3).
Commit messages
Conventional prefixes (fix:, feat:, docs:, refactor:, release:),
imperative subject, body explains why. No AI-attribution footers.
Security policy
Reporting a vulnerability
Use GitHub’s private vulnerability reporting on this repository (Security → Report a vulnerability). Please do not open public issues for security reports. You will get an acknowledgment within a week; fixes ship as patch releases with a CHANGELOG entry.
Supported versions
The latest released 0.x version. Pre-1.0, fixes are not backported.
Threat model
proef is a test tool with a deliberately modest threat model: it protects secret material at rest and keeps it out of every output, and it does not attempt to defend a compromised host.
What proef guarantees:
- The secret store (
.proef-secrets.json) holds only XChaCha20-Poly1305 ciphertext (enc:v1:envelope) — safe to commit and share. - Secret values never appear in any sink: artifacts carry
{{name}}placeholders, events/logs/reports are value-redacted at the sink boundary (property-tested) — the event stream through one exhaustiveapply_event, and the CI sinks that render from the run summary (JUnit, CTRF, TAP,timings.json, the GitHub summary and annotations) through its twinapply_outcome, so a new text field cannot ship unmasked — and asaveAs: globalcapture whose value equals a known secret is refused rather than persisted to the plaintext.proef-state.json. - Sensitive files (
.proef-secrets.json, the key file,.proef-state.json) are created0600, private from the first byte.proef doctorwarns when permissions have drifted. - Request file bodies are confined: every
file,…;asset is staged from beside the source that names it into the run’s per-scenario asset root, which is the engine’scontext_dirsandbox. A reference must be a plain relative path (no leading/, no..); a symlink already sitting at a staging destination is replaced rather than written through; and two references that are one file to a case-insensitive filesystem are refused rather than last-writer-won. - A fragment corpus is read, never written (ADR-0018). Pointing
[run] fragmentsat.hurlfiles somebody else owns is one-directional:proef fmtrefuses them in both discovery branches, and the declared root is the confinement boundary — nothing outside it is scanned. Files come back byte-identical, which an integration test asserts. - A renamed secret is still never materialized.
bind: { token: "${secret:x}" }lets a foreign corpus keep its own variable name; the value still travels viainsert_secretand never enters the artifact. Mixing a secret into a larger bound value is refused (lower::secret_in_composite_bind) rather than quietly written out, because injecting the joined string would require putting it in the artifact. - TLS verification is on unless a profile says otherwise, and saying so is
loud.
[http] insecure = trueexists because staging environments really do present self-signed certificates, but a suite that goes green without verifying one has not proved what a green suite normally proves. Every run with it active prints a warning naming the profile that set it. The run record deliberately carries no config, so that warning is the whole audit trail — which is why it cannot be suppressed. - mTLS credentials are file paths, not values.
[http] client-cert/client-keyname files; proef reads no key material into its own memory and writes none into any artifact. Aclient-keywithout aclient-certis exit 2 rather than a silent pass-through: libcurl would accept the pair and then present nothing, so the failure would surface at the server as an authentication error naming nothing about the cause. - Release binaries are built with
cargo auditable(dependency trees stay scannable) on cache-isolated CI runners.
What proef does not defend against:
- A compromised host or user account: the key file lives on disk, decrypted
values live in process memory (no zeroize — hurl holds its own copies),
and
PROEF_KEY/PROEF_SECRET_*are readable from the process environment. - Malicious suites: packs execute arbitrary HTTP requests by design; run suites you trust.
- Credentials written into
proef.toml. There is deliberately no[http] userornetrckey — a password belongs in the secret store, where it is encrypted at rest and masked out of every sink.[http] proxyis the one edge: a proxy URL embedding credentials is plaintext in a file you probably commit, and proef cannot mask a value it was never told is a secret.
If your environment needs more than this, inject secrets per run via
PROEF_SECRET_<NAME> from a real secret manager and skip the store entirely.
Changelog
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog;
versioning follows SemVer (policy in docs/RELEASING.md).
Each release groups its entries under one heading per kind, in the order
Added · Changed · Fixed, with Breaking / Internal / Documentation after
them where a release used those. Three shipped releases carried the same
heading two or more times — a release cuts by moving [Unreleased] wholesale
(RELEASING.md), so whatever shape it had at the time shipped verbatim.
Regrouping preserved every entry and its order within its kind.
[Unreleased]
Fixed
-
The sidecar’s two scanners now use hurl’s own grammar rather than an approximation of it. Both were read off
hurl_core’s parser and corrected against it, and both errors cost rows in.map.json— a normative artifact whose contract is that no legitimate row is dropped and no invented one appears (ADR-0010).is_method_linedemanded three characters whilehurl_core’smethodparser takes one or more ASCII uppercase letters, so a short method opened an entry proef’s capture scan did not see: the previous entry’s[Captures]run stayed open across the boundary, and a header of the next entry (X-Trace: abc) was recorded as a capture nobody wrote. The same predicate allowed-, which hurl’s grammar does not, so a dashed uppercase word could end a capture run on a line hurl would refuse to parse as a request — the two errors pulled in opposite directions and hid each other, which is how both survived from 0.1.0.A capture name was matched against
[A-Za-z0-9_-]while hurl’skey_string_textadmits anychar::is_alphanumeric— Unicode, not ASCII — plus_ - . [ ] @ $. Souser.id,items[0],@type,total$andprécisall parse as captures and were all silently missing from the sidecar. A leading[stays refused, matching hurl, and{/}stay out deliberately: a name written as a template has no statically knowable text to write a row for.
Internal
-
The secret-resolution order is pinned where it is decided. “An unset
PROEF_SECRET_<NAME>still reads the store” is load-bearing — every run that keeps its secrets in the committed store depends on it — and was asserted only by an integration test that stands up a fixture server, executes a suite and checks exit 0. That test does catch a regression, indirectly and by way of an exit code.secretstore::resolve_allnow carries unit tests for both directions of the precedence and for the neither-source error, each verified against a mutation that breaks it.One of those tests had to be rewritten first. It claimed to prove the
from_store.is_empty()early return by pointing at a corrupt store, and deleting that return left it green:load_store’s error only reaches the caller through names that needed the store, and there were none. The early return is an IO saving, not an observable behaviour. The corrupt-store case is kept as its own test, saying that. -
A located-lines undercount can no longer reach a caller looking complete.
locate::key_line_spansreturned a bareVec<Span>, and it sees only block-stylekey:lines — a flow-style- {use: base}item is valid YAML, parses to a real step, and contributes no line. So the list can be shorter than the items it describes, and pairing them positionally attributes every span after the gap to the wrong item: a go-to-definition landing on the neighbouring line, a diagnostic pointing at it.Both callers already knew, and each had written its own length comparison in its own words from a prose warning. Both were correct; neither was enforced, and a third caller would have had to rediscover the hazard and the remedy together. The scan is now private behind a
KeyLinesvalue whose only accessors arepaired_with(parsed)— which yields the spans only when the counts agree — andsole(), for a key likematch:that occurs at most once and has no sequence to pair against. No behaviour changes; what changes is that the guard is the only way through.
Documentation
- 0.19.0 is recorded where the corpus says it should be. The release
History in
RELEASING.md, the milestone Status inCLAUDE.md, the corpus index’s “through vX.Y.Z” line, andOPEN-FINDINGS’ own re-check date. This is the set that drifted after 0.15.0–0.17.0 — three tags with no History entry, concealed by a fourth filed out of order — so it is done in the same session as the tag rather than left for the next reader to discover.
[0.19.0] - 2026-09-11 (the checks that could not see what they claimed to cover)
Fixed
-
The Homebrew formula installs the man page and the shell completions. Its
def installwasbin.install "proef"and nothing else, so from 0.16.0 — the release that started shippingproef.1and fivecompletions/files in every archive — until 0.18.0,brew install proefgave noman proefand no tab completion, while binstall and a direct download gave both. The formula is a heredoc insiderelease.ymland the archive is staged in a different job, so nothing tied the two together and no gate could see the gap; the render step now fails if the archive lacks a file the formula installs, and the formula’s owntest doasserts the man page and completion landed. Takes effect on the next tag: a tag runs the workflow from its own commit. -
--dry-runrefuses afile,…;asset that is not there. A suite whose asset had been deleted reporteddry-run OK, and the failure arrived later from a different command, against a live backend — from the one gate CI runs before standing an environment up. Whether an asset resolves is statically knowable, so it is answered there now. The checker is staging’s own (assets::resolve_assets, split out ofstage_assets) rather than a second walk over the same artifacts, so validation and the run cannot disagree; the message and the diagnostic code are the ones a run already gave. -
file,inside a JSON or assertion body is no longer mistaken for a file asset. The emitter found the files an artifact reads by scanning its text for the literalfile,and a closing;, so a request body containing that substring —{"note": "see file,notes.txt; for details"}— produced a phantom asset, and staging then failed the run over a file the request never reads. The claiming engine now reads its own AST, where a body reference and six characters of prose are different things.
Breaking
-
proef_core::emit::emittakes the registered step kinds, andStepKindSpecgains anassetshook. Asset recognition was hurl’s body grammar living inproef-core:emit::file_refs_inscanned for the literal"file,", which ADR-0002’s amendment forbids and — worse — which the guard pinning that amendment could not see.engine_grammar_kindclassifies fences,HTTP,[Section]headers, method lines andkey: valueoptions; a body constructor is none of those, so the literal was never sanctioned and never reported missing. The ADR’s own measurement said thirteen literals; it was fourteen.The scan moves behind the seam as
StepKindSpec::assets, the fourth engine-contributed hook besidevalidate,fragmentsandoptions, and the guard gains abodyarm so the shape is classifiable whether or not anything currently uses it.emit()takes&[StepKindSpec]to reach it;FrontEndcarrieskindsbeside thekind_to_enginetable it is built with, whichregistryalready documents as a pair that must not be re-derived separately.emit::file_refs_inis gone.
Internal
-
A fragment’s assets stage from where its file was read, not from where its name points.
AssetRoots::source_dirrebuilt a fragment’s directory by splittingfile.hurl#nameand joining the file half onto the project root — the naming boundary run backwards, without the canonicalize fallback that boundary carries precisely because a lexical-only version already shipped a bug (a suite reached through a symlink silently failed to match, R11-9). The two agreed only because both were seeded fromconfig.root()and discovery walked from that same root, so only the lexical case was ever exercised, and nothing made them stay inverses. The corpus reader now records the directory it read each file from (front::CorpusDirs, carried onFrontEndbesidekinds), and staging looks it up — the fragment-side twin of whatLoadedFeature::read_fromalready does for features, so both halves of the naming boundary are one-way in the same way.AssetRootsloses itsprojectfield andbuild_specsitsproject_rootargument: with nothing to recompute, the project root is no longer staging’s business. -
--rerunreads its base record once. It calledrecord::read_eventsfor the JUnit overlay and thenrecord::rerun_candidates, which read and deserialized the sameevents.jsonla second time — two full passes bounded only by the 256 MiB record ceiling, over a file another process may still be writing.rerun_candidatesnow takes the&[Event]its caller already holds, which is the ruleread_record’s own documentation had already stated for exactly this case. The read error is handled once as well: the first call swallowed it with.ok()and the second rediscovered it a line later. -
The one doc check that reads only files now runs in the half that reads files.
no_current_behaviour_doc_spells_a_format_as_an_output_pathlived intests/docs.rs, whose stated charter is the checks needing a built binary to ask clap — this one only scans markdown, so it never ran in the fast doc-only CI step. It is nowxtask docs-check’scheck_output_path_spelling, reusingliving_docs()instead of carrying a second directory walk. Its allowlist-shrink guard got stricter on the way: it counted ADRs into the same total, so a renamed entry could be masked bydocs/adrbeing larger than the shortfall — which is the one failure that guard exists to catch. All three paths were checked by mutation: a stale spelling planted in an allowlisted doc, one planted in an ADR, and an allowlisted doc renamed away.
Documentation
-
The worklist stops contradicting what shipped. Three entries in
OPEN-FINDINGSstill called CTRF declined or its trigger unfired — the 2026-08-31 external re-test, the RF audit’s deferred list, and R3-5 under “deferred, with the trigger named” — for the eight days after--ctrfactually shipped (#160). R3-9, four bullets below R3-5 in that same list, was annotated the moment it shipped — the convention the three missed. Two more claims had outlived their facts: the shipped-changelog duplicate headers (no release carries one now, andcheck_changelog_kindsfails if one returns) and the machine-side note about Homebrew’s Rust shadowing rustup. Filed at the same time:a_second_interrupt_hard_exits_with_130failed once on Linux CI and passed on a re-run of the same commit, so the evidence, the mechanism and the fix shape are written down instead of left to the next re-run. And the stance that a scenario-level@retryis deliberately absent — retry-until-green hides a one-in-four defect 99.6% of the time — is stated inTESTING-STRATEGY§5, which the worklist asked for and nobody had written. -
The runbook records that the registry skips three versions. 0.15.0–0.17.0 were tagged and GitHub-released but never published, so crates.io moves 0.14.0 → 0.18.0. Noted in
RELEASING.mdso the gap does not read as a failed upload. The long-standinghomepagequestion inOPEN-FINDINGSis also resolved: the field reached the registry with 0.18.0, exactly as that entry predicted;documentationremains unset and still open. -
The release history records every release again.
RELEASING.md’s History section carried no entry forv0.16.0orv0.17.0and filedv0.15.0betweenv0.13.0andv0.14.0; the order is repaired and all three versions are present,v0.18.0included. The corpus also stops calling the 0.18 series unreleased, and anIMPROVEMENT-PLANpointer into CHANGELOG[Unreleased]now names the releases that actually carried the work —[Unreleased]has been cut several times since that sentence was written.
[0.18.0] - 2026-09-09 (the CI-consumer surfaces: output proef could not deliver never looks like success)
Added
-
SIGTERM and SIGHUP now take the graceful path (ctrlc’s
terminationfeature): a CI job timeout ordocker stopcancels the run — in-flight batches finish, the rest record as skipped, teardown runs, the reports are written, and the record closes with acancelledrun_finished— where it used to kill the process mid-write and leave a truncated record with no tail. A second signal still hard-exits 130 (the handler carries no signal identity, so the code is 130 for every second signal). Pinned bysigterm_cancels_gracefully_and_the_record_completesand — for the first time anywhere — an exit-130 assertion,a_second_interrupt_hard_exits_with_130. -
test --format jsonandexplain --format jsonnow reportwarnedandcancelled. A warned scenario (anoptional:step failed, or asaveAs: globalpromotion was refused) folded intopassed, andcancelled— in the record’srun_finished— was surfaced by neither, so a script could not tell a spotless run from one with warnings, nor a complete run from a cancelled one, and the two JSON surfaces disagreed on how to say “did not finish” (0.18 survey). Both keys are additive and always present.warnedalso becomes visible in JUnit (a<system-out>note, the status stayssuccesssince JUnit has no warned) and CTRF (anextra.warnedflag) — it was previously visible only in the HTML report. -
A tag that looks like a reserved one but is not exactly it now warns (
proef::tags::reserved_tag_typo).@quarantined,@skipped,@Skipmatched no reserved tag and silently did nothing — a scenario the author believed was quarantined gated the build. The warning names the spelling it likely meant, tuned to catch the real typos without firing on legitimate short tags (ship,slip,step). -
proef flakygains the 2026-field statistical guards (0.18 survey §6), each a pure fold over the JSONL history already retained — no new state, no gating mode (advisory stays the design):- A minimum-sample floor (
--min-samples/[flaky] min-samples, default 10): below it a scenario isinsufficient-datarather than classified, because a verdict on thin data is worse than none. - Hysteresis (
--recovery-runs/[flaky] recovery-runs, default 5): a flapping or latent scenario holds its flag until it earns a trailing clean run, so it cannot oscillateflaky↔healthybetween adjacent runs. - An environment-outage guard (
--outage-rate/[flaky] outage-rate, default 0.8): a run where over this share of suite scenarios failed is an environment incident, not evidence about any one scenario, and is excluded — so a single fixture or staging outage cannot mark the whole suite broken. - An input-fingerprint equivalence class — the default key. Each run
writes an
inputs.jsonsidecar carrying a hash of what it executes (feature sources + loaded macros/fragments + the resolved${url:…}/${vars:…}scope), so a pack, feature, orproef.tomledit correctly ends the comparison window instead of silently mixing runs of different inputs. It is a proef-computed fact about proef’s own inputs, not harvested from the environment (ADR-0020 unchanged — git-commit grouping stays handed-over via--meta commit=…andproef flaky --by commit).broken≠flaky, transition-counting, and the quarantine lifecycle were already present and are unchanged.
- A minimum-sample floor (
Fixed
-
A run-record write that fails now reaches the exit code. The JSONL reporter deliberately swallows write results (a reporter cannot report its own channel dying), and
events.jsonlwas handed a bareFile— so a disk filling mid-run truncated the record while the run still exited by its verdict, the exact class the v0.6–v0.8 series closed for the console. The record’s writer now latches its first failure (one stderr line, run continues) and the exit funnel turns it into a system error, the same shape as the stdout latch and the JUnit-write fold — unified in one pinned function,escalate_environment_failures.run.log’s mirror keeps its own contract (creation is warn-and-continue, so a mid-run failure warns once and leaves the verdict alone — previously it was silent). -
The GitHub step summary can fail again. It was the only CI sink that couldn’t: a failed open or write vanished while JUnit and CTRF failures re-classify the exit — so the page a reviewer actually reads could be missing on a green exit.
write_github_summarynow returns the error and the caller folds it into the samereports_failedpath as its siblings. -
A custom
--run-idno longer collapses the JUnit report identity onto the nil uuid. ADR-0021 made non-uuid run ids first-class, but the report uuid wasparse_str(...).unwrap_or(nil)— every--run-id cirun emitted00000000-…, colliding in any consumer keyed on it. A non-uuid id now derives a stable UUIDv5 from its bytes (a uuid id passes through verbatim). -
The interrupt window and the interrupt’s own words. The handler is installed at the top of
execute— before the front end, the run dir and the record exist — so no startup window takes the process default any more. Its installation failure is a printed warning (it was silently ignored, unlike--watch’s handler). The second-signal path no longer prints before exiting: the print took stderr’s lock, which a worker blocked on a full pipe can hold, wedging the escape hatch behind the very stall it exists to escape. And the teardown notice said “Ctrl-C again to skip” when a second interrupt actually hard-exits dropping every report — it now says what happens. -
Asset staging no longer depends on the working directory. A feature’s
file,…;assets were resolved by joining its portable name against the cwd — but a name’s anchor (the project root, or the caller’s own typed spelling) is not in the string, so a typed-absolute or config-written suite path run from any subdirectory failed staging with exit 2, blaming the author for a correct file (the feature-side twin of OPEN-FINDINGS H5). The resolved discovery path now travels beside the name (LoadedFeature::read_from) and staging resolves beside the file the parser actually read — the H5 prescription, applied to the feature side. Reproduced before the fix and re-verified after, from a subdirectory, against the reference corpus; a new integration test pins a project under a path with spaces and non-ASCII segments, which nothing in the suite had ever exercised. -
--sarifline numbers survive acd, and byte-match the parser. The SARIF writer re-read each source from disk by its portable name to count lines — from any subdirectory every read failed andstartLinesilently vanished, annotating nothing; the re-read could also disagree with the span by exactly the parser’s normalization. Lines now come from the diagnostic’s own carried source text — the same normalized bytes the span indexes. (On Windows, an absolute out-of-projecturialso spells its separators as a URI requires.) -
Staging’s two symlink edges. An existing symlink at a staging destination was written through —
fs::copyfollows links, so the bytes landed wherever it pointed, outside the root built to contain them; it is now replaced. A source symlink stays followed, deliberately: stockhurlfollows it too, and refusing would break the dual-runner rule (the module doc now says so). -
Asset names that are one file to the filesystem are refused. The duplicate-name guard keyed on the raw reference string, so
Data.jsonanddata.json— one file on macOS and Windows — silently last-writer-won, the very overwrite the per-scenario root was built to end. The check now runs on the canonical path the copy actually landed on, which is exact on every platform: a case-sensitive volume keeps both files legitimately, and nothing fires. -
Artifact slugs cap at 120 bytes. The slug flattens the feature’s whole directory path into one filename component, and
assets/<slug>/repeats it as a directory — so path depth became filename length, and a deep tree or a long scenario name (multi-byte scripts at a quarter of the visible characters) sailed past NAME_MAX and failed the write. Over the cap, the tail is a hash of the whole uncapped slug, so two names differing only past the cut still name two artifacts; every slug the existing corpus has is under the cap and unchanged byte-for-byte. -
The ADR-0007 budget family is closed over its inputs, and bounded as a product.
[Options] max-time:was read by the budget calculator (as the entry’s timeout) while invisible to the lint —max-time: 100000hwas lint-clean and produced a multi-year watchdog budget; it now carries the duration cap, and a test pins the rule the hole broke (every option the budget reads must be one the lint can see).retry-interval:— the one uncapped multiplicand — carries the cap too. And because individually capped values still compose into an unbounded product (retry: 10_000× a 30 s timeout is ~83 lint-clean hours, saturating toDuration::MAX, whose deadline addition panicked as a phantom “scenario thread panicked” fault), the computed batch budget now clamps to an absolute four-hour ceiling and the dispatcher’s deadline arithmetic can no longer overflow. ADR-0007 carries the amendment. -
[http] timeout-ms = 0is refused. libcurl reads zero as no timeout, so the value opted a suite into exactly the unbounded hang the default exists to defend against — while reading like “immediately”. Exit 2, in whichever table it appears. -
Every sink that renders run values now routes identities through the secret masker. The event stream masks
scenario,file,tagsand the skipreasonunder an explicit no-exemptions rule (“a field exempted because it can’t contain one is how that stops being true later”), and five sinks bypassed it for the same fields (0.18 survey): the GitHub annotationtitle=/file=lines (written to CI stdout), TAP’s skip reason and scenario name, CTRF’sname/suite/filePath/tags, JUnit’s suite/testcase identity andfileattribute, andtimings.json— the one sink that took noRedactionsat all, in the file whose documented workflow is being archived and shared across a CI matrix. Structural mitigations (secrets lower to{{name}}; the engine pre-redacts details) made a live leak unlikely, but the boundary rule was unenforced; a per-sink leak test now pins each, and a whole-run sweep asserts a reflected secret reaches no file any sink writes. -
proef lsphonours--env. The global flag was parsed and then silently dropped forlsp, soproef lsp --env staginganalysed the default profile while runs used staging — the editor/runner drift R10-1 closed for--config. And the workspace-root re-resolution (for an editor launched outside the project) re-loaded the config to find the root but dropped the${url:…}/${vars:…}scope it had computed, analysing the right tree against the wrong directory’s config; the scope now travels with the root it belongs to. -
A CTRF report cannot gain a key the spec would reject. CTRF §4.4 makes consumers reject any key outside the defined set (unless under
extra), and the spec moved five times in 2026 — so an additive field is a hard break. A test pins the exact allowed key sets. -
De-flaked three tests (0.18 survey): the abandoned-scenario record-gate test waited on a 500 ms blind sleep that passed vacuously on a loaded runner — it now polls a drop latch set strictly after the worker’s final emit attempt, so it tests the dropped event on every machine; the bounded-runtime smoke test’s wall-clock assertion is widened and documented as the “generous upper bound” class TESTING-STRATEGY §7 sanctions (distinct from the
#[ignore]d ratio guard); and a watch test’s fixed shared temp path (temp_dir()/proef-watch-alias-test+remove_dir_all) — the one cross-process race nextest cannot cover — moved to a uniquetempdir. -
proef doctorno longer prints fourteen literal spaces mid-sentence (a lost line continuation in the “hurl not on PATH” note).
Breaking
-
Library:
proef_lsp::RootResolvernow returns aResolvedRoot(root+disk+config_vars) instead of a(PathBuf, Box<dyn SourceProvider>)tuple, so the re-resolved config scope reaches the server.proef-cli’slsp::runtakes the--envvalue.timings::rendertakes a&Redactions. -
proef flaky’snewverdict is renamedinsufficient-data(its--format jsonverdictkey and human label), matching the 2026 vocabulary and the new sample-floor meaning — a MINOR break for a consumer keyed on the old spelling. The default per-scenario floor also rises from 2 to 10 runs, so a scenario with fewer than 10 runs now readsinsufficient-datawhere it previously received a verdict (--min-samples 2restores the old behaviour).
Internal
-
Post-0.18
/simplifycleanup — duplication the wave programme left behind, collapsed with no behaviour change (outputs byte-identical, no public API moved):- The input fingerprint’s FNV-1a loop and
fake’s were the same loop and constants twice; now onefingerprint::fnv1a_withprimitive, withfake::fnv1aa thin alias at the canonical offset basis. proef testandproef --watchduplicated the whole two-stage-interrupt skeleton (the once-latch, the second-signal hard-exit, the stderr-lock rule); now oneinstall_two_stage_interrupttaking the divergent first-signal action as a closure.proef flakyrecomputed each scenario’s verdict at ~8 sites — twice per comparison inside the sort; now classified once into a stored field, andrender_tableno longer threads the thresholds through to recompute it.- The reserved-tag typo warning derives its edit-distance threshold from
each reserved word’s own length instead of hardcoding
quarantine, so a future reserved tag earns fuzzy protection automatically, and it builds its diagnostic once rather than twice. is_outagecounts without a throwawayVecand drops a dead precision-loss suppression; JUnit redacts a scenario’sfileonce, not twice;emit::cap_slugdrops a redundant rebinding.
- The input fingerprint’s FNV-1a loop and
-
Redaction centralized at one exhaustive boundary (ADR-0005 hardening, no behaviour change on clean output). The CI sinks (
JUnit, CTRF, TAP, timings, the GitHub summary) render fromRunSummary, not the event stream, and each masked its identity and failure strings field by field — correct today, but a new field or sink could slip past unmasked. A newRedactions::apply_outcomemirrors the event stream’s exhaustiveapply_event: it destructuresScenarioOutcome/StepOutcomewith no.., so a new text field fails to compile until it is masked, and each sink now redacts an outcome once instead of the ~10 scatteredapplycalls it used to sprinkle (a scenario’sfaultmessage, which reaches only this path, is masked with the rest). Additive to the library surface (pub fn Redactions::apply_outcome). -
Redaction masking deduplicated to one primitive (a follow-up
/simplifypass, no behaviour change).apply_outcome/apply_step_outcomewere written as siblings ofapply_eventbut re-spelled itsArc<str>masking idiom inline and dropped its clean-field optimization; a sharedmask_arc/mask_step_refnow backs all four maskers, so a clean field reuses itsArcinstead of reallocating (andapply_step_finishedinherits the same win).timingsreverts to masking just the two identity fields it renders, rather than cloning the whole outcome graph to read them. -
Coverage is measurable on demand, and deliberately not a gate (the local half of P13, 0.18 survey).
just cover/cover-html/cover-lcovruncargo-llvm-covover the workspace (~90 % line coverage of the unit + integration suites today,xtaskaside); TESTING-STRATEGY §3 records the policy any CI half must follow — a ratchet that fails only on a drop, never a fixed threshold. The gating CI adoptions the survey also listed (acargo-mutantsjob, a coverage-service job, immutable releases) are a maintainer’s cadence/cost call and stay open in OPEN-FINDINGS. -
The Homebrew tap only moves forward. The release workflow’s tap job is gated on nothing but “a tag was pushed” and rewrites
Formula/proef.rbwhole, so a tag pushed late or out of order would regenerate the formula for an older release and downgrade everybrew upgrade. Not hypothetical: v0.15.0 was released and never tagged, so backfilling that tag would have walked the tap from 0.17.0 back to 0.15.0. The job now compares the tag against the version the tap carries and skips green when it is not newer — green, because publishing an old release’s binaries is legitimate and the correct outcome there is an untouched tap. It guards future tags only: a tag runs the workflow from its own commit, so one cut before this fix still runs the unguarded job, and RELEASING now says to disable the workflow around such a push. v0.15.0 was backfilled that way on 2026-09-09 — tag, and a not-latest Release carrying the changelog section without archives.
[0.17.0] - 2026-09-06 (the environment a suite runs in, and the guards that keep its claims true)
Added
-
[http] cookie-store = falseruns the whole suite cookie-less — hurl 8.0’s--no-cookie-store, surfaced through the table built for exactly this class of setting. NoSet-Cookieis retained and none is replayed, which is how a stateless API is proven stateless: the fixture-backed test is green only because its steps assert the 403 a missing session cookie earns.This is the one
[http]key with no per-entry[Options]spelling at all (OptionKindhas no cookie variant — verified against the enum), so run-wide is not a compromise but the only place it can be said. With the store off, the engine also skips both halves of the batch-split cookie round-trip: hurl reads acookie_input_fileonly when enabling the engine, so injecting one would be silently ignored — and there is nothing to write. hurl’s own FIXME (a handle once given cookie storage cannot lose it) never reaches proef, becauserun_entriesbuilds its client per call (TECH-SPEC §5) — a handle never transitions on → off.Breaking (library):
HttpDefaultsgains thecookie_storefield, so a struct-literal construction needs the new line (..Default::default()sites are untouched, and an absent[http] cookie-storekey changes nothing). -
--ctrf <path>— the run’s verdicts as a CTRF report. CTRF (https://ctrf.io) is the emerging JSON successor toJUnitXML for CI dashboards, and it models in the schema whatJUnitcan only smuggle through extensions — which is exactly the data proef already tracks: a pass-after-retry carriesflaky,retries, andretryAttemptslisting the real failed attempts with their (redacted) messages; every test carries its tags and file path. One serializer off the same fold asJUnit, so the two files cannot disagree — most visibly for a quarantined failure, which both report as skipped with a message (ADR-0019), because a dashboard reading “failed” beside exit 0 would contradict itself. AUser/Systemfault staysfailedeven under a quarantine tag: quarantine is for flaky tests, not broken input.The R12-3 contract applies from day one: a
[run] setupabort still writes the file, carrying the setup scenario itself — a job gating on the report must never see no file at all. The schema’s required wall-clockstart/stopare measured at the CLI edge like every other clock read (ADR-0015); the sans-IO core and the JSONL record are untouched — the record remains the only record (ADR-0008). -
The HTML report answers “what is slowest”. After “what failed”, it is the question a test report is most often asked, and the page could not answer it: the timeline showed that workers were busy, never which scenarios to attack. Every number needed was already in the fold.
A ranked section, slowest first, each row linking to its own block, with the heading reporting the share of run time the listed scenarios account for — “3 of 40 · 71% of run time” is a decision, where a column of durations is homework. Capped at eight: a ranking long enough to scroll has stopped answering the question.
Cost is the sum of a scenario’s step durations, the same definition
timings.jsonuses for shard weights — one notion of what a scenario costs across the whole tool. Not the wall-clock span, which includes time waiting for a worker: a property of how the run was scheduled, and not something the reader can go and fix.Absent when there is nothing to rank — fewer than two timed scenarios, or a record with no injected durations at all.
-
--shard-weightsbalances a shard matrix by measured duration.--shardassigns by a frozen hash, which guarantees that adding one scenario never re-buckets the others but cannot balance by time — and a CI matrix finishes when its slowest shard does, so a count-split routinely leaves runners idle. Every run that reaches its suite now writes a smalltimings.jsoninto its run directory; CI archives that one file and each matrix job points--shard-weightsat the same copy.The obvious design is silently wrong, and the module says so at length. proef already retains records carrying every step’s duration, so “weight by the newest local record” looks free. But matrix jobs run on different machines, each with its own (usually empty)
runs-dir— every job would compute a different weight table, therefore a different assignment, and scenarios would run twice or not at all while the suite reported green. Nothing about that announces itself. One named file shared by every job is what makes the split a pure function of (selected scenarios, that file).Two rules place scenarios and they partition rather than compete: a scenario the file mentions goes through longest-processing-time-first placement, and one it does not mention falls back to the frozen hash. So a test added after the timings were captured still runs exactly once. That is pinned by a test that runs a whole three-way matrix — with a weights file covering only five of nine scenarios, so both rules are exercised at once — and asserts set equality both ways; mutating the placement by one bucket drops two scenarios and the test names them.
The weight is the sum of a scenario’s step durations, not its wall-clock span. The span includes time spent waiting for a worker, which is a property of the run’s scheduling rather than of the scenario, and feeding it back would let one crowded run’s queueing distort the next split.
What this gives up is exactly what hash mode was chosen for: a balanced split is not stable under insertion. That is what balancing means, which is why the flag is opt-in. A missing or malformed weights file is exit 2 — falling back silently would hand back the unbalanced split the flag was passed to avoid.
-
The editor tells proef’s two variable tiers apart. A pack’s
hurl: |block is the centre of the authoring experience and, to every editor, a plain YAML scalar — inside which${…}(resolved at lower time, by proef, before any request exists) and{{…}}(resolved at run time, by hurl) look identical. That distinction is ADR-0005’s whole model and the thing authors most often get wrong, and no generic grammar can see it: a YAML highlighter sees a string, and a hurl highlighter never runs because the block is not a file. proef is the only party that knows.The server now answers
textDocument/semanticTokens/full, lighting${…}as macro — a substitution performed before execution, which is what a macro is — and{{…}}as variable. Both are coloured differently by every mainstream theme, so it works without anyone configuring anything. The$${escape stays dark, because telling an author proef will substitute text it will in fact leave alone is worse than no highlighting.The
${…}scan isproef_core::resolve::reference_spans, walking the samefirst_referencethe resolver itself uses — a second implementation of the escape rule would drift, and the drift would show as an editor confidently colouring literal text. The{{…}}scan lives inproef-lsprather than core, because that spelling is the engine’s and ADR-0002’s amendment is that engine syntax does not accumulate in the core.Collapsing the seven-arm request dispatch behind a local macro came with it: the chain crossed clippy’s line limit the moment an eighth feature landed, and the honest fix was to stop repeating an identical frame seven times rather than to suppress the lint that noticed.
-
The linear-validation claim is now a test, not a sentence. #138 made pack validation linear and recorded the result as a shape: “the curve changed shape — 4× per doubling before, ~2× after”. That number lived only in the changelog, where nothing could re-run it — so a future span locator scanning the whole pack file again would have restored the quadratic behaviour silently, a regression that costs seconds rather than correctness and which no gate measured.
The guard asserts the ratio between 1000 and 2000 macros, because the claim is a ratio. It observes ~2.05× against a bound of 3.0; mutating
locate::MacroIndexto re-index per lookup — the exact pre-#138 shape — measures 4.01×, matching the changelog’s own prediction of 4× and turning a 0.4-second test into a 73-second one. The failure message names the cause rather than reporting a number.A ratio rather than a benchmark, for a reason now written into
TESTING-STRATEGY.md§7: load on a shared runner inflates both measurements together and cancels, where an absolute threshold has to be loosened until it means nothing.iai-callgrindwould be the better CI gate — instruction counts ignore runner noise entirely — but it needs valgrind, so it would be a gate the maintainer cannot reproduce on macOS;criterionanddivansit in the same noise regime as this test while adding a dependency tree to a workspace that audits every edge. No new dependency was added. -
Every diagnostic code is now named by a test, and a guard keeps it that way.
DIAGNOSTICS.mdcalls codes “a contract: they never change meaning”. Twenty-three of seventy-five had nothing holding them to it — reachable in production, documented, exercised by nothing at all: not a seeded corpus directory, not a unit test, not even an assertion on their message text. They existed only at their definition site.The catalogue itself was found exactly honest — 75 codes defined, 75 documented, and its corpus column matched disk in both directions with zero drift. The gap was never documentation; it was that a documented promise had no enforcement.
Nineteen new tests close it, each reaching its code through a real path rather than constructing the diagnostic directly. Two of them exercise guards that are unreachable in normal operation and were therefore the most valuable to test:
lower::kind_unroutedfires only when the engine registry and pack validation disagree, so the test makes them disagree on purpose; andlower::expansion_too_deepsits behind pack validation’s identical depth limit, so the test bypasses validation withload_collecting— the only way to hand lowering a graph validation would have stopped, and therefore the only way to prove the second line of defence is still there.Two codes are exempted by name, with reasons recorded in the guard:
source::unreadableandconfig::unreadableneed a file the process may stat but not read, a permissions state CI runners do not reproduce because they run as root. The guard also checks its own exemption list, failing if an exempted code is deleted or renamed — an exemption that outlives its code silently excuses nothing.The guard joins the four in
source_guards.rsand is mutation-verified: rewriting one test to match a code by suffix instead of naming it turns the guard red, which is the point — a test that matches the prose pins the wording, and only one that names the code pins the contract. -
[http]now carries the settings that describe an environment: TLS, proxy and mTLS. The table exposed two of hurl’s runner options —timeout-msandfollow-location— while the embedded engine has supported the rest all along;TECH-SPEC.md:235even listedinsecureamong whatRunnerOptionscarries. So a suite that had to run against staging’s self-signed certificate, or through a corporate proxy, or against an mTLS-protected API, could not say so anywhere: the only route was repeating an[Options]block inside every macro’s raw hurl, which defeats environment profiles exactly where they are most useful, since these settings are the difference between environments.Eight new keys —
insecure,proxy,no-proxy,cacert,client-cert,client-key,max-redirs,user-agent— each merging field-wise through the existing[http]<[env.<name>.http]chain, so a staging profile turns verification off without production inheriting it. No new concept: only more of one that already worked.Three deliberate edges.
insecure = truewarns on every run, naming the profile that set it — a suite that goes green without verifying a certificate has not proved what a green suite normally proves, and since the run record carries no config by design, the warning is the entire audit trail. Aclient-keywithout aclient-certis exit 2 rather than a pass-through: libcurl accepts the pair and then presents nothing, so the failure would otherwise surface at the server as an authentication error naming nothing about the cause. And credentials are excluded on purpose — there is nouserornetrckey, because a password belongs in the secret store where it is encrypted at rest and masked out of every sink.The three path-valued keys resolve against
proef.toml, the one-path rule every other config path follows; core still reads no filesystem and receives them already resolved (ADR-0012). Each option is applied to hurl’s builder only when actually set, so a project with no[http]table runs byte-identically to one built before the keys existed — pinned by a test. Per-entry[Options]still override all of them exceptuser-agent, for which hurl has no per-entry option at all; that exception is documented rather than papered over.Breaking (library):
proef_core::engine::HttpDefaultsgains eight fields and losesCopy— it now carriesStrings.Defaultstays hand-written, and the reason is now stated in the type: a derive would maketimeout_mszero, which libcurl reads as no timeout at all, silently converting ADR-0007’s budget into an unbounded wait at every existingdefault()call site.
Changed
-
The toolchain policy is stated in the spec that
rust-toolchain.tomlcites. R18-2 corrected the policy to latest stable, adopted at itsx.y.1point release, and the correction reachedRELEASING.mdandCLAUDE.mdwhile TECH-SPEC §15 — named byrust-toolchain.tomlas its authority — still said “always latest stable”. R18-2’s own conclusion was that an unwritten policy contradicting the written one is a docs defect; fixing it in two files and leaving the source of truth contradicting itself reproduced the defect one level down. Now consistent across all four. -
An artifact is named by its feature’s path, not its stem — two scenarios can no longer claim one file. Slugs were
{stem}--{scenario}, dropping the directory, sofeatures/x.featureandfeatures/sub/x.featureeach with asame namescenario both producedx--same-name: the second artifact silently overwrote the first while the CLI reported writing two. Silent loss of the hand-off ADR-0010 calls a contract — and the project already treats same-named scenarios across files as real, which is what--scenario-fileexists for. The same slug drives the HTML report’s anchors and artifact links,reproduce:lines, and harness trial names, so all of them move together off the one helper.Names are now
features-sub-x--same-name. Derived from the path rather than disambiguated on collision, deliberately: a counter or hash appended only when two names clash would make one scenario’s artifact name depend on whether some other file exists, so adding a feature would rename an unrelated artifact — the instability--shard’s frozen hash exists to avoid. The path fed in is the portable suite-relative name the record carries, never a path off the running machine.Breaking, and quietly so for library callers:
emit::artifact_slugkeeps its(&str, &str) -> Stringsignature while its first argument changes meaning from stem to feature path, so the API gate cannot see it — passing a stem still compiles and now yields a different name.emit::emit’s second parameter changes the same way, andemit::feature_stemis removed (it had no remaining consumer). Artifact filenames and report anchors change for every suite; the snapshot corpus was regenerated under the new names and reviewed.The unification that made that a one-line change came first: the stem expression (
file_stem, falling back to"feature") had existed four times across both crates — the emitter’s caller, the dispatcher’s spec naming, the HTML report’s anchors, the editor’s analysis — and thestem--scenariocomposition twice, with the report’s links to artifact files resolving only because both sides happened to derive the same name. Worklist item Q6 called the four sites a future-drift risk; collapsing them to one helper is what let the collision above be fixed in a single place instead of four. In the same pass, Q2 (the editor’s per-request walks) was found already closed by the #146 analysis cache, and its entry now says so with the evidence. -
“What a scenario costs” is defined once, as
ScenarioOutcome::cost. The sum of a scenario’s step durations was computed in three places on the same type —JUnit’s per-suite time,JUnit’s per-case time, and the newtimings.jsonweights — plus a fourth over the record-fold shape in the HTML report. Four surfaces free to drift apart about a number they are supposed to agree on, and the argument for summing steps rather than taking a wall-clock span was written out twice.Now a method on the type that owns the steps, with the rationale stated there and referenced from the rest. The one behaviour change is a fidelity gain: the weights file used to truncate each step to whole milliseconds before summing and now truncates the sum, so its numbers agree with the times
JUnithas always reported. Additive to the library surface. -
A run whose setup aborted no longer leaves shard weights behind.
timings.jsonwas written from inside the CI-report block, which a setup abort also reaches — with the setup phase’s summary. The file that came out named setup scenarios, and a weights file naming them is worse than no file: those identities never appear in a suite run, so they absorb bucket load on behalf of scenarios that never run and skew the very split--shard-weightsexists to balance, silently. The write moved to the one site where the summary is the suite’s, pinned by a test that reproduces the old file. -
lower.rsstops threading the same three values through twelve functions.out,refsandsinkstravelled as separate parameters everywhere, and five functions —expand_macro,expand_step,expand_ref_step,expand_payload_step,finish_step— carried 8 to 11 parameters each behind individual arity suppressions. Adding one piece of lowering state meant editing five signatures and five call sites, which is the shape of change that drops a parameter at one site.Two bundles, both of them types that were already implied by the code:
Emit { out, refs, sinks }(the mutable outputs, always passed together and never independently),StepScope { step_ref, ctx, at }(what stays fixed for one authored step however deep expansion recurses), and a smallFinishedfor the four values that describe a step being completed.What was not done matters as much. The obvious refactor — hoist the state into a
selfand make the five methods — would have broken the reason they are parameters at all:resolve_inand friends take them explicitly so they remain callable while other state is mutably borrowed, and a method on&mut selfcannot be called whileselfis borrowed elsewhere. The threading discipline is load-bearing, so it stays; only the arity changes.Arity suppressions across the workspace: 13 → 6, with
lower.rsat zero. No behaviour change, and the 241 core tests say so. -
ADR-0002 now names the core’s hurl entry grammar, and a guard keeps it closed. “Adding an engine leaves
proef-corediff-empty” was true of engine-types and never of engine-syntax: the core does text surgery on entries — splicing[Options]in, merging anexpect:block’s asserts into the previous entry — so it has to find an entry boundary in text hurl will later parse. The worklist carried the gap for two rounds as “~290 lines of hurl grammar in core”, a figure that counted#[cfg(test)]fixtures.Measured: twelve literals across four files. Seven the core writes, four it recognises to find a boundary, and one it quotes — a hurl snippet inside a did-you-mean help string in
bind.rs, which generates nothing and parses nothing but drifts like any other copy. The four boundary recognisers are already one sharedpub(crate)set. proef’s own pack keys are shaped like option lines and are excluded by name rather than listed as sanctioned rows.The guard lexes whole files. The first version scanned line by line and so could not see a literal that spans lines — which is where a larger piece of engine syntax would naturally be written, and where the one entry above that nobody had counted was in fact sitting.
The amendment sanctions that set and closes it. Deferred with a named trigger — a second engine being scheduled — is moving the written half behind the seam, where the reading half already lives:
StepKindSpec::optionsexists precisely so an engine’s option spellings stay out of the core, and it covers recognising them only, soretry:,retry-interval:,delay:andvariable:are still core literals. Until a second engine exists that migration relocates seven literals that exactly one implementation will ever supply, at the cost of a public-API break.crates/proef-cli/tests/source_guards.rs(renamed fromstderr_hygiene.rs, which had not been only about stderr for two rules now) pins the set: a new token, or an existing one spreading to another core module, fails the test and names both remedies. A claim of this shape decays the moment it is only prose — this one already had, by an order of magnitude, in the direction that made it look worse than it is.
Fixed
-
A
file,…;body in aref:fragment resolves where its author put it. hurl resolves a file body against the directory of the file that wrote the reference — its--file-rootdefault, and the same rule Karate, pytest and Jest use for fixtures. proef resolved every asset against the feature, and a fragment lives in another tree entirely ([run] fragments), so the same bytes passed under stockhurland failed under proef, as exit 2, blaming the author for a path that was correct. Nothing worked around it: moving the file beside the feature breaks the standalone run ADR-0018 exists to guarantee, a reaching../path is refused by hurl’s own sandbox, and the advice that refusal prints — “check –file-root option” — names a flag proef does not expose.Each asset is now staged from beside the source that referenced it, feature or fragment, into that scenario’s own asset root, which is what the engine gets as its context dir. Staging rather than two roots because hurl offers one context dir per run of entries and no per-entry override, while a single batch may mix both body forms — measured, not assumed: an inline step and a
ref:step in one macro lower to one batch. Copying fixtures into the build output is the standard answer to exactly this, and it adds no copy operation: the record already copied these files once per scenario, just into a shared directory instead of the right one. What it does change is the footprint — an asset N scenarios read is now N files in the run record rather than one, which is the same fact as the collision below, seen from the disk’s side rather than the reader’s. -
Two scenarios’ assets no longer overwrite each other. Staging was flat and keyed by the asset’s bare name, so two features that each keep a
data.jsonbeside them staged to one file — last writer wins, with “0 warning(s)” — and the loser’s artifact replayed against the other’s bytes.artifact_slugalready refuses that trade for the.hurltext, deriving from the feature’s whole path so two same-named scenarios cannot collide; the files it reads now get the same treatment. An artifact that reads a file says so in its replay line (--file-root assets/<slug>); one that does not is byte-identical to before. Two sources claiming one name inside a single scenario — the case a per-scenario root cannot separate — is refused rather than narrowed.A missing asset is also an error now instead of a silent skip. It had to become one: the staged root is what the engine reads, so a file that quietly failed to arrive is no longer an incomplete record but a request reading nothing.
[Options] output:resolves through the same root, so the root is created for every scenario rather than by the staging loop — which never runs for a scenario that reads no file body. A response written that way now lands inside the run record, where a run’s outputs belong, instead of in the feature’s own directory.Breaking (library):
emit::file_referencesis replaced byArtifact::assets, aVec<AssetRef>carrying each reference with the source that wrote it — the provenance a whole-artifact text scan destroys, and the whole reason the bug was expressible.emit::asset_rootnames the staging directory for the three call sites that must agree on it, andpack::split_qualifiedis now the one reader of thefile.hurl#nameformFragment::qualifiedwrites — there were two, resolving aref:and ause:, and staging assets was about to make a third in another crate. New diagnostic:proef::run::asset_unstageable. -
A
--run-idrecord is findable again (ADR-0021).--run-id pr-1234writes a perfectly good record, and every command that resolves the latest run —explain,diff,flaky,report,--rerun— enumerated by the uuid shape, so that record was invisible to all of them.--rerunwas the sharp edge: it silently continued some older run instead of the one just produced.One predicate had been answering two questions whose risks point in opposite directions — may I delete this? is unsafe when broad, is this a run I can show you? is unsafe when narrow — so the deletion-safety choice had silently become a visibility choice. They are now separate: a directory is a record because it holds an
events.jsonl, while rotation still deletes only uuid-named directories, so a custom-id run is discoverable and still never reclaimed by[run] keep-runs. Ordering stopped riding on the name too — uuid-v7 sorted chronologically until a custom-id directory joined the set and sorted by its first letter — and now takes the timestamp a uuid-v7 name carries (48 bits of unix milliseconds, which is why the lexical sort worked), falling back to directory mtime for a name that carries none. -
A run with a failed
optional:step no longer prints exactly like a spotless one.ConsoleMode::Failed’s own doc comment states the requirement and theWarnedarm implementing it was unreachable: a scenario’s aggregate status was only everFailed | Skipped | Passed, so a real optional failure was invisible under--console failed, showed a.rather than the documentedwunder--console dotted, and left the HTML report’s warned count and its filter-bar warned button permanently empty — four consumers and three docs describing something that could not occur. Steps carriedWarned; scenarios never did. The aggregate now promotes, which changes what a run says and never whether it gates:Warnedcounts as passing in the exit code, the totals andJUnit. -
explainprints a step’s authoredname:, like its five siblings.step_label’s own doc enumerates the six surfaces that must render it — console, HTML,JUnit, TAP, the job summary,explain— andexplainwas the one that never called it, so the post-mortem tool showed one sentence repeated where the live console had told the steps apart. -
The HTML “Slowest” section no longer counts
[run] setup/teardowninto “% of run time”. Every other aggregate on the page excludes phases (ADR-0014), including the tag table directly above it, so the share meant something different in that one section. A slow phase stays visible in the timeline and in its own block. -
.cargo/audit.tomlno longer suppresses advisoriesdeny.tomldeliberately un-suppressed. It carried the quick-xml pair (RUSTSEC-2026-0194/0195) with a comment claiming it mirroreddeny.toml— which had removed them, precisely because the reason had expired (quick-junit0.7 moved to the patchedquick-xml0.41, which the lockfile is on). So the nightlycargo auditjob was suppressing for no reason, and would also have silenced any new advisory filed against that line. -
A merged report covers the whole suite again, and its headline agrees with its page. Two independent failures in
--reruncomposition, againstdocs/CI.md’s promise that “one report stands for the composed result”. The overlay followed only the immediatererun_of, so the ordinary fix → rerun → fix → rerun loop — the workflow the feature exists for — silently dropped everything from before the last link, with no banner saying so; the page just got smaller. It now walks the chain, newest verdict winning, with a cycle guard becausererun_ofis a string read out of a record and records travel. And the headline took its numbers from the tail totals, which belong to the re-run, so one page read2 passed · 0 failedabove a tag table summing to eight and a siblingJUnitsayingtests="8". The composed stream now declines those totals rather than inventing new ones, so the headline counts the scenarios actually rendered. -
Run metadata reaches the two ADR-0020 §5 consumers that never received it. The GitHub job summary — named in the ADR, and the page a CI reader actually opens from the job — carried none, so the commit under test was in the record and the HTML report but not there. And
diff --format jsoncarriedenvbut notmetadatawhile diff’s human output printed metadata differences, leaving the machine surface a CI gate reads missing exactly the context the ADR was written to provide. Still handed over, never harvested. -
The ADR-0002 grammar guard can now see the shapes the ADR names. The amendment claims the core’s hurl vocabulary is closed and pinned; the guard classified four shapes, and method lines — one of the four boundary recognisers the amendment’s own Measurement section names — was not among them. Teaching it surfaced one unenumerated token immediately:
GET ${url:base}/PATH, sitting inbind.rsin the same literal as the already-pinnedHTTP 200. The ADR’s table and the pinned set both now carry it, and the vocabulary is thirteen literals rather than twelve. The scan also stopped truncating at a file’s first#[cfg(test)] modand now excises every test module: production code placed after one was silently unscanned, and two core files already carry a second test module. -
--rerunon a truncated record no longer reports success over a suite that never ran.Record::scenariosis built fromscenario_finishedevents alone, so a run killed mid-flight — SIGKILL, OOM, a full disk, a container eviction — leaves its unreached scenarios absent rather than recorded. The candidate list built from such a record named nothing, the “no failures” branch fired, and--rerunexited 0 having executed no scenario at all.explainsaw the truncation the whole time;--rerundid not, and CI is exactly where truncation happens. The same class as the cancelled-run bug fixed in 0.14.0, which this code’s own comment describes.A truncated base inverts the question: not “what did the record say to re-run” but “what can the record prove finished” — everything else in the selected front runs, announced with a warning naming the truncation. That distinction now lives in a
RerunFilterpredicate rather than a list, because only the record reader knows which of the two questions applies. -
diffno longer reports a scenario skipped in both runs as “now skipped … (was passing)”. Both halves were false — it did not become skipped, and it was not passing — and it fired for every@skipscenario on every diff, including two runs of an unchanged suite, handing--format jsonconsumers the same wrong pair. The bucket exists for transitions (ADR-0019 §7); the guard makes that true of the code and not only of its name. -
A tab in a bound value is refused where every other control character already was. The lower-time guard exempted
\t, which hurl’svariable:grammar rejects like any other control character, so exactly one character kept taking the late path the guard exists to close — dying asemit::invalid_artifactagainst generated text the author never wrote, rather than as a refusal naming their ownbind:. -
--shard-weightsno longer piles every zero-cost scenario into shard 0. Costs are whole milliseconds, so anything sub-millisecond stores as0— routine for a fast suite — and adding0never moved a shard’s load, so shard 0 stayed the minimum forever. An all-zero weights file put the entire suite in one shard and left the others selecting nothing: the flag doing the exact opposite of its purpose, silently, with the partition still exact so nothing complained. Assignments are now a tie-break alongside load, which also gives the right answer when weights genuinely cannot separate scenarios: equal cost, equal share. -
A disk filling mid-run now reaches the exit code. A stdout that was already broken at start has failed loudly since the correctness series — but the human report’s own writes go through the console reporter, which swallows write errors (a reporter cannot report its own channel dying), so a disk filling during the run truncated the report while the run still exited by its verdict. The
Teeunder the reporter is the last place the failure is visible; it now latches the same stdout-failure flagoutln!uses, and the exit funnel turns lost output into exit 3. Same closed-pipe exemption as ever —proef … | headis the reader ending the pipeline, not a failure — and a stderr console (machine mode) does not claim stdout failed. Pinned by a three-case test, mutation-checked. -
The complexity guard added moments earlier was itself flaky, and now runs alone. It shipped in the ordinary suite on the reasoning that a ratio cancels out runner load. Measurement disagreed on its second full-suite run: 2.05× isolated, 3.09× under nextest’s full parallelism, against a bound of 3.0. The larger input has the larger working set, so memory-bandwidth contention penalises it more than the smaller one — the ratio drifts rather than cancelling, and interleaving the samples cannot fix a systematic effect.
nextest’s
test-groupsbound concurrency within a group and do not isolate one from the rest of the suite, so the only mechanism that actually delivers isolation is#[ignore]plus a dedicated invocation: a CI step of its own andjust perf. The samples are interleaved as well, which removes the one skew that ordering alone creates.TESTING-STRATEGY.md§7 previously asserted the opposite in as many words — that a ratio “survives a shared runner” — and is corrected with the numbers. The claim was reasoning, not measurement, which is the failure this whole section of the changelog exists to record.
Internal
-
A fixture that spells the record by hand can no longer drift off the schema.
explain’s truncated-record test wrote its stream as three JSON string literals, and all three had drifted: ascenarioscount on the head, aschemaon the body events, alineon the close.Eventcarries none of them. Nothing failed and nothing could — the reader has nodeny_unknown_fields, so a stale key parses cleanly and is dropped, and a fixture built to assert “a record holding one passed scenario” was three-quarters describing a format proef has never written. It is typed now, through the helpers its two neighbours already use.The class is closed by a sixth
source_guards.rsrule: every string literal in the workspace that parses as a JSON object taggedeventmust deserialize as anEvent, and every key in it must matter — a key is phantom when deleting it yields the sameEvent. Inertness rather than an inventory, so it stays correct through renames,#[serde(default)]andskip_serializing_if, none of which a key-set comparison survives. Substring assertions against records proef actually emitted ("event":"run_finished","passed":1) are skipped by construction — they are not objects, and they check the opposite direction. -
Each doc check now lives in the half of the gate that its own rule names.
tests/docs.rsholds the checks that need a built binary (they ask clap, rather than parsing help text into a model that could drift);xtask docs-checkholds the ones that read files. The changelog-heading check added moments earlier read one file and parsed headings, so it sat in the wrong half — and the cost was concrete rather than tidy: it never ran in the fast doc-only CI step, only under a full nextest that had to build a binary it did not use. -
A PR that changes source now has to record itself.
RELEASING.mdhas always said that every landed change adds an[Unreleased]line in the commit series that lands it, and nothing checked it — this very entry is the one that was missed. Measured before being written: across the previous 21 source-touching merges the rule would have fired exactly once, on exactly the commit that broke it, so the check earns its place by count rather than by argument.
[0.16.0] - 2026-08-31 (the surfaces tell the truth: an eight-wave improvement programme, and the round that found what it missed)
Supersedes 0.15.0, which was cut (
release: v0.15.0, 2026-08-25) but never tagged or published — its changes are all here, and crates.io goes 0.14.0 → 0.16.0 with nothing skipped.
Added
-
explain,diffanddoctorspeak--format json. They were the three commands with no machine output, and the three a consumer reaches for after a run. A run directory isartifacts/ + events.jsonl + run.logand carries no structured summary, so anything analysing a run it did not launch — a CI job reading another job’s artifact, a script, an agent — had to foldevents.jsonlitself. That is the fold proef’s own two internal copies disagreed on three ways beforereport::suite_totalsunified them; handing the canonical answer over is cheaper than inviting everyone to re-derive the one proef got wrong.Each object mirrors its prose field for field rather than modelling a richer view — the prose is the contract a reader already knows, and a machine surface that says something different is a second answer to one question.
diff’sflaky/slowerstay the rendered sentences for the same reason. The flag is the existing single-variantjsonenum the listing commands already use, renamed fromListFormattoJsonFormatnow that it serves non-listing commands too. Machine mode owns stdout: notes whose content the object already carries are suppressed rather than repeated on stderr.doctorneeded a real change to get there — it printed each check as it ran, so the verdict was the only thing a caller could see. Checks are collected before rendering now, which makes the JSON a second rendering rather than a second walk: the failure mode where one surface gains a check the other never learns about. -
--console failed— thefullBDD tree, but only for scenarios that failed or warned. A clean run prints the run line and the summary; a dirty one prints exactly whatfullwould. The gap it fills is the CI one:fullis a wall of green on a large suite,dotteddrops the detail you need when something breaks, andquietdrops everything.Warned scenarios are shown, which the name does not say and the code explains: a warned scenario is one whose
optional:step failed,RunSummary::passedcounts it with the passes, and the summary line has no warned column — so a mode that showed onlyFailedwould let a run in which something did fail print exactly what a spotless one prints. A fourth value on the existing flag rather than a new one. -
proef flaky --by <key>splits flakiness by run context.--by env, or any[meta]/--metakey (--by runner), folds the history per context instead of pooling it. A scenario that flaps in one environment and is solid in another is not flaky but context-dependent — the fix is in the environment, not the test — and a merged history cannot reach that conclusion, because pooled failures and passes look exactly like one flapping test. The command names the scenarios whose verdict changes with where they ran, which is the finding the flag exists for. A run that never set the key becomes its own(unset)bucket rather than being folded in with runs that did; the context also rides in--format json. Reads theenv/metadataprovenance the record has carried since ADR-0020 — no new recorded field. -
proef schema configpublishes theproef.tomlJSON Schema. TOML language servers (Taplo, tombi) validate against JSON Schema, so one file buys completion, hover documentation and typo detection in the config — before a run rather than after one. Generated from the same Rust model that parses the file, so it describes keys as they are written (runs-dir, notruns_dir) and inheritsdeny_unknown_fields, making an editor refuse exactly what proef refuses.proef schemakeeps printing the pack schema, so one command answers “what may I write in this file?” for both authored formats rather than two verbs answering it once each. -
An assertion that fails on values looking identical now says why. When the actual and expected values differ solely in whitespace, the failure carries a note repeating both with every whitespace character drawn —
·for a space,\t/\r/\nfor the usual escapes,\u{a0}for the exotic ones. hurl’s own message was already correct; the defect was simply invisible, so a trailing space, a CRLF fixture leaking\r, or a non-breaking space pasted out of a browser read as “the tool is wrong”. Taken from hurl’s structuredactual/expectedrather than parsed back out of its prose, and emitted per error so it sits beside the values it explains; silent whenever the difference is already visible. -
proef flakyaudits quarantine, which nothing else could. A@quarantinescenario’s failures gate nothing by design, so no exit code, no summary and no CI job reports them — which makes the tag’s own failure mode invisible: a quarantined scenario failing every run has been switched off and left in the suite. It now readsDISABLEDrather than sharing thebrokenverdict with untagged always-failures, which wrongly implies someone is watching. The opposite case gets its own verdict too: green throughout the window isrecovered, a tag that outlived its problem and is now suppressing the next real regression. Both print what to do, and--format jsoncarriesquarantinedplus the verdict key so a scheduled job can gate on either.This needed the record reader to stop dropping data it was already given:
scenario_finishedhas carriedtagssince 0.15.0, butScenarioRunnever parsed them, leaving every record consumer tag-blind. -
Document symbols and hover. A feature outlines to its scenarios (with their tags), a pack to its macros (with the pattern each matches) — the vocabulary chosen by what discovery found in the file, never by its extension. Hover answers the question go-to-definition charges a round trip for: what a step binds, what a
use:targets, what aref:resolves to and which of its variables still need abind:. Every fact is read from the same analysis the diagnostics come from, so a hover cannot contradict the squiggle on its own line.SuiteAnalysisgains ascenariosindex, taken from the parse rather than from binding — an outline that hid exactly the scenarios you are debugging would be worse than no outline. -
A panic no longer ends the editor session silently. Only the recompute was guarded, so a panic inside completion, definition or references escaped the message loop and killed the server — leaving an editor that shows nothing, which reads as “proef has no opinion here” rather than as a failure. Both entry points (a request, the debounced recompute) now wrap everything they do, the request is answered with
InternalErrorrather than dropped, and the user is told once per suite state throughwindow/showMessage— the channel an editor surfaces, unlike the stderr line that was the only report before. The next edit clears the report, because whether the new state also fails is news. -
The editor can apply a “did you mean”, not just print it. Every misspelled-name diagnostic that already suggested a nearest spelling now carries the structured half of that suggestion — a span and a replacement — and
proef lspserves it as aquickfixcode action:use:andref:targets,with:andbind:keys, step kinds, Examples placeholders, and data-table columns. The suggestion is computed once and rendered twice (prose for a reader, an edit for an editor), so the message and the fix can never disagree.A fix is attached only when the edit is certain: the suggested name is near enough, and the misspelling occurs exactly once, as a whole token, in the diagnostic’s own file. Each of those failing means no fix rather than an approximate one — notably, a lowering error anchors on the feature step that invoked a macro while the typo lives in the pack, so it finds nothing to replace and offers nothing rather than editing the healthy file. The action is reachable from either the diagnostic or the token, because the two are regularly lines apart: a
use:error carets the macro’s name key. -
README answers the comparison a prospect actually runs: a “When something else fits better” section maps raw
hurl(the exit stays open in both directions), Karate (choose it for embedded JS and whole-body fuzzy matching — the two mechanisms proef deliberately refuses; choose proef for one binary, deterministic reproduction, and files that run with no framework at all), and Postman/Bruno-class clients. The quick-start also points atproef initas the start that demonstrates theref:body form —tests/features/is deliberately fragment-free (the reference corpus is config-independent by design, and[run] fragmentsis a config key; the runnableref:demo lives in the scaffold, pinned green against the fixture). -
The docs site can get a visitor to a binary, and CI to a green workflow. New Installing page — install lived only in the repo README, outside the published site’s source, so the site’s first step sent visitors back to GitHub — and a new CI page with the paste-ready workflow the docs never had (zero
runs-onblocks existed anywhere): install, secrets viaPROEF_SECRET_*,--junit auto, a--shardmatrix,--metaprovenance, thediff --fail-on-regressionbaseline gate,--reruncontinuation, andflakyover retained records. Nav reordered visitor-first (Installing → Getting started → Writing scenarios). -
AUTHORING gains the three recipes every real suite needs: login-then-use-the-token (the docs’ most-asked absent question — zero “login” hits existed), waiting for an eventually-consistent result (finite
retry:as the polling primitive, and why it must be finite), and test-data seeding/cleanup across its three scopes (Background:,[run] setup/teardown,saveAs: global). -
Every release archive ships a
.sha256sidecar (basename inside, sosha256sum -cworks from a download directory). Attestation covers the provenance story forghusers; the sidecar covers everyone who installs withcurl— the half that was missing against the ripgrep/uv/starship baseline. -
A broken
proef.tomlis a located diagnostic, not a bare sentence. The file is edited as often as any pack, and it was the one authored input whose errors carried no code, no source excerpt and no caret — whilepack::yamlhad all three for the structurally identical failure. New codesproef::config::toml(with toml’s own error span under the caret) andproef::config::unreadable, in the catalogue (73 → 75) and pinned by an integration test;proef lsp’s boot warning anddoctor’s config row carry the same message. -
Five help-less refusals gained their missing action.
feature::parse(the shape of a feature file, and the most common way one stops parsing),bind::ambiguous_step(make one pattern more specific or retire the duplicate),bind::table_conflict(one source per param),pack::use_cycle(pull shared steps into a third macro), and the rawretry: -1message now says why infinite retries are refused (hurl cannot be interrupted mid-call) and what to write instead — it used to cite “ADR-0007”, an internal document id with no in-band route to it. -
A miss below the did-you-mean threshold names the valid set instead of going silent. All eleven suggestion sites ended
closest(…).unwrap_or_default()— when nothing was near, the tail vanished, andunknown_step_kindsaid “not claimed by any registered engine” about a registry with exactly one member it never named. Onematcher::suggest_or_enumeratenow serves every site: the nearest spelling when one is near, else the set verbatim (small), else a count with the command that lists it ((9 known —proef macroslists them)).unknown_fake,unknown_variableandmissing_config_varcarry the same rendered tail through their typed errors. -
unknown_placeholderfires once per authored defect, not once per Examples row — a 500-row outline with one typo’d<column>pushed 500 byte-identical diagnostics at one span (the console collapsed them; SARIF, one-result-per-site by design, did not). It also now names the header’s columns. -
Every rendered error links the diagnostics catalogue. The stable codes were greppable and led nowhere — the catalogue was linked from every doc and reachable from no error.
RenderedimplementsDiagnostic::url()and the LSP setscode_description, so editors show a clickable link on the code; on a terminal miette renders an OSC-8 hyperlink, and into a pipe or snapshot the URL prints as plain text beside the code (links ride the same TTY/NO_COLORgate as color — an escape sequence a non-terminal sink must never see). -
Pack diagnostics point at the defect, not the macro’s name. Every pattern-family and
defaults:error anchored on the macro-name span — thirteen of the nineteen seeded pack snapshots underlinedlogin:while the broken{rol}sat on a line outside the excerpt (one excerpted the previous macro). Thematch:-line span was computed since the pass was written and never reached a diagnostic; it does now, with the name span as fallback.locate::macro_spanalso stopped matching pack-rootbind:entries (a macro sharing a name with a bind key anchored every diagnostic on the config line). -
Parser errors speak hurl’s and gherkin’s prose, not Rust’s. A pack author was shown
ResponseSectionName { name: "Wrong" }andMethod { name: "" }—{:?}of internal enums from crates they never heard of. All three engine sites now render through hurl’s ownDisplaySourceError(“the section is not valid. Valid values are Captures or Asserts”), and gherkin’s expectation-set tail is sort-normalized: it renders from aHashSet, so the same broken file printed two different messages across processes (observed live) — breaking snapshot determinism and the duplicate-collapse alike. Pinned. -
A
resolve::*error names the pack it lives in. The span is the feature step (the invocation), but${nope}is written in a pack YAML the message never named — the reader was sent to a healthy.featureline while the sick file stayed anonymous. Every resolve error now carries(pack <file>). -
The HTML report is triageable, linkable, and filterable. Every scenario block carries an
id="s-<slug>"anchor (the samestem--nameslug as its artifact, so the two cannot disagree) — a failure is now a URL a colleague can be handed. A “failed:” jump rail under the summary links straight to each failing block (blocks keep completion order — the rail is how a reader skips the green between failures), and a status-filter bar (all/failed/skipped/warned) toggles block visibility through a ~15-line inline script: progressive enhancement over classes the blocks already carry, no framework, still one self-contained file. Snapshot reviewed deliberately. -
--watchreads like an inner loop. A visual rule with a rerun counter separates iterations (twenty edits used to stack twenty trees with nothing marking where the current one begins), and the post-run line says the verdict in words (“failures — details above”) instead of an exit number to decode. -
Shell completions and a man page, generated by the binary itself: hidden
proef completions <shell>(bash/zsh/fish/powershell/elvish) andproef mansubcommands, and every release archive now carriescompletions/plusproef.1— generated during packaging by the exact artifact they ship beside, so they can never drift from it. -
--envis global, like--config:proef --env staging testandproef test --env stagingboth work — five commands read the profile, and the position-sensitive spelling was a lesson nobody needed. -
doctorexamines the project, not just the engine: suite resolution (feature-file count, or the failure),hurlon PATH (a warning when absent — the engine is embedded, but ADR-0018’s stock-replay promise and the emitted# replay:hints need the binary), and runs-dir writability (probed with cleanup — the first-runcreate_dir_allfailure was invisible to the one command whose job is diagnosis). -
A typo’d
--tags/--scenarionames the nearest real spelling. The refusal held every scenario name and tag at the moment it printed “check –tags/–scenario” and used none of them; it now suggests the closest name and tag (glob atoms excepted — a glob selecting nothing is a fact, not a typo) and points atproef flows, the treatment[run] exclusive-tagsalways had. -
proef fragmentssays why a listing is empty when no[run] fragmentsroot is configured — previously indistinguishable from a configured-but-empty corpus, though the reader’s next move differs. -
The console speaks in color, and every run ends on its identity. The status vocabulary (
✓/✗/∅/⚠, the dotted glyphs, the summary’s verdict half) is ANSI-colored on a terminal —NO_COLOR, a dumb TERM, or a non-terminal stream turns it off, and therun.logmirror strips the paint either way (content verbatim, paint never). Color is paint on identical bytes: the record, the exit code and every text assertion see the same output. Each run’s final stderr line is nowrun <id> · <seconds>s— the run id is the reproduction key--shard,--shuffleand${fake:…}all hang off, and it previously printed only at the top of the scrollback; a red run’s trailer adds theproef explainpointer. Wall-clock stays console-only, never entering the record.
Changed
-
Secret redaction no longer runs inside the reporter mutex. The sink masked each event while holding the lock that fans it out to the reporters, so every scenario thread queued behind work none of them share — and masking is the expensive half, a scan per text field per needle with roughly nine needles derived per secret. It reads the event and the needle set and writes neither, so it never needed the lock; the critical section now covers only the fan-out it exists for.
Order is unaffected and the tests say why: a scenario is one thread, so its own events still reach the lock in the order it emitted them, and order across scenarios was never guaranteed. A new test emits from eight threads at once and asserts nothing is lost or doubled, everything arrives redacted, and each emitter’s own events keep their order. No timing assertion — the flake rule forbids one, and the change is justified structurally rather than by a stopwatch.
-
Pack validation is linear in the macro count, not quadratic. Every span locator scanned the whole pack file to find its macro’s block, so validating N macros scanned the file N times. A single indexing pass (
locate::MacroIndex) records each macro’s name span and block region, and the locators became lookups into it. Measured on a release build over generated packs: 3200 macros went from 1.96 s to 0.03 s (~65×), and the curve changed shape — 4× per doubling before, ~2× after — so 6400 macros now cost 0.06 s where the old scaling predicts ~8 s.It also fixes an inconsistency the split readers hid:
macro_spanaccepted a quoted"macro name":header while the region scan behind every other locator accepted only the bare form, so a quoted macro got a caret on its name and silently no span for itsmatch:,use:,ref:or payload lines. One reader now gives one answer. -
The editor stops re-analysing the suite on every keystroke. Completion, go-to-definition and find-references each ran the whole pipeline from scratch — read every pack and feature off the provider, parse, bind, lower — and threw the result away; between two keystrokes none of those inputs have changed, so the second run could only reproduce the first one’s answer. The server now holds the analysis and drops it exactly where an edit lands (the same notification path that already marks the suite dirty), so one recompute serves the debounced diagnostics publish and every request until the next edit. Measured on the two-file test suite: 10 provider reads per request before, none between edits after — pinned by a read-counting provider rather than by timing, per the flake rule.
-
The LSP’s type layer moved to the maintained generator:
lsp-types0.97 (unmaintained since; the crate that shipped its ownfluent-uriUrinewtype) is replaced bygen-lsp-types0.11 under the samelsp_types::name — rust-analyzer’s own aliasing pattern, so everyusepath is unchanged. Itsurlfeature aliasesUritourl::Url, which the embedded hurl engine already pulls in, so the swap adds no new crate and drops three (lsp-types,fluent-uri,serde_repr).Url::from_file_path/to_file_pathare the native-path bridgedocuments.rshad to hand-roll under 0.97 — drive letters, segment joining, percent-encoding, ~90 lines — so the bridge is now a wrapper that only pins the source-name identity rule. Behaviour visible to an editor is unchanged; the one difference is what counts as a malformed URI (urlpercent-encodes a raw space wherefluent-urirejected it), and request dispatch now compares a method enum rather than strings, so an unknown method lands inCustominstead of matching nothing.Breaking (library):
proef_core::report::percent_encodeis private. It was public solely soproef-lspcould encode URI path segments against the identical unreserved set; that hand-rolled encoder is gone, and redaction needles — its only remaining caller — live in the same module.
Fixed
-
The record-size ceiling reached two of its four readers. 0.13.0 bounded the run-record read at 256 MiB because records travel —
diffreads a downloaded baseline,flakyreads every retained run — and the read, the line split and the parsedVec<Event>are resident at once, so a corrupt or hostile file was an OOM rather than an error. The bound lives inrecord::read_events, andexplainandreporteach openedevents.jsonlwith a bareread_to_stringinstead, so neither had it.reporteven used the guarded reader for the base record two dozen lines below the raw read of the primary one.Both now go through
read_events, which returns the parsed events — exactly the read-once/parse-once its own comment asked for. A source-scanning test makes the next reader go through the same door, the shape this project already uses for the raw-print and malformed-plural rules: a guard added in one place and left for the next call site to rediscover is how it went missing the first time. -
cargo denyfailed on a yanked transitive crate.rand 0.10.2resolvedchacha20 0.10.1, which was yanked from crates.io; the lock now takes0.10.2. Not the secret store’s copy —chacha20poly1305pins0.9.1, which is unaffected — so nothing about encryption changed. Found by the gate, which is what it is for. -
proef report -owrote the author’s home directory into the file built to be shared. With the report inside the run dir the artifact links are a bareartifacts/…; with-opointing anywhere else they were made absolute, which resolves only on the machine that produced them — and-oexists to put the report somewhere it will be published, which is exactly where that path is dead. 0.13.0 scrubbed machine identity out of the run record (R12-1); this put it back, twelve times over, in the HTML uploaded beside it. The href is now relative to the report, which resolves everywhere the absolute one did plus wherever report and artifacts travel together, and in the CI shape (-o public/report.html) names nothing outside the workspace. The href is built from path components joined with/, not fromPath::display— Windows renders\, which is not a separator in a URL, so a Windows-generated report’s links would have been dead either way (the absolute path it replaces had the same flaw). A report written somewhere sharing no ancestor with the run dir still names the directories between them — that is what a correct relative path from there is, and it is no worse than what it replaces. -
The report’s
--skipcolour failed WCAG AA, and every status pill failed it in dark mode.--skipwas the one palette token the dark block did not redefine: a grey chosen against#0d1117(5.48:1 there) left carrying white text on white at 3.45:1, against a 4.5:1 threshold — on the status a reader scans for after an interrupted run. It is now#59636e(6.11:1).Writing the guard rather than the fix found a second defect nobody had measured:
.pillpaintedcolor:#fffon the status colour, and the dark palette’s colours are tuned as text on a dark ground, so all four dark pills sat between 2.52:1 and 3.45:1. The pill foreground is now a palette token — white on light, the page ground on dark — putting all four between 5.48:1 and 7.5:1. A test asserts the ratio rather than the hex, so a future palette change is free to move a colour and not free to move it below AA, and a second test pins that both palettes define the same token set (the absence that caused this). -
The HTML report had one heading and no outline. The timeline carried an
<h2>; the tag table and the scenario list — the body of the page — had none, so there was nothing to navigate by and no anchor to link a section with. Both gained one, sharing the class the timeline already used (renamed from.timeline-hto.section-h, since it now serves three). Pinned structurally, so a section added without a heading fails the test. -
A step’s
name:label reached the artifact and nothing else. A macro with more than one step turns one feature sentence into several engine steps, and they share aStepRefexactly — same file, same line, same text. The emitter has always written the authoredname:into the artifact’s entry comment, which is why the.hurlcould tell them apart;StepRefnever carried it, so the console, the HTML report,JUnit, TAP, the job summary andexplainall printed the same sentence once per step, with nothing but the status glyph to distinguish a warning from the failure beside it. The reference corpus demonstrated it: threestep_finishedevents for the cookie session is exercised, byte-identical in the pinned snapshot, are nowobtain the session cookie,optional probe (forces a split)andcookie survives the split.StepOutcomeandstep_finishednow carrylabel, exactly as they carryfragment— the two answer neighbouring questions (which file did this request come from / which step of the sentence is this) and travel the same channels. Oneproef_core::report::step_labelrenders it for every sink, so the six cannot drift. Additive on the wire: absent when a step has noname:, so every pre-existing record still parses and re-renders unchanged, and the event schema stays1.This retires two claims that were not true when written:
AUTHORING.md’s “they anchor artifacts, events, and failure output” andLoweredStep::label’s own “(events/console)”. Same class asreproduce_hintin the R18 wave — computed all along, printed all along, dropped by the record. -
A fragment’s text ran on into the comments introducing the entry below it. hurl attaches the blank and comment lines above a request to that request, which is exactly what makes the
# @proefbinding reliable — but it also means an entry has two different starts: where its lines begin and where its request begins. The scanner used one value for both, ending each fragment at the next entry’s request line, so every comment a corpus author wrote to introduce the next request was copied into the previous fragment and from there into the emitted.hurl. An artifact could carry# Destructive. Operators only.while containing no destructive request at all, andtrim_endcould not help — a comment is not whitespace. The same applied at the end of a file, where a trailing note became part of the last fragment. A fragment now runs from its annotation to the end of its own request and response; the gap between two entries documents the one below it and belongs to neither. Nothing executed differently, because hurl permits only comments and blanks between entries — which is why it survived: the only damage was to what the durable record says a request is.The property covering this asserted one request line per fragment, which is blind to comments; it now also asserts that no fragment holds any of the generator’s inter-entry filler.
-
explainand the HTML report disagreed about a truncated run’s totals. A record with no tailrun_finished— a run killed mid-flight — is reconstructed by counting, and each surface carried its own version of that fallback. On the same bytes they differed three ways: the report droppedWarnedscenarios from every column, counted[run] setup/teardownscenarios into a headline its own page labels “excluded from totals above”, and read a pre-0.6.0 record’s per-phase totals as the suite verdict whereexplaincorrectly declined to. Oneproef_core::report::suite_totalsnow holds the rule — prefer the tail event unless it cannot be trusted, else count suite scenarios withWarnedriding along withPassed, exactly as the live path reports — and both surfaces call it.Also un-splices three doc comments in
html.rsthat an earlier change had merged into one, leavingrender_tag_tableandrender_timelineundocumented andrender_provenance_and_summarycarrying all three. -
A parse error pointing at a non-ASCII character produced a span that split the codepoint. gherkin reports a char-counted column, so the span’s start was correct; its end added one byte to that, landing inside a multi-byte character whenever the error pointed at one — a span that is not a valid slice of its own source. Nothing crashed, which is how it survived: miette tolerated it and drew the caret slightly to the left, and the LSP’s converter snaps to a boundary defensively, so every consumer defended itself instead of the producer being right. Found by the new
fuzz_feature_parsetarget within a minute of first running. -
A long
--tagsexpression aborted the process instead of failing.and/orchains parse iteratively, and the module said so as though that settled it — but an iterative parse still builds a left-leaning tree as deep as the chain is long, and bothevaland the derivedDropwalk that tree recursively. A--tagsexpression of roughly twenty thousandand-joined atoms therefore overflowed the stack and died on SIGABRT: a signal, not one of the four exit codes ADR-0009 promises, and well within what a command line accepts. Expressions are now capped at 512 tokens, which bounds the tree and so bounds both walks, and past the cap you get a message naming the limit. (The test that was meant to cover this built 5 000 atoms and asserted success — one order of magnitude below the cliff.) -
EDITORS.md no longer under-promises on built-in macros. It said the
expect*family has “no jump target and no hover”; the first half is true and structural (their pack is compiled into the binary, so there is no file to open), the second is not — a built-in is in the analysis like any other macro, so hover answers with its pattern and params and names the pack asbuiltin:…, which is exactly why the jump is unavailable. Pinned by a test, since the page now claims it. -
The tutorial’s
ref:invitation no longer self-destructs. §3.6 showed a second[run]table that, pasted beside §3.5’s, was a TOML duplicate-table error naming a directory the tutorial’s layout doesn’t have; the fragments key now lives (commented) in §3.5’s one config block. “A suite is two things” undercounted its own mandatoryproef.toml— it says three files now, and the tree shows all three. TROUBLESHOOTING stops listing hurl’s[Options] repeat:as if it were a proef step key. -
The
proef initscaffold goes green against the dev fixture. The advertised fastest path (init→ fixture →test) ended 1 pass / 2 fail: the scaffold calls/searchand/version, and the fixture served neither — a red first run that read as a broken tool. Both routes exist now, the whole path is pinned by an integration test, and the scaffold’sref:fragment thereby executes against a live endpoint — the body form’s first runnable demonstration. -
A failure no longer prints its detail twice. An engine fault quotes the failing step’s own detail, and the located step line just below printed the same ~200 characters again; when the fault message contains a failing step’s detail, the fault line now keeps the scenario identity and the step line carries the detail once.
-
JUnit failure and skip detail reaches every platform. The detail — assert diff, fragment provenance,
@skip:reason, the quarantine notice — lived only in themessageattribute; GitLab parses only the element text, and Azure maps the text to its stack-trace field, so half the platforms showed a bare failure (or a reasonless skip). Every non-success now carries both, and a failure’s text node additionally carries each failing step’s redacted reproduce hint — the content channel has the room the one-line attribute does not. Pinned alongside two library guarantees that were verified rather than assumed: quick-junit strips ANSI escapes and XML-1.0-illegal control characters on every setter (one binary response byte used to be the classic whole-report killer on Jenkins/GitLab), andtimeis plain three-decimal seconds; both now have tests so a dependency bump cannot shed them silently. A third pin: composed reports (suite + rerun-carried + teardown) yield eachclassname+nameidentity exactly once — GitLab silently drops duplicates. -
The GitHub job summary can no longer vanish at the 1 MiB cap. The documented failure mode at GitHub’s limit is silent disappearance (and oversized writes have aborted jobs in shipped first-party actions); a failing rerun-overlay suite with per-tag tables crosses it more easily than it looks. The summary now truncates deterministically at a line boundary under a 900 KB budget, saying how many lines were cut and where the full detail lives.
-
::errorannotations budget for GitHub’s real limit. GitHub keeps ten error annotations per step and silently drops the rest — an uncapped emission made a forty-failure run look like exactly ten. The budget is now one annotation per failing scenario (its first failing step with detail, else its fault) capped at ten, with a closing::noticenaming what the ten are out of;title=is clipped under GitHub’s 255-character cap before encoding. -
saveAs: globalrefuses a secret it can find, not just a secret it can equal. The gate lived in the hurl engine and matched whole-value equality against raw secret values — a capture merely containing one (Bearer <token>) or carrying an encoded reflection (base64/hex/percent/ JSON-escape) promoted to.proef-state.jsonin plaintext. The refusal now lives on the store’s owner (World::set_global), armed once per scenario by the runner with the same derived-needle set redaction uses (ADR-0005) — one needle list for both invariants, and every engine a scenario dispatches to is covered. The invariant is now genuinely property-tested (any composite carrying a guarded secret never enters the store), as CLAUDE.md had claimed of the single example test. -
The SLA gate honors
@quarantine.sla::checkmeasured every scenario while the exit code excluded quarantined ones — so a quarantined, timing-marginal scenario (exactly what gets quarantined) could not fail the run on its assertions but still turned it red on latency. The latency population now applies the same non-gating list as the exit code. -
A record that travels can no longer lie, crash, or steer. Reading a record predating
scenario_finished.file(or any foreign baseline whose closes key under the serde default""), the step buffer never attached: every scenario read as step-less,flakycould never see a retry or a duration, anddiff --fail-on-regressioncertified green over empty step maps — the close now adopts its steps’ file when exactly one pending scenario matches by name (pinned by test). The head fold’s “first head wins” guard tested emptiness rather than position, so a secondrun_startedin a concatenated or legacy record overwrote the run’senv/metadata/rerun_ofwholesale (pinned by test).rerun_of— a string read out of the record — was joined onto the runs root unvalidated, so a crafted"../../elsewhere"spliced a foreign file’s events into the rendered report; it must now be a single path component, and--run-idgets the same rule at the CLI edge (a typed clap error on separators or.., on all four commands that accept one). Record reads gained a generous 256 MiB ceiling — the one input loaded with no bound — and every duration sum over record-suppliedu64s (HTML report, tag table,flaky) is now saturating instead of a debug-build panic on a corrupt file. -
A
[tag-links]template can no longer be subverted by a tag’s spelling. The GitHub-summary sink substituted the tag into the URL raw, so@JIRA-1)[x](yclosed the markdown link early and injected content into the job summary; the tag is now percent-encoded in the URL slot. Both sinks (HTML report and summary) also render non-http(s)templates as plain text rather than mintingjavascript:-class links. -
An inverted
Spandegrades instead of exploding:Span::lenand the SARIFbyteLengthare saturating — B1’s shipped class, closed in the type rather than at one construction site. -
Eleven sites that swallowed an error and reported success now speak. The class the v0.6.0–v0.8.0 series was named for, still present at the edges: a poisoned store lock silently skipped persisting the World (every
saveAs: globalpromotion of the run lost — now recovered, matching the runner’s own policy, which also stops failing an innocent scenario for another thread’s panic); an unreadable subdirectory silently shrank the suite to a confident “0 failed” (now warned, per entry too);doctorreported a clean “no packs” over a tree it could not read (now a Fail row) andfmtformatted nothing while reporting success (now warned); a non-UTF-8 environment value read as “not set” — the wrong cause — for${env:…}(now named up front); a.map.jsonserialization failure was the one silent write in the run record (now warned);proef lspbooted with defaults over aproef.tomlthat exists but does not parse, silently diverging from the runner (now says so on stderr); one unreadable run aborted all ofproef flaky(now skipped and counted, with the two-run floor re-applied over what was readable);xtask docs-checkprinted “aligned” when it could not read the directories it checks (now a failure); and a mis-severitied diagnostic pushed into the lowering error sink vanished entirely (any error-sink entry now fails the scenario). -
fmtnormalizes every literal-block spelling. The scan required the key line to end with|, sohurl: |-,|+, an indent indicator, or a trailing comment — all loadable — were silently skipped and--checkcertified them canonical. Folded scalars (>) stay out deliberately: YAML folding rewrites the line structure there is nothing line-preserved to normalize. -
A Ctrl-C landing in
--watch’s debounce window no longer launches one more full suite run. The ≥300 ms drain between “change detected” and the rerun never checked the interrupt, and the rerun then minted a fresh cancellation token — so the handler cancelled the finished run’s token, printed “leaving watch”, and a whole suite executed anyway. The interrupt is now checked inside the drain and again after the new token is stored, so a Ctrl-C from any point forward cancels the token the run actually carries. -
--watchcan no longer go silently deaf. A delivered watcher error and notify’s rescan signal (the kernel-queue-overflow event agit checkoutburst produces) were both discarded by the event filter — the watch kept printing “watching … for changes” while missing every change. Both now retrigger a run, saying why. Two adjacent silent paths gained voices too: a runs dir whose path has no final component now warns that its writes cannot be excluded from the watch (the self-feeding-loop shape), and a failed Ctrl-C handler registration now says the two-stage interrupt is unavailable instead of silently dropping the contract. -
A comment on a section header no longer blinds the scans that gate on it. hurl’s own
section_nameparser leaves the rest of the header line to the ordinary comment terminator, so[Options] # tuningis a real section — but proef’s scans required whole-line equality. Behind a commented header, validation pass 6 was off entirely:retry: -1dry-ran clean (the abandoned-thread hole ADR-0007 exists to refuse), the delay cap and the double-declaration check with it, in inline blocks and fragments alike. The same equality bug made[Captures] # idsdrop every capture under it from.map.json, and[Asserts] # noteopen a second section under anexpect:merge. Oneis_section_headerrecogniser now serves every section scan. -
delay: 5his refused likedelay: 90malways was. The duration table knewms/s/mbut not hurl’sh, so an hour-spelled delay five times over the 1-hour cap fell through the suffix parse and validated clean. The table now mirrorshurl_core’sDurationUnitin full. -
A pack-scope
bind:value resolves in the pack’s scope, not in whichever macro reached it first. The table resolved through the first ref-using macro’s argument scope and was then cached for the scenario — a bare${param}silently took that macro’s value everywhere (or vanished, blaming an innocent macro). The pack table now resolves arg-free and default-free: namespaced references (${url:…},${vars:…},${secret:…},${fake:…},${env:…}) are its vocabulary, and a bare${name}is a deterministic error attributed to the pack’s ownbind:in every macro order. -
A star-heavy tag atom can no longer hang selection or abort the process. The glob matcher was naive recursion: backtracking was exponential in the
*count (a 19-character atom took seconds per tag per scenario) and recursion depth grew with pattern length (a long enough atom in--tags,[run] exclusive-tagsor[tag-links]overflowed the stack — SIGABRT, outside the exit contract). Rewritten as the standard two-pointer match: linear-ish, iterative, oracle-property-tested against the old semantics. -
A bound value carrying a lone
\ris refused at lower time.lower::multiline_bindtested\nalone, so a carriage return (a value read off a CRLF file) sailed into the emitted[Options] variable:line and died one stage later asemit::invalid_artifact— blaming generated text the author never wrote. The guard now refuses any control character except tab. -
An
HTTP2-Settings:request header no longer mis-slots anexpect:merge. The last-entry scan recognised a response line by the bare prefixHTTP, which the emitter’s own recogniser was already hardened against; both now share oneis_response_line(HTTP/HTTP/).
Breaking
--outputsplit by meaning:--formatchooses a format,-o/--outputnames a path.testtakes--format json|tap; the listing commands (flows,macros,fragments,flaky) take--format json— each through its own enum, so clap’s help can no longer advertisetapon four commands whose runtime rejected it (the old shared enum lied about a quarter of the surface, and-ochanged category between siblings: format on five commands, directory onartifacts, file onreport).--output json/--output tapno longer parse on those five commands — clean break, no alias;artifacts/reportkeep-o/--outputfor their paths, unchanged. The runtimejson_onlycheck is deleted: the type system does its job now.- Library:
World::set_globalreturnsbool(#[must_use]) —falseis a refused promotion — andWorldgainsguard_secrets;Redactionsgains thetaintsprobe. The hurl engine’s private equality-only gate is deleted in favor of the World’s. - Library:
ConsoleReporter::newtakes a fourthcolor: bool— the TTY/NO_COLORprobe stays at the CLI edge; the sans-IO core takes the answer as a plain value.
[0.15.0] - 2026-08-25 (the Robot Framework audit: visible skips, tag verdicts, explicit metadata)
Breaking
- A quarantined test-failure reaches JUnit as
<skipped>with a message, not<failure>— Jenkins marked builds UNSTABLE while proef exited 0; every dashboard now says what the exit code says (ADR-0019). Library:ScenarioSpecgainsskip,ScenarioOutcome/ScenarioRungainreason,Event::ScenarioFinishedgains additivereason,write_junittakes the non-gating list. --shardassignments re-deal: the hash gained a mixing finalizer. Raw FNV-1a’s low bit is the XOR-parity of the input bytes, so a scenario named after its feature file — the commonest Gherkin convention — collapsed to one shard atN=2and left odd shards empty atN=4, silently (the empty shard exits 0).shard_bucketnow finalizes with Murmur3’sfmix64; every scenario re-buckets, so all jobs of one matrix must run the same proef version (already true in practice). Round-18 finding, reproduced and mechanism-verified before fixing; the balance test gained the name-mirrors-file corpus it was structurally blind to.- Tag atoms glob.
*and?in a--tags/[run] exclusive-tagsatom are now anchored wildcards (@FRD-*selects the family;?is one character) — previously they were literal characters that silently matched nothing, the trap this closes. Metacharacter-free atoms are bit-identical to before, property-pinned. Case stays sensitive. - JUnit test identity is
classname+name.classnamecarries the feature file,namethe scenario alone; the old singlenameembeddedfile:line, so an edit above a scenario re-identified every test below it in Jenkins history and GitLab’s MR diff. Anything keyed on the oldfile:line namestrings must re-key. The suiteskippedcount is now spelledskipped(wasdisabled, which no consumer reads).
Added
-
[tag-links]turns tag cells into tracker links (RF’s--tagstatlink, reduced to one mechanism): tag glob → URL template with{tag}substituted, honored by the HTML report’s by-tag table and the GitHub summary; the pattern language is the same anchored glob--tagsuses. Library (Breaking):render_htmltakes the link map;tags::atom_matches_publicexposes the one matcher. -
--console dotted|quiet(RF wave 3): one glyph per scenario (.pass,Ffail,sskip,wwarn — lowercase is non-gating, the pytest/RF convention, flushed per glyph, wrapped at 80) or just the frame. Purely presentation: the record, every report, the post-pool failure details and the exit code are identical in every mode;run.logmirrors the console verbatim, dots included —events.jsonlis the full truth. Library (Breaking):ConsoleReporter::newtakes aConsoleMode. -
A
--rerunnow produces the one JUnit and the one report that cover the whole suite (E2’s rerun half; Robot Framework’srebot --mergeshape, done as composition): the run head recordsrerun_of, the JUnit carries the base’s not-re-run scenarios as ordinary testcases, andproef reportoverlays the base into a merged page (banner named, base timestamps stripped so timelines never mix, rotated-away base degrades loudly). Exit code and totals stay the rerun’s own. -
--meta key=valueand[meta]/[env.<name>.meta]record explicit run metadata (ADR-0020, RF wave 2): commit, build URL, team — recorded in the run head, shown by the HTML report, GitHub summary,explain,diff(which now also warns on cross-env comparisons) and the--output jsonbody (additive keys). The active--envprofile name and the--shufflemarker ride the same head. proef never harvests: no git, no hostname, no CI env sniffing — the shell harvests, proef records. Everything passes the sink-boundary mask, keys and values both. Library (Breaking):RunRecord::openandexec::executetake the head inputs. -
Per-tag verdicts in the HTML report and the GitHub summary (RF wave 2): tags now reach the record — additive
tagsonscenario_finished(finished-only: the cancel-skip path emits no start), additiveexclusiveonscenario_started(closes R11-6, the scheduler’s own bool) — and both reports roll them up per tag (suite-only, Warned counts with passed). Requirement-tagged suites (@FRD-3.1) get their traceability matrix for free. Tags are deduped at the one accumulation point (first occurrence wins); the quarantine list is now derived from the outcomes’ own tags — one owner, same behavior, pinned by the exit suite. Library (Breaking):ScenarioSpec/ScenarioOutcomegaintags. -
@skipand@skip:<reason>park a scenario visibly (ADR-0019, RF wave 2): counted in every total, reasoned in the console, JUnit, TAP, the record, the HTML report,explainandflows --output json; the harness maps it to libtest’s ignored flag. All-selected-skipped exits 0; the empty-selection refusal stays exit 2.--tags "not @skip*"unselects both spellings; an authored skip is never re-queued by--rerun, anddiffgives skip transitions their own bucket instead of reading them as fixed. -
flowsshows the feature description. The prose block underFeature:was parsed and then dropped — the one paragraph written for exactly the readerflowsserves never reached them. Human output prints it under the feature header;--output jsonrows gainfeatureDescription: string|null(additive). Library:FeatureFilegainsdescription. -
--shufflere-deals the execution order, seeded by the run id — one determinism knob for order and fakes alike, so--shuffle --run-id <id>reproduces an order-dependent failure exactly (Robot Framework’s--randomize, minus the parallel seed it threads separately). Applied after--shard, so membership never moves; under--watchevery unpinned rerun re-deals, deliberately. The permutation is version-stable and pinned. Recording ashuffledmarker in the run head is deferred to the plannedRunStartedadditions (env/metadata), one wire change instead of two. -
The failing step’s
reproduce: curl …reaches the record. The engine always computed the redacted curl and the live console always printed it — and the record dropped it, soexplainand the HTML report knew less than the console did.StepFinishedgains additivereproduce_hint(absent on passing steps and every pre-field stream);explainand the report print it; the sink-boundary mask covers it likedetail. -
README documents every flag the binary exposes, enforced. v0.14.0 shipped
--shardand--max-failwith no README mention; the docs gate gains the flags direction (same vacuity guard as the command half), and the measured gap — those two plusschema --add-to— is closed. -
JUnit carries what GitLab and Jenkins actually read (R3-6, specced from GitLab’s parser docs and Jenkins’
SuiteResult.java):fileon each testcase (GitLab source linking),timeon suite and root.timestampandhostnamestay absent deliberately — ignored or substituted by both consumers, and a hostname would undo R12-1’s provenance fix. -
The docs corpus is a website: https://emrecdr.github.io/proef/. mdBook renders
docs/on every push tomainthat touches it; the nav isdocs/SUMMARY.md, which the existing docs gates link-check like any other doc, and the pages workflow refuses a corpus doc that is not on the site. The cratehomepagepoints there from the next release.
Fixed
-
A failure detail is bounded before it reaches any sink. hurl’s rendered assert error quotes the actual response, so a failed assert on a large body rode full-size into the record, JUnit, the HTML report and the GitHub summary at once. The engine now middle-cuts past 40 lines / 8 KiB with a marker naming the elision; the artifact pointer survives outside the cut, and the full output is one re-run away (Robot Framework’s 40-line rule, adopted at the boundary where all sinks are covered at once).
-
The machine-body contract closes its last two paths: an empty selection (
--scenario/--tagsmatching nothing — loud exit 2 by design) and a corrupt global-state file both emitted zero stdout bytes under--output json. -
Identical errors collapse like identical warnings — a broken macro usually fails to lower everywhere, so the error wall was the more common fifty-block wall; distinct errors still render separately, and SARIF keeps every site.
-
Injected
[Options]lines respect every section-ending shape. The section-end move covered one shape of five: an unfenced JSON/XML body after an author[Options]swallowed the injected lines into invalid hurl (exit 2 on input that worked before), and an entry with an author section but no response line leaked its pending lines into the next entry, where hurl parsedretry:as an HTTP header and the artifact validated green. The section now ends at the first line that could not sit inside it. -
A
#inside abind:value no longer hides the reads after it. The template probe parsed the value in an unquoted position where#opens a comment; it now probes the quotedvariable:position bake actually injects into, so"{{a}} # {{b}}"reports both. -
A setup that fails to load still emits the machine body — the last terminating path returning zero stdout bytes under
--output json. -
SARIF keeps one result per site again. The warning collapse shipped at the front-end aggregation, which also feeds SARIF — a code-scanning consumer lost every anchor but the first. The collapse now happens at console rendering only; SARIF carries all sites, the console one line with the count.
-
Every terminating path emits exactly one machine body (R17-2.3/2.4). An empty shard printed its prose note as the
--output jsonbody —jqfailed on the very path a sharded matrix guarantees one job takes — and a setup abort printed nothing at all while JUnit carried the failure. The note now goes to stderr under machine output (and lost a stray-space run); never-ran paths report ADR-0014’s suite-only zeros with the exit code carrying the verdict. -
A failed teardown reaches JUnit as its own suite (R17-2.5) — #78’s rule made symmetric: a phase appears in the reports when it fails. A gated pipeline used to read a fully-passing report on an exit-3 run.
-
A repeated warning is one warning with a count. One authored mistake in a macro shared by fifty scenarios rendered fifty times; identical warnings now collapse to their first occurrence plus “(N sites across the suite)”.
-
explain’s truncated-record fallback counts the suite only — a record that died mid-setup folded the phase scenario into the totals three lines above the label saying phases are excluded (ADR-0014). -
bind:values are read by hurl’s parser, not a text scan (R17-2.2). A hurl function ({{newUuid}},{{newDate}}) no longer counts as an unbound variable — proef refused input stock hurl runs — and a sibling literal bind whose name sorts earlier now counts as a supplier, since injected[Options] variable:lines are written and evaluated in name order. The seam answers the question once:FragmentSupport::template_reads. -
A fragment’s own
[Options] variable:lines now evaluate before the injected ones. Injection used to land at the section head, so the fragment-supplies-it route the unbound check accepts was assigned too late to be read at run time — accepted at dry-run, wrong at execution. -
A
{{x}}inside abind:value is validated at--dry-run, not at run time. hurl templates the injected[Options] variable:line when the entry runs, so a name nothing supplies used to pass dry-run and die mid-run;proef::lower::unbound_placeholdernow names the placeholder and the bind key at lower time, where the capture set is known. What legitimately supplies it: an earlier step’s capture, the fragment’s own[Options] variable:, a secret in scope, or a sibling literal bind whose name sorts earlier (injected lines are written and evaluated in name order). -
A literal
bind:that shadows an earlier capture is named, not silent. hurl’svariable:assigns into one shared set, so the bound value replaces the captured one from that entry on — sometimes intended, so it is a warning:proef::lower::bind_shadows_capture. A secret bind cannot shadow (it skips the[Options]path) and draws no warning. -
A failed
[run] setupreaches JUnit, the GitHub summary, and PR annotations. The abort used to return before the CI-report block, so a job gating on--junitsaw no file at all — indistinguishable from proef never running. The reports now carry the setup scenario itself (suite named by the setup feature file); exit codes are untouched (ADR-0014), and nothing is fabricated for the pool that never ran.
Internal
- The machine body has one exit.
execute’s six terminating paths each pasted the same empty-body emission; they now return through a single funnel with the oneemit_machine_bodycall after it, so a new path cannot forget the contract — and the empty-selection body takes its exit code from the refusal itself instead of restating it. Post-merge cleanup pass over the deep-audit cycle; behavior pinned by the existing path tests. - The probe and bake share the whole
variable:line.template_readsre-spelled the injected[Options] variable:line by hand around the shared escaper;proef_core::lower::variable_option_linenow builds it for both, andquote_optionreturns to being private (library-surface swap; unreleased either way). The section-end flush inbake_entry_optionsalso drops its fence-branch duplicate — one check covers all shapes — and the console collapse builds its annotated message without cloning the diagnostic on the common single-site path. - The canary stopped trusting the index’s tail twice over: it skips
prerelease versions (the sparse index is publish-ordered, so a
9.0.0-betawould have become “latest”), and refuses a backport older than the pin by semver ordering (an8.0.2published after9.0.0would have produced a green about a downgrade). deny.toml’s advisory ignores were dead and are gone. The quick-xml pair was ignored under “the patched release is unreachable” — the quick-junit 0.7 bump made it reachable and the workspace has been on the patched line; the stale ignores would also have silenced any new advisory against it.quick-xmlitself re-pinned to quick-junit 0.7’s in-tree copy (=0.41.0, one lock generation);lsp-serverrides to 0.10,tomlto 1.x.
Documentation
- The corpus tells the truth again, audited claim-by-claim: CONFIG
documents the
[env.<name>.run]jobs-only rule a reader used to discover as a parse error; GETTING-STARTED can produce its own output (it now states the fixture token its §5 requires, and its reproduce command names the--secretthe replay needs); EVENTS carries the provenance, totals, and field facts consumers implement against; TESTING-STRATEGY describes the CI that exists; RELEASING’s gate list predicts CI; TROUBLESHOOTING’s exit-1 row covers the--checkfamily. The#Nin an outline instance is documented as positional, with the column-placeholder naming that keeps identity stable across--shardand JUnit history.
[0.14.0] - 2026-08-18 (proef at CI scale)
Fixed
--rerunafter a cancelled run continues it, instead of a false green.--max-fail(and Ctrl-C) stop a run early with the never-reached scenarios honestly recorded as skipped — but--rerunfiltered to failures alone, so stop → fix → rerun ran only the old failures and reportedexit 0with most of the suite never executed in either run. Reproduced live before fixing (found by round-15 external review): stop at 2 of 6, fix, rerun →2 passed · 0 failed, green, four scenarios untested. On a cancelled base record--rerunnow runs failures plus the cancellation-skipped tail, and says so (note: the last run was cancelled before N scenario(s) ran…); scenario-level skips only exist under cancellation, so a completed base keeps the old semantics exactly. This also changes--rerunafter Ctrl-C — continuing the unfinished work is what stop → fix → continue always meant. Mutation-tested: reverting the union fails the continuation test.
Added
-
proef test --shard I/N— stable hash-mode sharding (R3-3). A CI matrix runs--shard 1/N…N/Non separate machines; scenarios are assigned by a frozen FNV-1a hash of the run-wide(file, scenario)identity, so adding a scenario never re-buckets the others — the measured stability argument that rejected index-slicing at triage (inserting one scenario re-bucketed the whole shifted tail under slicing, nothing under hashing; the shard tests pin both directions, and the assignment itself is frozen by literals — the hash is a published contract, and changing it would be breaking). Sharding applies after every other selector (the pinned filter→shard order), so each matrix job partitions one agreed-on set. An empty shard of a non-empty selection is a note and exit 0 — a small suite over a big matrix is a fact, not a mistake — while an empty selection keeps the loud typo’d-filter refusal, sharded or not. -
proef flaky— flakiness verdicts over the retained run history (R3-2). The 2026 discipline is detect → quarantine → resolve, and proef already owned the middle step:@quarantineruns a scenario without gating the exit code. This is the missing detect, a fold over the recordsruns-diralready retains — the window is[run] keep-runs, and no new state is written. Three signals from fields the record already carries (ADR-0008): flapping (verdict changed between consecutive observed runs more than once — transition-counting, not fail-rate, which is what separates flaky from broken: a scenario failing every run is consistently broken, a different problem), passes only on retry (green, but some step needed more than one attempt — the latent flake pass/fail-history tools structurally miss; the record keeps per-step attempts), and always failing. A cancellation-skipped row is not evidence and does not count toward a scenario’s history; phases are excluded (ADR-0014).--output jsonemits one object per scenario with the counts behind each verdict. Fewer than two runs is refused (exit 2), the same answerdiffgives. -
proef test --max-fail Nstops the run after N suite-scenario failures (1= fail fast) — the convention Playwright (--max-failures), pytest (--maxfail) and cargo-nextest (--max-fail) share, with the shared honest semantics: in-flight scenarios finish, the never-run rest record as skipped (not absent, never passed), and teardown still runs on its own token. The stop rides the graceful-cancel path Ctrl-C already exercises, so the record is a complete cancelled run — whichdiff --fail-on-regressionalready refuses to certify, exactly right for a deliberately-partial one.[run] setup/teardownfailures never count toward the threshold (a broken fixture is not a failing test, ADR-0014).
Documentation
- The R3 enhancement registry is triaged (OPEN-FINDINGS):
--max-failbuilt; a flakiness verdict over the run history and hash-mode sharding validated as build-next (the 2026 flaky pipeline is detect → quarantine → resolve, and the@quarantinetag already owns the middle step); CTRF, pack doc and the pre-M6 seam refactors deferred with named triggers; OTel and Cucumber-Messages exporters declined under ADR-0008’s one-record rule; items defined only in the absent v1 research document held for a spec.
[0.13.0] - 2026-08-17 (a record that travels, and a secret that stays one)
Added
-
proef difftakes a path. Each side is now a run id, a record directory, or an events.jsonlfile under any name — the stream is the record (ADR-0008), so all three must mean the same thing. The file form is the CI baseline flow an adopting suite asked for: download the base branch’sevents.jsonlartifact andproef diff baseline.jsonl <new> --fail-on-regressiongates the PR, with no shared record store. Previously every argument was joined ontoruns-dir, so a path produced.proef-runs/<your path>/events.jsonl: No such file— the argument mangled into the complaint. A path that does not exist now names itself; a--baselineflag was considered and declined as a second spelling of the same positional. -
[run] keep-runsbounds how many past run recordsruns-dirretains. The policy already existed as a hard-coded 200; it just could not be expressed, so a suite re-run on every save accumulated records for a day with nothing signalling a ceiling.0keeps none but the run in flight. Rotation still only ever deletes directories named by a generated run id —runs-dirmay be.— so a--run-id <name>record sits outside the budget and is never rotated, now stated in CONFIG.md rather than left to be discovered. Filed as R12-2.
Fixed
-
A run record no longer names the machine that produced it.
[run] suiteresolves against the config directory (0.12.0), so a path-lessproef testhanded the front end an absolute path — and every emitter printed it: the.hurl# source:header,.map.json’sfeature.file, everystep_finishedevent, the console, and pack diagnostics. Two checkouts of one suite stopped producing equal artifacts, which is the property ADR-0010 exists to guarantee; an adopting suite hit it as/Users/…in 133 artifact lines and 64% of its event stream by bytes.The resolution rule was right and stands. What was missing is its naming dual: resolve against the project, then name against the project again.
front::SourceNamingis now the one boundary that answers “how is this path spelled”, for features, packs and fragments alike — replacing the fragment corpus’s separate cwd-relative strip, which was a second anchor for the same question. The four ways to name one suite — derived from[run] suite, typed, typed absolutely, or reached from a subdirectory — now emit one artifact, byte for byte.A path that arrives relative is recorded exactly as it arrived; a suite or corpus genuinely outside the project keeps its absolute name, there being no project-relative spelling of it. Filed as R12-1, and it closes R9-6, which had described the same defect as safe from the project root — it no longer was.
Breaking, by the rule in
docs/RELEASING.md: it changes emitted artifact bytes, which is inherently breaking and takes a MINOR bump. Migration: nothing to do for a suite invoked with a typed relative path — those bytes are unchanged. A tool readingstep.fileorfeature.fileout of a record produced by a path-less run now sees a project-relative path where it saw an absolute one; join it onto the directory holdingproef.toml. Records written by earlier versions are not rewritten.
Security
-
An encoded reflection of a secret is redacted (S1). Redaction was exact-match on the raw secret bytes, and a server that reflects a bearer token encoded — an OAuth introspection endpoint, a debug echo, a JWT claim — defeated it: a failing assert quoted the base64 form in its detail, and a string trivially
base64 -d-able back to the live credential reached the console andevents.jsonl, the retained record CI uploads. Demonstrated live against 0.12.0 by an external research pass and reproduced here before fixing.Redactions::newnow derives each secret’s common encoded forms as additional needles — base64 (standard and URL-safe alphabets, with and without padding), hex (both cases), RFC 3986 percent-encoding, and the JSON-string escape — so every construction site (the CLI sink, the engine’s internal renderer, TAP) is covered by construction. This is the remedy GitHub’s own log-masking documents for the same limitation: register each transformed value too. The needle set covers the reversible transforms that occur at HTTP boundaries and does not claim completeness — a secret reflected hashed or re-encrypted matches no needle list. Over-redaction is the accepted failure direction. Property-tested over every derived form, pinned end-to-end by a fixture route that echoes the bearer base64-encoded, and recorded as an ADR-0005 amendment. -
The fragment corpus read is bounded.
[run] fragmentsnames a directory proef did not write and does not control, and it was read with no per-file or total cap: a 279 MB file cost 601 MB of resident memory onproef flows— a command that never looks at a fragment — because the text is read whole and then copied into anArc<str>. A file over 8 MiB is now skipped (proef::pack::oversized_fragment_file) and the reader stops past 64 MiB total (proef::pack::fragment_corpus_too_large). The size comes from the directory entry, so an oversized file is never allocated at all; the same bound applies inproef lsp, where the corpus is held between requests rather than for one command. Skipped, never fatal — a corpus is foreign by design, so one bad file must not sink the ones beside it. An unreferenced corpus still costs nothing: the scan stays lazy, so nothing is reported unless a pack actually names a fragment. Filed as R9-3.
Internal
-
A hung test is now a five-minute failure, not a five-day zombie. The nextest config had
slow-timeoutwith noterminate-after, which only labels a test SLOW and never kills it — anlsp_stdiotest wedged on an unboundedchild.wait()ran for five days with itsproef lspchild alive. Both layers fixed: the two barechild.wait()sites got the file’s own bounded-watchdog pattern (a server that fails to exit now fails the test in 10s, naming what did not exit), and the runner gainedterminate-after = 2(120s), sized from a cold-cache census of the whole suite (slowest ordinary test: 5.1s). Theharness_trio — which shellscargo testinside the test and measured 216s on a fully cold cache — gets a per-test override to 600s, the nextest docs’ own tight-global-plus-overrides pattern. The process-group kill (a spawned server dies with its test) was verified empirically with a deliberately hung test holding a live child. -
Cleanup pass over this cycle’s four PRs (reuse/simplification/efficiency/ altitude review). The corpus-bound decision moved into core as
pack::CorpusBudget— it was abstracted in the CLI and hand-copied in the LSP, agreeing by copy rather than by construction; both readers now share it and only measurement stays reader-local.Redactionsstopped allocating on the miss path (nearly every call: per string field per event under the reporter-stack mutex, with the needle list ~9× larger since the encoded forms) — clean fields now hand back their originalArc. A relative source path is left exactly as it arrived, per its documented contract — it had been falling through to a per-file canonicalize that could rewrite a../-typed spelling. The LSP’s percent-encoder folded onto core’s (byte-identical copies, one character set to drift). The fixture’s hand-rolled base64 became the crate call — its dependency-surface rationale died when this same cycle madebase64a workspace-wide compile.diff’s path-or-id resolution moved beside its sibling inrecord. Adeny.tomlhome for the curl floor was tried and reverted by mutation test: cargo-deny 0.19.8 mismatches build-metadata versions (curl-sys@<0.4.90banned the good0.4.90+curl-8.21.0); the floor stays a unit test, now scanning every lockfile entry rather than the first. -
The bundled libcurl cannot silently regress under the June-2026 CVE batch.
curl-sys 0.4.90+curl-8.21.0in the lockfile is past the batch — but only as a transitive accident of resolution, and the usual gates are structurally blind here: RUSTSEC carries no advisories for CVEs in a*-sys-bundled C library, socargo audit/denystay green however stale the bundled curl is. A test now asserts the lockfile floor, and each release build prints the libcurl actually linked into that artifact (proef doctoralready reported it; the release log now carries it per target). The hurl-8.1 watch items —variables-file:’s missing sandbox first among them — are recorded as a pin-bump checklist in the thin-fork runbook. -
Fuzzing reaches the fragment rules.
fuzz_pack_loadran against an empty corpus, soref:resolution,bind:keys nothing reads, abind:colliding with a variable the fragment supplies itself, and unbound placeholders were covered on paper and unreachable in fact. The newfuzz_fragment_bindingtarget is structure-aware: it builds a well-formed pack and corpus and spends its budget on the name space where those rules live. That shape was chosen from measurement, not taste — a byte-oriented version never once resolved aref:in 1.45 million runs, because reaching the rules meant discovering valid YAML and a matching corpus at the same time. The corpus is read by a synthetic scanner rather than hurl’s, which is what keeps the fuzz workspace free of native libraries: cargo dependencies are package-level, so one engine-dependent target would compile hurl for all of them. -
Hurl’s own annotation scanner is property-tested, in
proef-engine-hurlwhere the native libraries already are. The properties pin what the entry-boundary arithmetic is for: every reported line lies inside the file, every entry is accounted for exactly once, the starts are ordered and distinct, and — the one that matters — no fragment’s text runs into the entry after it. That last assertion exists because a first draft without it passed while the boundary was deliberately broken. -
The fuzz target list comes from
cargo fuzz list. It had been spelled out inci.ymlandnightly.yml, so a new target ran nowhere until both were edited, and nothing failed to say so.
[0.12.0] - 2026-08-14 (one path rule, and a watcher that stops lying)
Fixed
-
A
runs-diredited mid---watchno longer feeds the loop its own output. Reruns re-read the config (the fix below), so records went to the new directory while the watcher’s exclusion still named the one it had frozen at startup — and every rerun’sartifacts/*.hurl, now under an unexcluded directory, requeued the next run. One edit produced 39 runs in 12 seconds, firing real traffic. This was the third outing for the watch-feedback class, so the fix removes the second answer rather than resynchronising it: each rerun registers where it is about to write, and the exclusion is derived from the same config the run is. A directory a previous run wrote stays excluded too, since its events can still be in flight. Filed as R11-8. -
--watch --config <relative path>retriggers on config edits. The watcher compared the config by exact path whilenotifyreports events under the spelling the OS resolved them to, so--config proef.tomlnever matched and config edits produced nothing — silently, because feature edits kept working and the loop looked alive. Symlinked and/tmp-style aliased paths failed the same way and are also fixed: the flag is made absolute when it is stored, and identity is settled by comparing canonical paths, which is a stricter question than being absolute. The same relative-path flaw silently costproef lsp --config <relative>go-to-definition across the whole fragment corpus, sincedocuments::name_to_urlrefuses a relative name. Filed as R11-9. -
doctorreports aproef.tomlthat will not parse. The discovery arm had become a silentunwrap_or_default, so a malformed config leftdoctorreporting on invented defaults and printing “all checks passed”, exit 0 — with the parse error, which the previous code printed, discarded. It is aproject:row now, so it reachesworstand the exit code a CI script actually reads. Being absent is still not a finding:doctormust run outside a project. Filed as R11-10. -
proef fragmentsexits non-zero when a[run] setup/teardownphase fails to load. It printederror: setup feature failed to validate:and exited 0, because the phase half flattened its failure to “not measured” while the suite half kept its code. Withholding the counts was right; reporting success while printing errors was not. -
proef.tomlhas one path rule. A path written in the config now resolves against the directory holding the config; a path typed on the command line still resolves against the working directory.[run] fragmentsalready worked this way and everything else did not, so two keys in one table meant two different roots: from a subdirectoryfragments = "hurl"resolved whilesuite = "features"reported “neither a feature file nor a directory”. With--configthe split was worse than inconsistent — pointing at a config in another tree randry-run OKover whatever suite happened to sit beside the shell, and never looked at the configured one.The rule now covers
suite,setup,teardown,runs-dirand thetests/convention probe, plus two files nothing had inventoried:.proef-state.json(the persistent World) and.proef-secrets.json(the secret store), which were anchored on the working directory — so two shells in one project were two Worlds and two secret stores. It is the convention Cargo,tsconfig.jsonand pytest’s rootdir all follow. Absolute values are taken as written, and with noproef.tomlin scope written paths stay relative to the working directory, so the config-independent reference corpus is unaffected. Filed as R11-1. -
--watchrereads the config it retriggers on. Editingproef.tomlretriggered a run that still used the snapshot loaded at startup: changing[url] baseproduced a rerun that dutifully called the old host, and the same went stale forjobs,[env.*]andexclusive-tags. Watching a file whose contents you then ignore is worse than not watching it, because the rerun reports that the edit was taken. Each rerun now re-reads the file and re-resolves the suite from it; a config that no longer parses fails that rerun and leaves the loop watching, since half-typed TOML is the normal state of a file being edited. Which directories the loop watches is still fixed at startup, so changing[run] fragmentsor[run] suiteneeds a restart to be watched. Filed as R11-2. -
--configis honoured or refused by every subcommand.doctorprinted the error for a missing named file and then reported on defaults, exit 0 — the “fall back to defaults”CONFIG.mdforbids — whilefmt,init,schemaandsecretaccepted a nonexistent path silently, against the “global to every subcommand” claim inCONFIG.md,README.mdand this file. A named file that is not there is now exit 2 everywhere, including where nothing reads it;doctorstays lenient about discovery, which is a different claim.secretadditionally uses the flag, since the store is the project’s. Filed as R11-3.
Breaking: the secret store, the persistent World and the run records move with the config rather than with the shell. What decides whether this reaches you is where you invoked proef, not where
proef.tomlsits: runs started from the project root are unchanged, but a run started from a subdirectory used to write.proef-state.json,.proef-secrets.jsonand.proef-runs/beside the shell, and now writes all three beside the config.Nothing is migrated, and none of it announces itself. A World written from a subdirectory reads as empty, so
saveAs: globalvalues start over on the first run after upgrading; stored secrets read as absent; and the old run records are simply invisible toexplain,reportanddiff, which say “no run records” rather than erroring. To carry them over, move.proef-state.json,.proef-secrets.jsonand.proef-runs/from the directory you used to run from into the one holdingproef.toml. Otherwise re-runproef secret setand take a fresh baseline.Breaking (library):
proef_cliis not a published library surface, but for the recordfront::runtakes the state-file path,ProjectConfig::runs_dirreturns aPathBuf,setup/teardownreturnOption<PathBuf>,suiteis gone (fold intodefault_suite_path), and thesecretstoreentry points take the store path.proef_coregains one item:pack::FragmentCorpus::unreadable_file.
[run] exclusive-tagsvalidates itself.--dry-rundid not parse the expression at all, so a malformed one exited 2 fromproef testand passeddry-run OK … 0 warning(s)from the gate CI runs. And a well-formed expression matching no scenario was silent:@solozagainst a@solosuite put every scenario back in the shared pool, exit 0, nothing said — the exact silent degradation the key was designed as a config expression to prevent, and one that reads as flakiness rather than as a typo. Both paths now parse it, and a zero-match expression warns, naming it and pointing atproef flows. Judged over every scenario the suite loaded rather than the ones selected, so a--tagsfilter that removes the matches from one run is not reported as a broken setting. Filed as R11-4 and R11-5.
Changed
-
proef fragmentssays which half it could not measure.--checkreported “needs a suite that binds” when the suite had bound perfectly well and a[run] setup/teardownfeature was the thing that failed to load, sending the reader to inspect the half that was fine. The degraded listing also now carries the notemacrosprints, so withheld counts read as “not measured” rather than as a corpus nothing uses. -
proef fragments --checkrefuses to pass with no corpus configured. With[run] fragmentsunset it printed0 entriesand exited 0, indistinguishable from a fully-used corpus — so a CI gate disarmed silently the day the key left the config. The listing still works; only the gate is now a user error. -
proef fragments --output jsoncarriesannotatedon both row shapes. The annotated and unannotated rows differ in eight fields, and consumers had to probe for the absence of one to tell them apart.
Documentation
CONFIG.md’s “everything else keeps running atjobswidth” was false: queueing is strict FIFO, so nothing new starts while an exclusive scenario waits at the head. The cost is bounded, not absent, and is now described.- The one caveat
[run] exclusive-tagscarries is written down inCONFIG.mdand ADR-0007: exclusivity is enforced against the dispatcher’s active set, which a watchdog-abandoned scenario leaves while its detached thread is still issuing requests (hurl cannot be cancelled mid-entry). TECH-SPEC§10 gainedproef fragmentsand the global--config; §11’s[run]inventory listed three of seven keys.DIAGNOSTICS.mdcarried apack::loadrow nothing emits — a reader who grepped it found a plausible cause that could never be one — and filedlower::multiline_bindunderproef::pack::*. Both fixed, and the two-way agreement between the file and the emitted codes is now a test, since this drifted twice.OPEN-FINDINGSR9-2 still saidfuzz_tag_expr“sits in neither fuzz loop” three sections after recording that it is in both.
[0.11.1] - 2026-08-12 (the gaps 0.11.0 shipped with)
Fixed
-
An output path creates the directories it names.
--junit,--sarifandreport -ofailed when the parent directory did not exist, whileartifacts -oand the run directory created theirs — no rule, four sites deciding separately, with the two used most in CI on the failing side. Every adopter paid the samemkdir -p.pytest --junitxml,jest-junit,cargo-nextest’s JUnit store and thehurlproef embeds all create them. This does not weaken the “side effects should be explicit” principle: that is about writing files the user did not name, and here they named exactly this path. -
proef fragmentscounts[run] setup/teardownusage. A fragment only a phase feature reached was reportedUNREACHABLE — no macro refs it, which was false, and failed--check— a false CI failure in the workflow--checkexists for. The verdict also depended on where the phase file sat: inside the suite directory it was discovered as an ordinary feature and counted. The listing’s universe now matches the runner’s, and a phase that fails to load withholds every count rather than guessing. Filed as R10-2. -
One predicate answers “is this a fragment file?” (
FragmentSupport::claims). Three answered it before — CLI discovery viaPath::extension, the core scan viarsplit('.'), and the LSP’s corpus invalidation case-insensitively — so they disagreed aboutapi.HURL(the editor rebuilt its corpus for a file nothing would scan) and about a dotfile named.hurl. Filed as R10-3. -
--configreachesproef lspand--watch. The flag bypasses the upward search so aproef.tomlbeside the suite becomes usable — but the editor re-discovered its own config and--watchwatched whatever a fresh search found. So in exactly the layout the flag exists for,proef test --config …ran green while the editor reported everyref:as unknown, and editing the config driving the run never retriggered it.ProjectConfignow keeps the file it was read from (withrootderived from it rather than stored beside it), and both consumers use the config actually in force. Forproef lspthe flag also outranks the client-announced workspace root. Filed as R10-1.
[0.11.0] - 2026-08-12 (the adoption response)
Added
-
[run] exclusive-tags— a tag expression selecting scenarios that run with the pool to themselves. Real suites contain scenarios that cannot run beside anything: one asserting absolute positions (items[0]) needs a store no concurrent scenario writes to, and the only workaround was several CLI invocations driven by tag discipline in a Makefile, each producing its own run record, JUnit file and exit code to aggregate in shell.A matching scenario waits for the pool to drain, runs alone, and the pool refills after it, with discovery order unchanged so an exclusive scenario never loses its place. Queueing is strict FIFO, so nothing new starts while one waits at the head — the throughput dip around each exclusive scenario is the price, and it is bounded. A config expression rather than a reserved tag name, because with a bare convention a scenario added months later lands untagged in the parallel pool and breaks isolation intermittently — which reads as flakiness rather than as a missing declaration. A malformed expression is a user error, never a silently-ignored key.
This is exclusion, not ordering: a scenario that must run before the rest belongs in
[run] setup, which already runs once before the pool exists. Deliberately one axis of the twocargo-nextestsettled on — per-group concurrency limits (rate-limiting a shared dependency) are a real future need that nobody has asked for, and a group table can be added later without breaking this key. -
proef fragments— the corpus listing, symmetric withmacros. Until now no proef output stated how many fragments there were, so neither way a fragment can die had a denominator to be noticed against: one no macro references was unobservable, and one reached only through a macro no scenario binds looked covered because the macro was flagged. Both are now named apart, unannotated entries are listed by line (they have no name to list by), and--checkexits 1 when something never runs.--require-annotatedextends that to unannotated entries and is deliberately opt-in: an unannotated entry is inert by design (ADR-0018), so “not done yet” is a porting team’s meaning, not every adopter’s. Reachability is read off the lowered scenarios, so a fragment reached through a chain ofuse:counts as reached. -
--config <path>, global to every subcommand, naming theproef.tomlto read instead of searching up from the working directory. Discovery only goes up, so a config beside the suite is unreachable from the repository root — a layout an adopting team planned and abandoned after it failed. A named file that does not exist is a user error rather than a fall back to defaults: discovery finding nothing means “no project here”, but a named path that is not there is a typo, and a silently unconfigured run is what that used to buy. -
proef doctorsees the fragment corpus — a row reporting how many fragments loaded from[run] fragments, warning when the configured root is not a directory. A misconfigured path used to surface much later aspack::unknown_ref: an error about a name when the cause is a path. -
proef initscaffolds both body forms — a one-entry.hurlfile with a# @proefannotation,[run] fragments, and a pack macro of each kind. The newcomer with most to gain fromref:is the one who already owns a hurl corpus, and a scaffold teaching onlyhurl: |reads as “proef wants your files transcribed into YAML”.
Fixed
-
A
bind:key nothing reads is refused (proef::pack::unread_bind_key), with did-you-mean over the names actually in scope.bind_without_refonly caught a table with noref:at all, sobind: { token: …, toekn: … }validated clean — the one authoring mistake in the fragment path that produced no signal whatsoever. Checked as a union over the scope, never against one fragment: a pack-scope table is the plumbing every macro in the file needs, so a key serving one macro and not its siblings stays correct. -
duplicate_fragmentno longer says “in bothxandx” for two entries in one file, and stops offeringfile.hurl#nameas the remedy there — that qualifies by file and cannot separate two entries inside one. Annotating a corpus adds many names to few files, which makes same-file the likely collision. -
unbound_placeholdernames all three supply routes. The omitted one was the fragment’s own[Options] variable:— the route that makes a corpus file runnable standalone, which is the property ADR-0018 exists to preserve. -
A fragment’s
[Options]escaped the ADR-0007 value caps.retry: -1,repeat: -1and an unboundeddelay:were rejected in an inlinehurl:block and accepted in aref:fragment — byte-identical text, exit 2 one way and “dry-run OK, 0 warning(s)” the other, then written verbatim into the executed input. The scan lived inside the inline-only linter; only the twinned-option half of pass 6 had crossed to fragments. It reads the text alone, so it now runs against a fragment’s too, anchored on theref:line and naming the fragment file and line. This is the case the caps exist for: hurl has no cancellation, so an infinite retry makes the batch budget unestimatable and leaves the watchdog abandoning a thread it cannot stop. -
A step declaring both
ref:and a payload was told, falsely, that its pack had noref:at all. The conflicted step is reported and dropped, so the loaded bodies stop showing everyref:the author wrote — and the pack-scopebind_without_refcheck then drew a conclusion from the gap. It now infers nothing from a pack whose steps did not all normalize. -
A pack-scope
bind:with noref:anywhere was silently dropped.AUTHORING.mdsaidbind_without_refapplies “at every scope” while only the macro and step scopes were checked — and a setting ignored in silence is the bug those two exist to refuse. The check was the better half of the disagreement, so the pack scope now has it too. -
A multi-line
bind:value blamed the artifact. A hurl[Options] variable:value is a single-line scalar, so a newline could never reach the entry — but it surfaced one stage later asemit::invalid_artifact, pointing at generated text the author never wrote. Refused by name at lower time aslower::multiline_bind, naming the inlinehurl: |form that is what splices a multi-line body (ADR-0018’s splicing-versus-binding boundary, enforced where it can be explained).
Changed
-
Breaking (library):
AnalyzeCtxtakes the fragment corpus instead of building one. Building it internally meant a fresh scan memo per call, so the LSP re-read and re-hurl-parsed the whole corpus on every request — each completion popup, each go-to-definition, each debounce tick. The server now holds one and rebuilds it only when a fragment file changes; editing a pack or a feature, which is nearly every keystroke, leaves it alone. It is also what core purity already required: the caller does the IO. -
Breaking (library):
StepKindSpecgainedoptions, an engine-contributed recogniser mapping a raw option key to what ADR-0007’s budget rules should make of it. The fragment half of that rule already crossed the seam while the inline half matched"retry-interval:"as a literal insideproef-core— one rule at two altitudes, and a second engine would have had its fragments linted and its inline blocks not. A kind contributing no recogniser is not linted, since the core has no way to know what its option keys mean. -
Breaking (library):
proef_core::engine::FragmentScannerreturnsScannedFile { fragments, unannotated }rather thanVec<ScannedFragment>. An engine’s scanner now also reports the 1-based lines of entries carrying no annotation — lines only, never built-then-discarded fragments, so a foreign corpus still costs a push per unannotated entry. Without it “which entries did I forget to annotate?” is unanswerable: a missing annotation produces a green run and a silently absent test, and the entry that would prove it was never built.FragmentCorpusgainsfragments(),unannotated()anddiagnostics(), because the scan is gated on some pack naming a fragment — soPackSet::fragmentsis empty for exactly the suite a listing has most to say about.
Documentation
-
Config discovery is a requirement, not a convention.
proef.tomlis found by searching up from the working directory, so a config beside the suite (tests/proef/proef.toml) is never found from the repository root — an adopting team planned that layout and discovered it by failure. CONFIG.md now says so, and notes that keeping the file at the root collapses the one place[run] fragments(config-relative) andsuite/setup/teardown/runs-dir(cwd-relative) differ. -
The release runbook could not work as written.
mainis a protected branch, and step 4’sgit push origin main --follow-tagsfails in the dangerous direction:--follow-tagsis not atomic, so the branch is rejected while the tag still lands — and the tag is whatrelease.ymltriggers on, starting a release build from a commit that is not onmain. It happened cutting 0.10.0. The runbook now routes the release commit through a PR and tags the merged commit, and thecargo publishsection carries the dry-run, tag-check and--lockedsequence plus why only four crates go ([workspace.package] publish = falseis the default). Also drops step 1’s reference to changelog “bottom links”, which do not exist.
[0.10.0] - 2026-08-12 (named hurl fragments)
Breaking (library):
proef_core::pack::loadtakes a&proef_core::pack::FragmentCorpusbetween the packs and the step kinds (&FragmentCorpus::empty()for the previous behaviour, orFragmentCorpus::new(sources, kinds)to supply fragment files), andPackSet::fragmentsis anArc<BTreeMap<…>>so one scan can be shared by every load;LoweredScenario::secretsis aBTreeMap<String, String>of engine-variable → secret name rather than aBTreeSet<String>;PreparedandScenarioCtxeach gain asecret_bindingsfield carrying that map to the engine; andSourceProvider::discover_fragmentsis a required method (returnOk(Vec::new())to serve none) — it was briefly defaulted, and the default silently disabled fragments for a provider that forwarded the other two; andScannedFragment::nameis aStringrather thanOption<String>, because a scanner now reports only the entries it found an annotation on; andScannedFragmentandpack::Fragmenteach gain asupplied_variables: Vec<String>(Vec::new()for none), which an engine’s scanner must fill from the entry’s[Options] variable:lines — leaving it empty reinstates the silent last-wins it exists to refuse; and bothLoweredStep,StepOutcomeandEvent::StepFinishedgain afragment: Option<String>field andanalyze::FragmentDefgainsplaceholders: Vec<String>, so a literal construction of any of them needs one more line (None/Vec::new()reproduces the previous behaviour). The wire schema is unaffected — the event field is skipped when absent, which is what keeps existing records byte-equal.
Added
-
The docs are checked mechanically, not only read.
xtask docs-checkgained two passes — every relative link resolves, and every fencedtoml/yamlexample parses with the product’s own parsers, so the check means “proef would accept this example” rather than “some parser would”. A third pass, whether a documented command or long flag actually exists, needs a built binary and so lives incrates/proef-cli/tests/docs.rs.All three were written against defects already in the tree: ADR-0018’s first example could not load (an unquoted
${…}inside a YAML flow mapping, where{opens a nested mapping), and a row marked shipped documentedproef report --html, a flag that never existed. Both had correct prose around wrong code — the failure mode review does not catch. -
Packs can name fragments:
ref:andbind:(ADR-0018). A macro step’s body may beref: <fragment>instead of an inlinehurl:block, andbind:supplies the fragment’s{{…}}variables at pack, macro and step scope, most specific winning. Fragment names are global, andfile.hurl#namequalifies one — the same two spellings, resolved the same way, thatuse:already accepts.Refused at load, each with its own code: a
ref:naming no loaded fragment (unknown_ref, suggesting the closest, and saying so plainly when no fragment file was loaded rather than implying a typo); two files declaring one name (duplicate_fragment); a file the engine cannot read (bad_annotation— its siblings still load); a step that is bothref:and a payload (body_form_conflict); andbind:on a step with noref:(bind_without_ref— an inline block takes${…}, so that binding would feed nothing, and a setting silently ignored is the bug this refuses to ship).A fragment declaring its own retry alongside a step’s
retry:is the sameoption_declared_twicean inline block gets, so the two body forms behave identically rather than differing by where the hurl text happens to live.A fragment may also supply a variable to itself with an ordinary
[Options] variable:line — that is how a corpus file stays runnable on its own, so it counts as an answer to that fragment’s own{{…}}and needs nobind:. Supplying and binding the same name is refused (option_declared_twice): both reach the entry asvariable: k=, hurl takes the last, and the fragment’s own line is last — so the bound value would silently never be sent, and would stay unsent for every later entry, since hurl’svariable:assigns into the run-level set rather than scoping.Discovery arrives below, so a
ref:resolves end to end. -
[run] fragments— the hurl files a pack mayref:. Names one root, scanned recursively for the extensions the registered engines claim, so discovery never learns a file type of its own. Unset means no fragments: there is no convention fallback, becauseunknown_refsaying “no fragment files were loaded” beats guessing at a directory.Relative paths resolve against
proef.toml’s own directory, not the working directory. The config is found by walking up from the cwd, so a path in a config three levels above must mean “relative to the project” — otherwiseproef flowsfrom a subdirectory reads the right config and then cannot find anything it names.[run] suitepredates this and stays cwd-relative; it is only consulted when no path was given, so the difference is not observable there.The LSP resolves fragments through the same root, so
ref:does not read as unknown in an editor while the suite runs green.--watchretriggers on.hurledits and watches the fragment root separately, since a corpus may live outside the suite.proef fmtstill refuses.hurlin both discovery branches — it locates hurl blocks inside YAML, and a corpus proef did not write is not proef’s to rewrite — now pinned by a test. -
Fragments lower, bind, and execute. A
ref:step emits the fragment’s own text with its non-secret bindings baked in as per-entry[Options] variable:lines, so the artifact stays the executed input and replays identically under the stock CLI (ADR-0010). Values are always quoted:variable_valuetries null/bool/number before string, so an unquotedrecords,2andtruewould become three different types by accident.Two refusals guard the parts that could otherwise pass silently:
lower::unbound_placeholder— a fragment reading a{{variable}}that nobind:in scope supplies and no earlier step captures, anchored on the.hurlline the variable is on rather than on the pack. hurl’s[Options] variable:assigns into one shared set rather than scoping, so an unbound name would inherit whatever a previous entry happened to leave and run green against the wrong value.lower::secret_in_composite_bind— abind:value mixing${secret:…}into a larger string. To inject that, the composite would have to be materialized into the artifact, which ADR-0005 forbids; bind the secret alone and let the fragment spell the surrounding text.
Secrets keep their own path: recorded as engine-variable → secret name and injected via
insert_secretat run time, never as an[Options]line. That indirection is what letsbind: { auth_token: "${secret:apiToken}" }give a secret the variable name a corpus proef did not write already uses.Bindings resolve once per scope instantiation — pack scope once per scenario, macro scope once per invocation, step scope per step — so one binding is one value and two bindings are two. A macro with no
ref:step resolves nothing, so an unused table never advances the${fake:…}counter. -
The engine seam can describe fragment files (ADR-0018, groundwork).
StepKindSpecgainsfragments: Option<FragmentSupport>, andproef-coregainsScannedFragment/FragmentScanError/FragmentScanner. The hurl engine implements the scanner over hurl’s own AST: the# @proef <name>annotation is read from the entry’sline_terminators, so the annotation↔entry binding is exactly as reliable as hurl’s parser and no text is scanned for structure. An entry’s required inputs and produced captures are read from the same AST, which is what will let an unbound placeholder be an error rather than a runtime surprise.Additive only — nothing was removed from
proef-core’s surface, and no hurl type appears anywhere in it. Discovery asks the registry for the extension instead of naming.hurlitself, so this stays ADR-0002’s “adding an engine leavesproef-corediff-empty” rather than an exception to it. Nothing observable ships yet: no pack can reference a fragment until the schema lands.StepKindSpec::fragmentsis oneOption<FragmentSupport>rather than a separate extension and scanner, so a kind that claims a format it cannot read is not expressible; a file no kind claims is skipped rather than handed to whichever engine happens to be registered first.ScannedFragment::declared_optionslists option families rather than flagging retry alone, so the core applies its double-declaration rule todelay:too — through the samebake_entry_optionspath, so leaving it out reproduced the very last-wins bug the rule exists to refuse.supplied_variablesis separate from it because the two clash on different keys: an option family family-to-family, a variable name-to-name.A note for whoever extends the scanner: hurl’s
Visitortreats templates as leaves, andvisit_template,visit_urlandvisit_filenameare three separate no-op defaults that do not forward to one another. Overriding onlyvisit_templatesilently under-reports an entry’s inputs — and a missing input reads as “needs no binding”. -
A run record says which fragment a step ran, and
explainprints it.step_finishedgains afragmentfield carryingfile.hurl#name(additive per ADR-0008: absent for an inlinehurl:block, so no pre-existing record changes a byte — the reference event-stream snapshot is unmoved), andproef explainrenders it under a failure asvia tests/hurl/admin.hurl#admin.search. A step that never ran reports it too: “not run” is exactly when someone is reconstructing what the suite was about to do.This closes a promise ADR-0018 made rather than adding a new one — three files per test was accepted on the condition that
explainand go-to-definition earn it back, and only go-to-definition had. The name is qualified at lowering rather than by the reader, because a record has to stand alone: by the time it is read, the pack that named the fragment may say something else.JUnit, the GitHub job summary and the::errorannotations name it too, as a trailing(via file.hurl#name)on the failure message, and the HTML report renders it under the reason. CI is where a reader is least able to go looking for themselves, so it is the last place provenance should drop out — and all three sinks share one helper rather than a format string each, because three copies is how one of them quietly stops agreeing with the run record. -
bind:completes against what the fragment actually reads. With the cursor in abind:table — flow or block style — the editor offers the{{variables}}of the fragments that packref:s, nearestref:ranked first, each labelled with the fragment that wants it. The names come off the engine’s own AST at scan time (analyze::FragmentDef::placeholders), so this is the file’s real interface rather than a second description that could disagree with it.Until now the only route to a foreign corpus’s variable names was to run the suite and read
proef::lower::unbound_placeholder— a lower-time error, so the names arrived only after a failure.bind:exists at three scopes and only the step one names a single fragment unambiguously, so the list is a union rather than a guess; the owning fragment rides in each item’s detail. -
The fragment corpus is scanned once per command, not once per pack load. A
proef testloads packs up to four times — the suite, then[run] setupand[run] teardown, each validated and then run — against different feature paths but always the same corpus, and each load re-read and re-parsed every.hurlfile. Measured on a 200-file / 15k-line corpus: 140 ms → 40 ms warm, with pack loading falling from ~28% of the run to a single pass. The win scales with the corpus, which is the direction adoption goes.The corpus is now read once per invocation (
front::fragment_corpus) into aFragmentCorpusthat scans itself lazily, at most once. Laziness is the part worth guarding:load_collectingstill scans only when some pack actually has aref:, which is what makes CONFIG.md’s “pointing at a corpus you did not write costs nothing” true. Hoisting the scan to the caller to share it would have bought the speed by breaking that promise, so the memo lives with the corpus instead — and a test proves the eager version fails, by pointing an unreferenced corpus at a file that cannot parse and asserting no diagnostic appears.Built per invocation rather than in a static:
--watchre-enters the same process after each edit, and a corpus outliving one run would serve pre-edit fragments to the next. -
Go-to-definition on a
ref:worked again, then briefly did not. Shortening the[run] fragmentsroot to a cwd-relative spelling — done so a run record would not carry an absolute, machine-specific path — also shortened the rootproef lsphands to its source provider. The LSP keys document identity on absolute names (name_to_urlyieldsNonefor anything relative), so everyref:go-to-definition returned null and.hurl-positioned diagnostics stopped publishing, while the suite still ran green. That is the capability restored two commits earlier.Resolution and spelling are now separate concerns:
ProjectConfig::fragments()returns a resolvable path, and the shortening happens at the naming boundary infront::fragment_sources, which only CLI runs pass through. Both properties hold at once — the editor resolves, the record stays portable.Covered by an end-to-end
proef lspstdio test with a realproef.toml, the seam the unit tests could not reach: they inject absolute names through a fake provider, so they never exercise config → provider → URI. The test canonicalizes its temp root deliberately — on macOS a tempdir is/var/…whose real path is/private/var/…, and without that the cwd comparison silently no-ops and the test passes vacuously. -
Every failure sink names the fragment, not just the CI ones.
via()moved fromci_reportstorender, and the console failure list and TAP diagnostic now carry it too. A helper scoped to one delivery channel was howproef testprinted no provenance on stderr whilereport.junit.xmlfrom that same run printed it — the drift the helper’s own comment says it exists to prevent.
Internal
-
The secret-name join has one home.
proef_core::engine::secret_variablespairs a scenario’ssecret_bindings(variable → secret name) with itssecrets(name → value) and is the only place that join is written. Doing it engine-side invited injecting under the secret name, which makes a renamed binding (ADR-0018) resolve to nothing — the request then leaves with an unresolved{{…}}and fails far from the cause. It yields borrows on purpose: an owned variable → value map would put a second copy of every secret value in memory per scenario, and ADR-0005 keeps values in one place. -
engine::OPTION_FAMILIESnames the vocabulary the double-declaration check compares against, andMacroStep::declared_optionsderives the other half of that comparison once for both body forms. The two sides were previously hardcoded lists that met by string equality with no test spanning the crates — a spelling only the engine knew would have matched nothing and quietly disabledoption_declared_twice, reinstating the hurl last-wins it exists to refuse. Aproef-engine-hurltest now asserts every family the real scanner emits is one the pack can declare;delaywas untested there entirely. -
Lowering’s two diagnostic sinks are one
Sinksvalue. They were adjacent parameters of the same type threaded through seven functions and a closure: transposing them at any of a dozen call sites compiled cleanly and routed every error intowarnings, so a scenario that should have failed lowered “successfully” and the run exited 0. No&mut Vec<Diag>parameter remains inlower.rs, which makes the mistake unspellable rather than merely unmade.
Documentation
-
AUTHORING says which body form to reach for, and why. A table contrasting splicing against binding — what each can substitute, whether it can be reused, whether stock
hurlcan run it, and when an unknown variable is caught — plus the rule that decides it: inline when you need to splice something hurl cannot template (${docstring}as a body has no binding equivalent),ref:when the request is shared, foreign, or must stand alone.CONFIG.mdgains[run] fragmentswith a worked three-file example. -
The hurl non-goal is about generation, not direction (PRD §3 amendment). It read “importing/round-tripping hand-written hurl files into Gherkin (artifacts flow outward only)” — a clause and a parenthetical saying two different things, the parenthetical forbidding hurl text from being an input at all. What the non-goal protects is that proef never authors a test for you, and that reasoning is untouched (ADR-0016 stays declined on it). It does not extend to hurl being an input source, which §1’s own framing — “there is no tool that joins the two” — describes as the product’s purpose. Recorded honestly: OPEN-FINDINGS M3 asked for this re-examination to arrive with a measured port cost, and it has not.
-
ADR-0018 — named hurl fragments. A macro step’s body may be
ref: <fragment>naming one entry in a real.hurlfile, annotated# @proef <name>, with proef values supplied by an explicitbind:map instead of${…}splicing. The file stays valid hurl, so the same file runs underproef testand under stockhurl. Inlinehurl: |is unchanged and stays: the two are splicing versus binding, with different capability envelopes, and the 844-line corpus port is recorded in the ADR as evidence the inline path is sufficient for real work. No behaviour ships with this entry — the ADR and the charter amendment land first, deliberately.
Fixed
-
--watchreran itself forever. ADR-0018 added the engines’ fragment extensions to the retrigger allowlist —.hurlamong them — while every run writes.proef-runs/<id>/artifacts/*.hurl. A watched tree containing its own runs dir fed itself: 49 runs in 15 seconds, firing real traffic in a tight loop and churning record rotation. The filter now excludes generated trees by directory name, reusing discovery’s ownskipped_dirso there is one rule with two consumers, and takes[run] runs-dirfor the case where it is not a dot-directory.OPEN-FINDINGSP5 had closed this “by inspection”, naming.hurlas a file that could never match; the note is corrected in place. -
One unreadable file sank the whole corpus. A fragment root is foreign by design, but a single binary or latin-1 file in it exited 3 from every command —
flowsincluded, which never looks at a fragment. Read failures are now per-file diagnostics (pack::unreadable_fragment_file) that never sink their siblings and stay silent until somethingref:s the corpus, matching what pack loading and the annotation scan already did. -
schema --add-torewrote fragment files. It prepended a yaml-language-server modeline to a.hurlcorpus file and dropped the pack schema beside it — violating ADR-0018’s “fragment files are inputs proef never writes”. It now refuses anything that is not a pack, reusing theis_pack_filepredicatefmtalready had. -
A
#in an annotation name was accepted but unreachable.#separates a file from a fragment inref: file.hurl#name, so such a name could be declared and never referenced — and the failure suggested the exact spelling that had just failed. Refused at scan time. -
proef lspanswered every URI-keyed request withnullon Windows. A source name is an identity compared as a string, and the two sides spelled it differently:Path::joinappends without rewriting what is already there, so aproef.tomlsayingsuite = "tests/features"— the portable spelling the docs use — producedC:\proj\tests/features\packs\api.yamlfrom discovery while the client’s document URI producedC:\proj\tests\features\packs\api.yaml. The two never matched, so go-to-definition, find-references and completion all found nothing while the suite itself ran green. Discovered names are now rebuilt in native form. Unix has one separator and was never affected, which is why every gate stayed green. -
A fragment’s path was absolute everywhere it was named.
[run] fragmentsresolves against the config file’s directory, sofragments = "tests/hurl"became/home/you/project/tests/hurl— and that spelling then named the file in every diagnostic and, once steps recorded their provenance, in the run record too. Feature and pack names are project-relative because the path the author typed was; a path the author never typed had no such luck. Records went machine-specific: the same suite on two checkouts stopped comparing equal, and a temp-dir path could reach a durable artifact. The root is now shortened back to a cwd-relative spelling when it is under the working directory — resolution is untouched, so which file gets read never changes. -
Every
ref:was an error in the editor while the same suite ran green.SourceProvider::discover_fragmentsshipped with a defaultOk(Vec::new()), and the LSP’s overlay provider — which forwards feature and pack discovery to disk — never overrode it. So the analyzer saw no fragments at all: go-to-definition on aref:did nothing,ref:completion returned nothing, and everyref:rendered asproef::pack::unknown_ref. Exactly the diagnostics-you-cannot-trust drift the fragment-aware analysis was added to prevent.The default is gone;
discover_fragmentsis a required method. Every implementation lives in this workspace, so the default bought no compatibility — it only let a forwarding provider inherit “no fragments” silently instead of failing to compile. An integration test now drives the real provider chain and asserts aref:jump lands on the annotation in the.hurlfile. -
A fragment file saved with a BOM failed at line 1, blaming the request. Every other text entry point (
feature::parse, the inline-payload probe) strips a leadingU+FEFF; the fragment scanner did not, so the mark reached hurl’s parser as the first character of the first request. The file is now normalized by the same rule, and the mark cannot travel into an artifact that has to be valid hurl. -
A macro-scope
bind:with noref:step was silently dropped. The step-scope version of this mistake has been a hard error sincebind:landed; one scope up it vanished at lower time. That is the half authors actually hit, because factoring plumbing upward is the habit — and the tempting reading, that ause:target will pick the table up, is wrong: the child resolves its own scopes. Nowproef::pack::bind_without_refat both scopes, with a message that says so. -
A
ref:step’sname:reported a${fake:…}value it never sent. A label is a replay of what the request was built from, not a fresh use of it: the inline path rewinds the${fake:…}occurrence counter, resolves the label, then restores it to the high-water mark. Theref:path reproduced that tail without the rewind, so a step binding${fake:email}and naming${fake:email}minted two identities — the console and the event stream announced one address while the request sent another, and every later step’s fake values shifted by one. Both body forms now end in one sharedfinish_step, so the rule is stated and enforced in a single place rather than copied. -
An escaped
$${secret:…}in abind:value was refused as a composite.$${is the escape (ADR-0005), so$${secret:token}is the literal text${secret:token}and names no secret — but the composite check searched for the substring"${secret:", matched at offset 1, and rejected the binding withsecret_in_composite_bind. Both the whole-value and composite tests now read the value through the resolver’s own reference scanner, so there is one thing that knows what a${…}is and$${stays an escape everywhere. -
A step that set
retry:twice ran the value it did not name. A pack could declareretry:(ordelay:) as a step key and again inside the block’s own[Options]. Lowering extends an author’s existing section rather than opening a second one, so proef’s baked line landed above the author’s; hurl resolves a duplicated option last-wins, and the raw value therefore won every time. The pack saidretry: 10, the run didretry: 3, and nothing anywhere said so — the finite-retry lint only ever looked for-1and over-cap counts, so a plausible finite value passed untouched. Declaring an option in both places is nowproef::pack::option_declared_twice, refused at load with the span on the raw line that used to take effect.The scan is deliberately scoped to
[Options]sections rather than matching anyretry:-shaped line:retryis a legal request-header name, and a header isname: valuelike an option is, so a line-shaped match would have turned an ordinary header into a hard error. Pinned by a test that a header namedretryon a step carrying a typedretry:still loads.
[0.9.0] - 2026-08-11 (tool-surface integrity & authoring guidance)
Breaking:
proef secret set --valuewas removed in favour of--stdin, andproef macros --output json’spatternfield changed from a boolean tostring|null.
Added
-
The run record says which scenarios were lifecycle phases.
phase("setup"/"teardown") is now onscenario_started/scenario_finished— additive and optional (ADR-0008), so older records read as “no phases”, which is what they had. Without it a teardown scenario was indistinguishable from a suite one except by feature path, so every consumer re-derived phase membership fromproef.tomland three of them got it wrong in different ways. Fixing them off one signal is what the three entries below have in common. -
proef doctorreports a missing pack schema.initinstalls it automatically, but noticing when it is absent never shipped — so a suite whose editor completion had been silently off had nothing telling it so. Reported as a warning, never a failure: it costs autocomplete and load-time validation in the editor, not a run, anddoctor’s exit is the environment verdict. Uses the same predicateinituses, so the two cannot disagree about what “installed” means. Runs outside a project too — no config or no suite is reported, not failed. -
bind::unbound_stepnamesproef macrosagain, from the CLI. The pointer was removed from the diagnostic in #25 for a correct reason — that text also renders in an editor’s diagnostics pane through the LSP, where the affordance is completion, not a command — but nothing put it back on the terminal side, so a terminal reader saw it zero times. It is now added by the CLI’s own renderer, which legitimately knows it is the CLI. The core diagnostic still names no tool. -
proef macrosanswers when the suite does not bind. Listing the vocabulary previously required every scenario to bind — so the command refused in exactly the situation that sends an author looking for it: a step that matched no macro. It now prints the diagnostics, then the vocabulary the packs offer, and keeps its exit code unchanged (2), so scripts see no difference. Pack loading precedes binding and does not depend on it, so the listed vocabulary is complete. Every count-derived verdict is withheld in that mode —calls/unusedrender as—/nullrather than0/false, because a feature that failed to bind contributes no calls and would otherwise make its own macros look dead.proef flowsdeliberately still refuses: its contract is to list every scenario, and a partial list that silently omits the unparsed feature is the wrong answer, not a degraded one. -
A failed run says when the suite is still the untouched scaffold. A freshly scaffolded project cannot pass — its target and its routes are both placeholders — and
initsays so once, two commands earlier, in a parenthetical the failure never referred back to. The run now names the situation and the remedy. It fires only on the conjunction ([url] basestill byte-identical to whatinitwrote and noPROEF_BASE_URL): an operator who set the override did name a target, so their failure is about their API and is not second-guessed. Exit codes are untouched — whether an unreachable target is a user or a system fault is a taxonomy question decided in the engine (ADR-0009), and the reader’s actual problem is vocabulary.
Changed
-
proef secret set --valueis gone; use--stdin. Breaking. A secret in argv is visible to anyone who can runps, and the failure path steered people to it — the hidden prompt’s error said “pass--valuein scripts”, which fires exactly in the non-TTY/CI case where the exposure matters. There is now no flag that takes a value:--stdinreads it from a pipe (same shape asdocker login --password-stdin), stripping the trailing newline the pipe added, and the prompt stays the default. Scripts using--valuemust pipe instead:printf %s "$TOKEN" | proef secret set NAME --stdin. -
proef macrosprints the sentence, not just the identifier. A test author writes prose that binds to a vocabulary somebody else maintains — and the one command that lists that vocabulary showedhealthwhere the author needsthe service is healthy. Thematch:pattern was already loaded and already linted; both renderers discarded it on the way out. It now appears in the text listing, and--output json’spatternfield carries the string itself (nullwhen a macro isuse:-only) instead of a bare boolean.
Fixed
-
proef fmtrefuses a file that is not a pack. It took an explicit path on trust, so it rewrote whatever it was pointed at:proef fmt src/main.rsstripped trailing whitespace from Rust source, printedformatted:, and exited 0. A mistyped path was a silent edit. Formatters parse before they write and refuse what they cannot parse; this one locates blocks textually, so the extension is the check available — and it is now the same predicate discovery already used, rather than a second opinion about what a pack is. Only the explicit-file path was affected: a directory was always filtered. -
proef fmtleaves the YAML skeleton alone, as it always said it did. Its documented scope is hurl blocks — the module doc promises the skeleton, comments included, is never touched, and the code claimed the trailing newline was the only normalization applied outside a block. Both were wrong: every line was trimmed. A pack whose blocks were already canonical failedfmt --checkon nothing but a trailing space in a comment, which is a CI red an author cannot explain from the documented scope. This is the same over-reach the line-ending fix removed in 0.8.0, in the same function, one line above where that fix landed. -
A truncated record no longer drops a warned scenario from its totals. With no
run_finishedto read,explainrecounts the scenarios present — and countedPassed/Failed/Skippedbut notWarned, so a scenario whoseoptional:step warned vanished from every column. The live path countsPassed | Warnedtogether (RunSummary::passedis “passed, warnings allowed”), so the reconstruction silently disagreed with the run it was reconstructing — andoptional:exists precisely so a scenario can warn and still pass. -
A failing run says when the scaffold’s routes are still placeholders. The scaffold has two halves to fill in, and a reader can have done either. Someone who follows
init’s instruction — point${url:base}at your API — then hits the other half:/healthand/search404, and the target-side note deliberately cannot fire, because they did configure a target. They had been told about the routes once, parenthetically, two commands earlier. Now they are told at the failure. Decided from the pack’s bytes, never from what the server answered: a 404 proves a route is missing, not that it is a placeholder, and inferring the second from the first is the class of claim removed in 0.8.0. The two notes are mutually exclusive — a reader with one unfinished half is told about that half, not handed a list. -
--dry-run’s “next” command is the run that was validated. After--dry-run --env prod --tags smokeit printed a bareproef test, which is a different run — another[url] basefrom the profile, and every scenario rather than the tagged subset. The operator could not tell: the command works and simply tests something else. Every selector that chose what ran is echoed now (--env,--tags,--scenario,--scenario-file, and the path), quoted so a tag expression or a scenario name with spaces survives a paste. Deliberately selectors only — a general “reprint the invocation” is how secret-bearing arguments reach stdout. -
--sarifemitsstartLine. GitHub keys inline annotations on it, so a log carrying onlybyteOffset/byteLengthuploaded cleanly and annotated nothing — the flag looked wired up and delivered none of what it advertises. Sources are read once each at the IO edge and only to count newlines;Diagkeeps carrying byte spans, and no column arithmetic is introduced. -
--watchretriggers onproef.toml. It watched the suite path recursively, and the config lives above it — so editing a[url]/[vars]/[env.*]value that every scenario resolves through changed nothing, which reads as the watcher being broken. Matched by exact path rather than by a.tomlextension, so an unrelated manifest in the tree still does not requeue. -
Three places interpolated a value into a format without escaping it. Same shape each time, so they are fixed together:
- LSP completion snippets.
$,}and\are LSP snippet syntax, and amatch:pattern is prose — prose carries$.the price is $5made the client read$5as tabstop 5 and drop the text, so accepting the completion inserted something the author never wrote. Literal characters are escaped now; the tabstops the generator writes stay syntax. - GitHub annotations.
file=was passed raw whiletitle=and the message beside it in the samewriteln!were encoded. A path carrying,or:— every Windows path carries a:— broke thekey=value,key=valueparse. - The GitHub job-summary table. The scenario name and file went into
Markdown cells unescaped; a
|in either ends the cell and shifts every column after it, and the row still renders, which is why it goes unnoticed.
- LSP completion snippets.
-
A templated
retry:/delay:/repeat:/max-time:no longer under-counts the batch budget. The estimator matched literal values only, so a{{var}}-driven option fell through and read as no retries — the budget was then computed for a single attempt, and the watchdog abandoned a scenario that was retrying exactly as authored, reporting it as an environment fault (exit 3). A placeholder resolves inside hurl at run time and cannot be estimated, so the engine now says so:batch_budgetreturnsNone, whose contract already routes the batch to the orchestrator’s default budget. An infinite count is treated the same way, since it is unbounded by definition.TROUBLESHOOTINGdescribed the old behaviour as if the budget could see these values; it now says what actually happens. -
--output json’sexit_codeis the code the process exits with. A failed JUnit write escalates the run to 3, and that escalation was applied by areturnafter the body had been printed — so a machine consumer read a verdict the program then exited past, with nothing to signal the disagreement. The escalation is now folded in before anything serializes it. -
proef fmtkeeps each line’s own ending. Its scope is hurl blocks, not line endings, but it split the whole file withstr::lines()— which throws the terminator away — and rejoined with a single one. A file mixing CRLF and LF was therefore homogenized, andfmt --checkcame back red on a pack whose blocks were already canonical. The earlier fix moved from “always LF” to “the dominant ending”, which still rewrote the minority lines. Terminators now travel with their line, so an untouched line is written back byte-for-byte; the only endingfmtstill supplies is a trailing newline on a file that lacked one. -
proef --helpdescribesmacrosas it now behaves. It still said “with its call count” after the command started printing the sentence each macro binds — the README table was updated and the clap text that actually produces--helpwas not. -
proef lspadopts the workspace root the client announces. The root was resolved at the process edge, before the handshake, from the working directory — so an editor launched anywhere but the project analysed the wrong tree, andnvim ~/proj/x.featurefrom$HOMErooted the analyser at$HOME. Theinitializeparams were bound and discarded. The server now readsworkspaceFolders, falling back torootUri(deprecated since LSP 3.16, and the spec is explicit that folders win when both are present) and then to the previous config-then-cwd resolution.proef-lspstill knows nothing aboutproef.toml: it calls back into the CLI, which owns config (ADR-0012). -
A mixed suite+phase failure kept the phase label.
explainchose the label from the whole report (failed == 0), so it appeared only while every failure was a phase failure — and vanished the moment a suite failure joined one, leaving1 failedabove two indistinguishable blocks. The disambiguation disappeared exactly where it was needed. Labelled per block now, from the record. -
--rerunafter a phase-only failure says there is nothing to rerun. It returned the failed teardown, whichbuild_specscannot match because the phase is excluded from the pool — producing a run that matched nothing and reported “no scenarios matched the filters (check –tags/–scenario)”, naming flags the operator never passed. Phases are invisible to--rerun(ADR-0014); it now exits 0 saying so. -
diffno longer counts a failing teardown as a test regression. A cleanup fault makestestexit 3, not 1, so blending phases into the regression buckets madediff --fail-on-regressioncontradict the run it was diffing. Phase scenarios are excluded from the verdict and the exclusion is reported. -
Records written before 0.6.0 no longer report the wrong verdict with confidence. They carry one
run_finishedper phase and their totals counted every phase; read under today’s suite-only meaning, a genuine suite failure was reported as1 passed · 0 failedand labelled setup/teardown. Theschemafield cannot distinguish them — that change was semantic and never bumped it — but the structure can.explainnow detects the multiple pairs, recomputes the totals from the scenarios present, and says the record predates 0.6.0. A reader must be able to consume a record or detect that it cannot; quietly doing neither was the one unacceptable option. -
proef initno longer destroys aproef-pack.schema.jsonyou wrote. The never-overwrite loop walks a fixed four-entry array; the schema is not in it, and is written afterwards by the shared installer. So the one unguarded path was pack-absent + schema-present:initscaffolded the pack, then the installer replaced an authored file — reported ascreated 5 file(s), skipped 0, while the README promised the opposite in as many words.initnow asks the installer to preserve what is already there and reports it as skipped;proef schema --add-tostill refreshes, since that is an explicit install and how the schema is updated after upgrading proef. -
The first-run note no longer fires on real suites. It keyed on
[url] basestill equalling the valueproef initwrites — which looks init-specific and is not:GETTING-STARTEDteaches that exact line to people building a suite by hand, and proef’s ownproef.tomluses it. So a hand-built suite whose server was up and whose assertion genuinely failed was told “this suite is still theproef initscaffold — its target and its routes are placeholders, so it cannot pass yet”: every clause false, moments after the suite reached a real verdict. The deciding evidence is now the run itself — the note appears only when nothing was reachable (no scenario passed and every outcome is a system fault). A suite that got an HTTP response, even a 404, has a target; whether its routes are placeholders was a guess, and the note stated it as fact. Wording softened accordingly. -
Suite discovery no longer walks build output, and one unreadable directory no longer empties the suite. The walk had no exclusions, no depth bound, and a
canonicalize()per directory — and it re-runs on every language-server request, so enteringtarget/cost that price over and over for a subtree that cannot contain a suite. It now skipstarget/,node_modules/,vendor/and dot-directories (tested on children only: a suite may legitimately be rooted at such a name), and refuses beyond 32 levels rather than recursing until the stack runs out. APermission deniedon one descendant used to abort the entire walk, andproef lspswallowed that error into an empty analysis — so a single unreadable subdirectory silently emptied the suite. Unreadable descendants are now skipped, the wayfindand ripgrep do; an unreadable root is still a loud error, because that path is the caller’s own. -
Ctrl-C no longer skips cleanup in silence. Teardown shared the run’s cancellation token, so an interrupt left every teardown scenario
Skipped— and because a skipped phase carries no fault, the worst-wins fold passed it without a word. Whatever setup created stayed created and nothing said so, against this ADR’s own premise that suite cleanup is reliable. Teardown now runs on its own, independent token (notchild_token(), which cancels with its parent and would have re-implemented the bug): the pool stops at its batch boundary, the operator is told cleanup is running, and it completes. A second Ctrl-C still hard-exits (130) — the escape hatch ADR-0007 relies on — and the announcement says so. Amends ADR-0014. -
A phase that only skipped is now a failure, not a pass. That silence was the shape that hid cancelled cleanup. A setup completing no scenario aborts the run rather than letting the suite execute against state setup never created — which is also what keeps teardown gated on setup-success, since the abort is the gate; a teardown completing no scenario is reported and fails.
-
--dry-runvalidates[run] setupand[run] teardown— which ADR-0014 always claimed (“validated like any other feature but never executed”) and nothing did:--dry-runnever read the keys. A broken teardown therefore surfaced only after a full suite had run — real requests, a run directory, artifacts — while the identical mistake insetupfailed in milliseconds. Both are now validated by one loader shared withproef test, which also pre-flights teardown before the pool. A bad phase path is a user error (exit 2) rather than a blanket system fault (exit 3), and creates no run record. -
proef schema --add-toandproef initnow announce the schema file they write. Both wroteproef-pack.schema.jsonsilently, soinitlisted four files and then reported “created 5 file(s)” — the first output a new user reads, not reconciling, with the unannounced file being the one that powers editor completion. -
proef initno longer sends you to install editor completion that is already installed. A re-run namedproef schema --add-tounconditionally, even with the schema sitting beside the pack. It now says which of the two situations you are in. -
The nextest harness no longer reports green having listed no tests. A
PROEF_HARNESS_SUITEset to bytes that are not valid UTF-8 read as unset, which the harness treats as “expose nothing” on purpose — socargo testpassed having run zero scenarios. APROEF_BINit could not read fell back toproefonPATH, silently invoking a different binary than the one named. Both now surface as a failingproef::configtrial, the same loud shape the harness already used for flows-contract drift, whose comment states the invariant this violated: never run zero tests green.
Documentation
-
AUTHORING shows how to write a validation-error catalogue. Two patterns that were reachable but not signposted, and that compose into one. A validation suite’s cases differ structurally — one omits a key, one empties it, one adds a key the caller may not set — so a single parameterised macro cannot express them and an
Examplescell cannot practically hold JSON; the answer is one named macro per malformation, whose sentence says what is wrong in business terms. The expectation side then does not grow with the catalogue: because anexpect:merges into the previous request entry, one parameterisedthe error code is {code}covers every case in the set, typically the largest de-duplicator in a validation pack. That merging was documented as a mechanism in two sentences and never shown as the pattern it is. The cost is stated rather than hidden — the pack grows with the catalogue, which is what buys feature files a non-engineer can review. -
An outline’s
<column>placeholders substitute into the docstring, and AUTHORING now says so. They always have — TECH-SPEC §4.4 specifies it and the code has done it since — but the author-facing guide named only step text and table cells, andStepDefn’s own doc comment named the substitution ontextandtablewhile describingdocstringas just “raw request bodies”. Naming it twice and omitting it once reads as a deliberate exception, so a reader concludes the opposite of the truth: this is exactly the capability an author reaches for to data-drive a request body without leaving the feature file. AUTHORING gains a worked example. Pinned by tests for the first time — every other outline test asserts on step text, so a regression would have emitted a literal<label>into an artifact with the suite green. -
The docs-drift backlog is closed.
EDITORS.mdsaid go-to-definition cannot land on amatch:line — it has since 0.5.1, anddefinition_on_a_step_lands_on_the_match_lineproves it; the bullet now names the gap that is real (built-in macros live in a pack compiled into the binary, so there is nothing to open). TECH-SPEC §10’s command surface gained--run-id/--rerun/--sarif.GETTING-STARTEDno longer shows a scaffold comment with a word the scaffold does not write. ADR-0015 described aworkeronScenarioFinishedthat is alwaysNone, because that event is emitted from the dispatcher thread rather than the worker — an errata records what shipped, whichEVENTS.mdhad right all along.Two entries did not reproduce and are recorded as such rather than dropped:
CONFIG.mdcarries no claim that[env.<name>.run]overrides any section, and the 0.5.2 changelog does mention the directory-valued-phase error. -
WRITING-SCENARIOS’s two sample outputs match the binary again. Themacrossample showed two builtins with no ellipsis and omitted the(builtin, unused here)marker and the trailing count; themissing_config_varsample dropped the(or in the active [env.<name>.url])clause. Both read as verbatim transcripts, so a reader comparing them against a real run found differences that were the document’s, not theirs. -
One worklist instead of four documents to cross-read. Four files read like backlogs and only one was:
OPEN-FINDINGSnow carries every open item, including the residue of both UX reviews (R1–R3) and the decisions taken against them, each entry self-contained. The two review documents were removed once their open items landed there — their transcripts and citations remain in git history, and a retired review left on disk is exactly the thing that reads as a backlog.IMPROVEMENT-PLANstays a separate file — five ADRs cite it by section number — but its master table gained a Status column, because its ✅/⚠️ glyphs mean “fits the architecture”, never “done”, and 13 of its 16 items had already shipped while the table gave no way to tell. Item 14’s cited mechanism (Refs::default()resetting perlower()call) was corrected: 0.6.0 replaced it, and only the cross-scenario half of that caveat still holds. -
A page for the persona the product is named after. PRD §4’s first persona writes prose against a vocabulary somebody else maintains — and every document labelled “test authors” taught pack authoring, so that reader had no route through the tool.
docs/WRITING-SCENARIOS.mdcovers only their loop: what a sentence is, how to list the ones available, the dry-run cycle, and the two diagnostics they will actually hit. The index now labels each author-facing page with the persona it serves instead of calling six P2 documents “test authors”. -
bind::unbound_stepleads with the action its reader can take. The help opened on “add a macro to a pack” — the pack maintainer’s move, which a scenario author cannot make — and buried theirs in a parenthetical. It now opens with matching a sentence the suite’s packs already bind. It names no tool:Diag.helpreaches an editor’s diagnostics pane verbatim through the LSP as well as the terminal, and each front end already has its own way to show the vocabulary (completion in the editor,proef macrosin a shell) —proef-coredoes not know which one is reading. The YAML stub is unchanged: it is load-bearing for the maintainer and stays verbatim. -
ADR-0014 now records the question it was silent on. It is specific about a failing setup and a failing teardown, so a reader reasonably infers the cancellation case was considered — it was not. What teardown does on Ctrl-C is unspecified, and today it silently skips: the phase runs with the already-cancelled token, every scenario resolves
Skipped, andphase_failedignores a phase that only skipped, so cleanup never runs and nothing says so. The ADR now states the gap and the two defensible answers, since an implementer working on teardown reads the ADR, not the findings list. -
The open-findings list is now in the repo, not on one machine. A v0.5.3 review was validated claim-by-claim (40 claims, 38 confirmed) and the record lived only in a gitignored scratch directory, so ~26 still-open defects — the Ctrl-C teardown gap, LSP rooting,
--sarifline numbers, several docs drifts — existed nowhere durable.docs/OPEN-FINDINGS.mdcarries them, plus what shipped against them, so a fixed finding is not re-reported and an open one is not lost. -
proef initis now in the command tables it was missing from. It shipped in 0.6.0 and was documented inGETTING-STARTED.mdand in the README’s prose, but not in the README’s CLI table orTECH-SPEC.md’s command surface — so the two places a reader scans for “what can this tool do” both omitted the command that starts a first run. -
CLAUDE.md’s status list now records the v0.6.0–v0.8.0 correctness series rather than ending at post-M5, so the three releases that closed the reports-success-on-wrong-output bug class are visible to anyone picking the project up.
[0.8.0] - 2026-08-09 (CLI output & exit integrity)
Changed
- A set-but-unreadable environment variable is now a loud user error, never
silence — breaking for a pipeline that relied on the old silent fallback.
std::env::varcollapses “unset” and “set to bytes that are not valid UTF-8” into the sameErr;.ok()erased that distinction at five call sites, so a value proef could not read was indistinguishable from one the user never set. A non-UTF-8PROEF_KEYfell through to the key file and decrypted with the wrong key, reporting tampering instead of the real cause (anddoctorreported the key source as the file instead of the override); a non-UTF-8PROEF_SECRET_<NAME>fell through to the store and reported a missing secret; a non-UTF-8PROEF_ENVran silently against the wrong environment, including inproef lsp, where it meant analysing against the wrong config profile. Four of the five sites now exit 2 (user error) naming the variable;doctorinstead reports it as a failed check alongside its other unready-environment findings and exits 3, the same as an unreadable key file. A pipeline that today tolerates a mis-setPROEF_ENV, or a non-UTF-8 key/secret, will start failing after this upgrade. - A failed stdout write now reaches the exit code — breaking for a pipeline
that tolerated truncated output. Writing to a full disk or other failed
stdout exited
0with truncated output; it now exits3. A closed pipe (proef … | head) still exits cleanly. A pipeline that captures proef’s stdout somewhere that can fail mid-write (a full disk, a device error) previously reported success over truncated output; it now gets a nonzero exit it can act on instead of trusting truncated bytes. Perdocs/RELEASING.md, any breaking change is MINOR — together with the environment-variable change above, this forces the next release to be 0.8.0, not 0.7.1.
Fixed
-
proef fmtrewrites line endings wholesale, violating its hurl-blocks-only promise.fmtsplit pack files withtext.lines()(which strips both\nand\r\n) and rejoined with hardcoded"\n", so CRLF files became LF. On anautocrlfcheckout (a supported way to clone this repo),fmt --checkwas permanently failing through no fault of the author.fmtnow detects the file’s dominant line ending and preserves it when rewriting. -
run.logcould gain duplicated fragments when the console accepted a short write, because the tee re-wrote the full slice on every retry. It now mirrors only the accepted bytes. -
proef report -ooutside the run dir wrote artifact links relative to the run dir, so every link 404’d from the report’s own location while the command reported success. The href is now absolute when the report is written elsewhere. -
proef diffreported a brand-new retried step as newly flaky, because a step absent from the base run was assumed to have run once. Steps with no baseline are now skipped, and the ordinal-shift caveat inherent to positional step keying is documented in TROUBLESHOOTING.
[0.7.0] - 2026-08-07 (record & artifact integrity)
Changed
${fake:…}values no longer repeat across a scenario’s steps. The occurrence counter restarted on every step, so two steps each asking for a fresh${fake:email}received the same address. Every independent${fake:…}reference within a scenario — across steps, and within one step’s payload/when:/label — now gets its own value and never collides with another, however many a single step ends up resolving. A step’sname:label (shown in artifact comments and events) is the deliberate exception: it is not independent of its own payload, so it replays from the start of the step’s own occurrence window instead of minting new ones, matched by position (the label’s Nth${fake:…}reference reuses the payload/when:’s Nth occurrence, regardless of generator kind) — so it reproduces the payload’s own value when the label’s references mirror the payload’s in kind and order, and shows a different generator’s output when they don’t. Even a label with more${fake:…}references than its payload still reserves each extra one, so a later step can never be handed a value the label already displayed. Values remain deterministic for a given--run-id, but suites using${fake:…}will see their emitted artifacts change. Known limitation, not fixed here: the counter resets at the start of every scenario, not the run, so two different scenarios that each resolve${fake:email}at the same position in their own step order still collide — that is a separate bug with its own snapshot-moving fix.proef_core::resolve::resolvechanged signature (public API break for downstreamproef-coreconsumers): it now takes an additional&mut usizeoccurrence counter supplied by the caller, andResolution::fakeswas removed —resolve()no longer owns the counter itself.
Fixed
-
run_finishedis once again the last line of a run record. A scenario the watchdog abandons keeps running on a detached thread and only notices its cancellation token at the next batch boundary, so it went on appending events after the sweep had recorded its outcome — and after the run itself was finalized.docs/EVENTS.mdhas always said the last line isrun_finished; it was not, so anything reading a record as a stream (the JSONL consumer,report,explain) could see events arrive after the terminal one. Late events from a finalized scenario are now dropped at a single gate rather than by asking every emitter to check. Abandonment itself is unchanged and stays cooperative (ADR-0007) — only the record’s tail is affected. -
.map.jsonno longer loses a request’s captures when the pack comments one of them. A comment inside an open[Captures]run is the author’s note about a capture, not the start of the next entry, so it no longer closes the scan — previously it dropped every capture after the comment. The entry that follows opens with a method or response line, and that closes the run on its own. -
.map.jsonno longer lists captures that were never made. The sidecar’s capture scan was fence-unaware — a literal[Captures]line inside a fenced (…) body re-armed it — and it recognised only the stock HTTP methods, so an entry opened by a custom method (PROPFIND, …) never ended the previous scan. Both let capture names that don’t exist in the emitted entry land in.map.json, a normative artifact (ADR-0010). The scan is now fence-aware and shares the lowering pass’s method recogniser (is_method_line) instead of carrying a second, weaker copy. -
pack::empty_expectnow also catches a whitespace-onlyhurl:fragment. The diagnostic already existed for anexpect:item with neitherstatus:norhurl:at all; ahurl:key present but carrying no non-blank assert line slipped past it, lowered to an empty asserts block. It also gains a remediation hint and the seeded corpus case it was missing. Scope: this check reads the unresolved pack text, so a fragment that is non-blank as authored but resolves to nothing at lower time (e.g.${vars:key}naming aproef.tomlvalue that is""in the active environment, or an unset${global:key}under--dry-run) still lowers to an empty asserts block — see the sidecar-emitter entry below for how that residual case is handled. -
The sidecar emitter can no longer produce an inverted
.map.jsonspan. AThenstep whose asserts all resolved to nothing — reachable even after thepack::empty_expectwidening above, since pack validation cannot see what a fragment resolves to, only what it says — lowered to a zero-line merged-asserts step, and the emitter’s line-span arithmetic underflowed: the start offset exceeded the end. Such a step now gets no sidecar row at all instead of an inverted one — nothing was appended to the artifact, so there is nothing to report a span for.
[0.6.0] - 2026-08-07 (first-run UX & run-record correctness)
Added
proef initscaffolds a working suite. It writes the filesGETTING-STARTED.mdteaches —proef.toml, one.feature, one matching pack — installs the pack JSON Schema for editor completion, and prints the next command. Nothing is ever overwritten, so a second run is a no-op and no--forceflag exists to destroy authored work. A test asserts the scaffold passes--dry-rununchanged.- The README now shows a parameterized macro and states the load-bearing non-goals, including the supported path for teams that already have a hurl corpus.
Changed
- A passing
--dry-runnow names the next command. Every failure path already named a remedy; the success path stopped talking at the moment a new user decides whether to continue. - A scenario with no steps is now an error, not a silent pass — breaking.
A
Scenario:with a commented-out or never-written body previously bound to nothing, ran nothing, and exited 0; it now exits 2, throughproef test,proef flows, the libtest-mimic harness, andproef-lsp(which re-analyzes ondidChange, so a half-typedScenario:now shows a live error while you’re still typing it). Perdocs/RELEASING.md, any breaking change is MINOR — this forces the next release to be 0.6.0, not 0.5.4.
Fixed
resolve::missing_config_varnow suggests the closest key defined in the same namespace, matchingresolve::unknown_variableandresolve::fake_unknown. Candidates are namespace-scoped, so a${url:…}typo can never suggest a[vars]key. The code also gains the seeded corpus case it was missing.proef initno longer rewrites a pack it declined to create. Installing the editor modeline ran unconditionally, so a hand-authoredsuite/packs/api.yamlreported as “already exists” was still modified; the schema install is now gated on the file having been created, and an existing pack gets a hint namingproef schema --add-toinstead.- Setup and teardown no longer corrupt the run record. Each phase bracketed
its own
run_started/run_finished, so one record held up to three pairs andproef explainreported the last phase’s totals — printing “1 passed · 0 failed” above a failure it had just listed. The record now carries one pair, and itsrun_finishedtotals are the main suite’s own verdict —[run] setup/teardownscenarios still appear as their own events in the record, but are never folded intopassed/failed/skipped, so those numbers agree with the consolesummary:line, JUnit,--output json, TAP, the SLA gate, and the exit code. The console run header also prints once per run instead of once per phase. reportandexplainflag a truncated run. Both rendered an incomplete record as if it were whole;explainalso derived its headline solely from the missing tail event, reporting all zeros for a record that held completed scenarios. Both now read through the same record readerdiffuses.explain’s step/attempt totals count a still-in-flight scenario. A step only attached to the record once itsScenarioFinishedlanded, so a scenario still running when a truncated record’s stream ended had its step evidence silently dropped from the headline — the one place a post-mortem tool most needs it. Totals now fold the raw events directly instead.explain’s failure detail is keyed(file, scenario), not scenario name alone. Two same-named scenarios in different files previously bled each other’s failure output together.workeris the slot a scenario occupied, not a per-scenario counter. The timeline drew one lane per scenario regardless of--jobs.- Run rotation only treats hyphenated UUID directories as run records. The
parser also accepted bare 32-hex,
urn:uuid:and braced spellings, which rotation could then delete when the runs directory points somewhere shared. - The nightly canary can fail again: its step piped through
teewithoutpipefail, so a red canary exited 0 and the open-an-issue step was unreachable. - The raw-print-macro guard now covers
proef-lsp, where stdout is the JSON-RPC channel and a stray print corrupts protocol framing.
Documentation
- The stdout/stderr macro rule is now written down where contributors look:
docs/CONTRIBUTING.md(“Rules that are easy to trip over”) andCLAUDE.md. 0.5.3 began enforcing it with a source-scanning test, so a rawprintln!oreprintln!inproef-clifailed the suite with nothing explaining the rule or namingrender::outln!/errln!as the sanctioned spellings.
[0.5.3] - 2026-08-06 (closed-pipe safety)
Fixed
- The CLI no longer panics when stderr is a closed pipe. Every remaining
raw
eprintln!inproef-clinow routes through the EPIPE-safeerrln!guard added in 0.5.2, soproef test … |& headends the pipeline with the contracted exit code instead of aborting with 101 — a code outside the typed 0/1/2/3 taxonomy (ADR-0009). The execution failure summary, which writes several lines per failing scenario, was the largest remaining exposure. A source-scanning test now keeps raweprintln!out of the crate. - The language server no longer dies while recovering from a panic.
proef-lspreports a caught analysis panic on stderr; that report used a raweprintln!, which panics when its write fails — so a closed stderr (EPIPE) took down the very server the surroundingcatch_unwindexists to keep alive. The write is now explicitly unchecked. Ships without a test: reaching the line needs a real analysis panic and a closed stderr, and the panic is not injectable without a test-only hook in shipping code; the mechanism itself is already covered by the CLI’s closed-pipe tests.
Changed
proef reportderives its output directory through the sharedfsutil::parent_dirhelper instead of an open-coded empty-parent fallback, so there is one spelling of that derivation. Internal consistency only — the emitted artifact links are unchanged.
[0.5.2] - 2026-08-05 (CLI correctness)
Fixed
- A directory-valued
[run] setup/teardownis now a loud user error. ADR-0014 defines setup/teardown as a single feature file; a directory ran every feature under it as the phase and again in the pool (a silent double-run) — that path is closed. - Diagnostics no longer panic when stderr is a closed pipe:
print_allandreport_front_error’s trailing"{errors} error(s)"summary line are now routed through an EPIPE-safeerrln!guard (mirroringoutln!’s stdout guard), soproef test --dry-run <broken suite> |& headexits cleanly instead of panicking (exit 101). diffstep records are now keyed by(text, occurrence ordinal)instead of text alone — macro-expanded steps that share text no longer collide in the last-write-wins map and silently drop out of the diff.diff --fail-on-regressionnow fails when the new run is incomplete or cancelled (was a silent pass), and banners any incomplete/cancelled record in the diff output either way. Its slower-step duration math is hardened against overflow (saturating arithmetic).- A bare-filename
[run] setup/teardown(or suite path) now resolves its packs and assets from the current directory. A path with no directory component (e.g.setup = "setup.feature"at the project root) has an emptyPath::parent(), which produced acannot read directoryfailure; it now normalizes to.(the current directory) via a sharedfsutil::parent_dirhelper at the pack/asset base-derivation sites.
Documentation
- The second-interrupt hard-exit code 130 (128+SIGINT) is now documented
for
testandwatch(TECH-SPEC §10, ADR-0009) — a deliberate escape hatch outside the typed 0/1/2/3ExitCodetaxonomy.
[0.5.1] - 2026-08-05 (LSP go-to-definition + correctness)
Added
- LSP go-to-definition:
use:references andmatch:landing (ADR-0017). Go-to-definition now jumps from ause:reference in a pack to the macro it targets, and lands on the macro’smatch:line rather than its name key (falling back to the name key for use-only macros with nomatch:).
Fixed
- LSP: the stdio server now exits cleanly.
proef lspdropped the connection after joining the transport threads, so the writer thread (holding the sole channel Sender) never ended and the process leaked. It now drops the connection before joining. Covered by a real stdio subprocess lifecycle test. - LSP: a malformed request no longer crashes the server. A bad document URI or
out-of-range position propagated a deserialization error out of the event loop
and exited the process; the request now gets an
InvalidParams(-32602) reply and the server keeps serving. - LSP: one broken pack no longer blanks the whole suite.
analyze_suitenow keeps the packs that loaded (and reports the broken one’s diagnostic) instead of zeroing all bindings, completion, and go-to-definition on any pack error. - LSP: analysis is scoped to the configured suite. The server roots at
[run] suite(else thetests/convention) under its launch directory rather than walking the entire working tree, sharing the CLI’s suite resolution. - LSP: unsaved edits are honored for paths with special characters. The
open-buffer overlay is keyed by source name instead of the raw file URI, so a
path segment containing sub-delimiters (
(,+,', …) no longer misses.
Documentation
- Documented
proef-lspand thelsp/macros/diff/reportsubcommands across the README, TECH-SPEC CLI/dependency references, and the RELEASING publish order.
[0.5.0] - 2026-08-04 (LSP language server)
Added
proef lsplanguage server (ADR-0017). A server-only, generic-LSP stdio binary — a second front-end over the sans-IO core — giving feature/pack authors live editor support: diagnostics (the whole--dry-runvalidation set, republished across the suite as you type), go-to-definition (Gherkin step → the macro that binds it), completion (macro-pattern step completions, prefix-ranked by relevance to the typed prose), and find-references (every step a macro binds). Wired into Neovim/Helix/Emacs via generic LSP config — seedocs/EDITORS.md. No VS Code extension in v1.proef.tomlconfig is a startup snapshot (restart the server after editing it). Works on Linux, macOS, and Windows. Pinnedlsp-server 0.7.9/lsp-types 0.97.0.- New
proef-corepublic surface enabling the language server: the injectableSourceProviderseam (proef_core::provider), the collect-allanalyze_suiteanalysis (proef_core::analyze) — the same headless analysis the CLI runs, driven over an overlay-then-disk provider so the LSP re-validates the whole suite on every edit — andmatcher::prefix_rankfor prose-prefix completion ranking. All keep the core sans-IO (the IO is injected).
[0.4.0] - 2026-08-03 (external config & environments; competitive-review breadth)
Added
-
Suite setup & teardown (
proef.toml [run] setup/teardown, ADR-0014). Each names a feature run once around the whole suite (the Playwright/JestglobalSetupmodel).setupruns before the parallel pool and merges itssaveAs: globalpromotions into the shared store before any scenario lowers, so it seeds fixtures/shared state every scenario reads via${global:…};teardownruns once after for cleanup. A setup failure aborts the run as a user/system fault (never a test failure, exit 1); teardown runs only if setup succeeded and its failure is a distinct exit 3 (never a silently green suite). Both are excluded from the pool, so a setup/teardown feature inside the suite never also runs as an ordinary scenario. -
proef test --output tap— a TAP version 13 stream to stdout, one test point per scenario, derived from the run’s own outcomes (not from hurl), forprove/tappyand TAP-native CI. The human report moves to stderr (as with--output json).@quarantinescenarios map to the# TODOdirective (their failure does not gate); skipped scenarios to# SKIP; failure detail rides in a redacted YAML block.--output tapis rejected onflows/macros(a user error, not a silent human fall-back). -
proef macrosnow flags near-duplicate pattern macros — two that differ only in their{capture}names (identical literal skeleton), which are confusable to authors. Advisory only (never gates the exit code);--output jsongains anearDuplicateOffield besideunusedfor a CI hygiene check. The heuristic is deliberately tight (skeleton equality), so a legitimately similar family with distinct literals is left alone. -
Localized Gherkin (
# language:) is now verified and test-covered — a localized feature parses, its dialect keywords are stripped, and a localized scenario outline withExamplesexpands like any other. Outline detection now keys primarily onExamplespresence (dialect-independent) with the English keyword as a fallback, so this no longer relies on an English-only heuristic. (A localized outline that omits itsExamplesstill degrades to an unbound-step error, since gherkin 0.16 does not expose its dialect keywords.) -
Built-in
expect:shape-macro library. The embeddedCorepack gains a curated, product-neutral set of response-shape assertions —the value at {path} is a string/… a number/… a boolean/… a uuid/… an ISO date/… present/… a non-empty list— each merging one hurl type predicate (isString/isUuid/isList+count, …) into the previous request. It is a convenience layer over the existingexpect:mechanism (no new engine capability, no marker DSL); the raw-hurl assert vocabulary still covers anything the macros don’t. -
Run-level SLA gate (
proef.toml [sla]). An opt-in latency budget: after a run, per-step wall-clock durations fold intop95-ms(95th-percentile ceiling) andmax-ms(slowest-step ceiling); a breach prints the offending metrics + the slowest steps and maps to exit 1 (a test failure). It is off by default (no[sla]table = no gate, run byte-identical to before), env-overridable via[env.<name>.sla], introduces no new exit code, and never downgrades aUser/Systemfault. Distinct from hurl’s per-requestduration <assert — the gate is an aggregate budget over the whole run. Skipped steps are excluded from the population. -
External config & environments (
proef.toml, ADR-0012). New[url]and[vars]tables hold non-secret suite variables, referenced in packs as${url:<key>}/${vars:<key>};[env.<name>.<section>]profiles deep-merge per-environment overrides over the base tables (url/vars/http/run).proef test --env <name>(orPROEF_ENV) selects the active environment.proef.tomlis discovered by searching up from the working directory (like cargo/git), so it is found from any subdirectory. Adds theproef::resolve::missing_config_vardiagnostic. -
Default suite path.
[run] suitesets the pathproef test/flows/artifactsuse when given none (falling back to thetests/convention), soproef testruns with no argument. An explicit path still wins. -
Documentation set completing the corpus:
docs/DIAGNOSTICS.md(all 57 diagnostic codes, corpus coverage marked),docs/CONFIG.md(proef.tomlreference),docs/EVENTS.md(theevents.jsonlwire schema for CI),docs/TROUBLESHOOTING.md(exit codes, glyph legend, frequent failures),docs/CONTRIBUTING.mdanddocs/SECURITY.md(threat model, private vulnerability reporting), and an IDE-integration section in AUTHORING. -
proef test --scenario-file <file>: scope a--scenarioname filter to one feature file (duplicate scenario names across files stay disjoint; the libtest-mimic harness uses it to keep the Trial↔scenario bijection). -
scenario_finishedevents now carry afilefield — the run-wide scenario identity alongsidescenario(additive, ADR-0008; absent in older records). -
Diagnostics
pack::pattern_duplicate_capture(a{capture}written twice) andlower::kind_unrouted(internal registry-drift safety net). -
proef macroslists every loaded macro with its call count and flags user-pack pattern macros that no scenario binds (dead prose bindings);use:-only helpers and unused builtins are listed but never flagged.--output jsonfor CI dead-code gates. -
proef test --run-id <id>pins the injected run id (likeartifacts --run-id), so a run’s${fake:…}data — which keys on the run id — is reproducible; the JSON summary echoes the id. -
proef test --dry-run --sarif <path>serializes validation diagnostics (unbound steps, pack lint, non-finite retries) to a SARIF 2.1.0 log — a shift-left gate that renders findings as inline PR annotations. The export is additive: the dry-run’s exit code is unchanged. -
proef test --rerunre-runs only the scenarios that failed in the last run (read from its JSONL record, keyed on the run-wide(file, name)identity); it composes with--tags/--scenario, and reports “nothing to rerun” (exit 0) when the prior run was clean. -
@quarantinetag: a scenario so tagged runs and reports normally, but its test-failure no longer gates the exit code (aSystem/Userfault still does — quarantine is for flaky tests, not broken input or infra). A note prints when a quarantined scenario fails, so it is never silently swallowed. -
proef diff [base] [new]compares two run records (defaulting to the previous and latest runs) and reports scenario status transitions — regressed, fixed, still-failing, new, removed — keyed on the run-wide(file, scenario)identity, plus per-step flakiness (rising retry counts) and perf deltas (steps diffed ontext, never the volatile authored line). It is a derived view overevents.jsonl, never a second record (ADR-0008);--fail-on-regressionexits 1 when a scenario regressed, for CI gating. -
Flaky-failure detail: a step that passes only after a retry now records the messages from its earlier, failed attempts as
attempt_detailson thestep_finishedevent (additive, ADR-0008); JUnit surfaces them as<flakyFailure>under the passing test case, so a green-on-retry run is honest instead of indistinguishable from a clean pass. The engine already collected the earlier-attempt errors — they were being discarded on success. -
proef report [run-id]writes a self-contained HTML report for a run — scenario tree with pass/fail pills, per-step attempts and timing, a per-scenario timing waterfall (each step’s bar offset by the steps before it and as wide as its own duration — the sequential cascade within a scenario, derived purely from step durations), a cross-worker timeline (a lane per worker, each scenario a bar on a shared run-relative axis, so concurrency is visible at a glance), failure detail, and deep-links to the executed.hurlartifacts (bodies are not inlined). -
Injected run timing (ADR-0015).
scenario_started/scenario_finishedevents gain optionaltimestamp_ms(run-relative) andworker(0-based index) fields, stamped at the CLI sink on the worker thread so the sans-IO core stays clock-free. Additive (absent on records without timing); they power the HTML timeline. Records without them degrade to the waterfalls alone. A pureproef_core::html::render_htmlderives it from the event stream (ADR-0008, snapshot-locked); the events are already redacted at the sink, so the page is too. Defaults toreport.htmlinside the run dir;-oredirects it.
Changed
--tagsis now a boolean expression, not a comma-separated list. It takes a single expression overand/or/notand parentheses (the@stays optional), e.g.--tags "@api and not @slow"; a bare tag still works. The grammar and evaluator live in the sans-IO core (proef_core::tags, deterministic and fuzzed); a malformed expression is a user error (exit 2), as is a selection that matches nothing. This replaces the old CSV OR-list — there is one selection mechanism, not two.--outputis a typed value: an unknown format (e.g. ajsonltypo) is a user error (exit 2) instead of silently degrading to the human report.--watchreruns only on.feature/.yaml/.ymlchanges — the watched tree can now contain proef’s own run output without a self-trigger loop.- The example corpus (
tests/features/) and the dev fixture use a neutral workspace / activity-board domain (record · note · event · attachment · session · channel) — no product-specific vocabulary. CHANGELOG.md,CONTRIBUTING.md, andSECURITY.mdmoved underdocs/(root keeps onlyREADME.mdandCLAUDE.md).- Pack root key renamed
templates:→macros:(ADR-0004 amendment): one canonical spelling for the prose→engine binding layer (the entry is a macro, the file a pack). Notemplates:alias — packs using the old key fail to load. - The dev-loop fixture (
cargo run -p xtask -- fixture) binds the advertised default port 8787 — falling back to an ephemeral port (and printing aPROEF_BASE_URLline) only if 8787 is busy;... -- fixture <port>overrides. Soproef.toml’s defaultbasereaches it with noPROEF_BASE_URLexport (ADR-0011 amendment). ItsGET /healthnow returns a versioned identity —name, a numericversion(1.0), and the RFC 3339timeit answered. - The unbound-step diagnostic (
bind::unbound_step) now prints a paste-ready pack-macro stub — quoted tokens in the sentence become{argN}captures — alongside the existing did-you-mean suggestion, so an author can add the missing macro without hand-writing thematch:/hurl:scaffold. - CI reporting surfaces failures and flakiness more honestly. Under GitHub
Actions the run emits a
::error file=,line=,title=annotation per failure (rendered in the PR “Files changed” gutter; gated off when--output jsonowns stdout). The job summary gains a flaky passes section and per-failure attempt counts, and the JUnit report records “passed on attempt N” for a scenario that only went green after retries — a silent green-on-attempt-2 is no longer invisible. docs/AUTHORING.mdgains an “Asserting responses” cookbook surfacing the hurl 8.0 predicate/filter/RFC-9535-JSONPath vocabulary that rawhurl:blocks already accept — documenting existing capability, not new engine work.- A failed step now prints a
curl:reproduce line — the redactedcurlfor the failing request, surfaced from the embedded engine via a new engine-agnosticStepOutcome.reproduce_hint— so a failure can be replayed request-by-request without leaving the terminal. Secrets are masked.
Removed
- The
# key: valuefeature-file directive mechanism (e.g.# baseURL:, ADR-0012 amendment). Variables now have exactly one home —proef.toml([url]/[vars]) — so a.featurefile can no longer define a variable (one-way-to-do-one-thing).#comment lines stay valid gherkin comments; they are simply no longer parsed. The env-override the directive provided is preserved by embedding${env:NAME:-default}in a config value (resolved recursively).${…}plain-name resolution is nowargs > defaultsonly.
Fixed
- Optional-batch error path no longer double-reports later batches into the
JSONL run record (ADR-0008);
saveAs: globalpromotions are no longer dropped when the store lock is poisoned; the event sink recovers from a poisoned lock instead of truncating the record. expect:merge scopes to the last entry (fence-aware);[Options]injection can no longer duplicate a section; theuse:graph walk is node-linear instead of exponential on multi-edge chains.- The embedded-hurl version lockstep is now asserted by a test; the encrypted
secret store maps user vs. environment faults to exit 2 vs. 3 (ADR-0009);
run.log / artifact-write / malformed-
proef.tomlfailures surface instead of being swallowed.
[0.3.1] - 2026-07-29 (secret-management hardening)
Added
proef secret rm NAMEremoves a stored secret (locked atomic rewrite; removing an absent name exits 2).PROEF_KEYenv override supplies the project key directly (base64) — a committed ciphertext store now decrypts in CI without shipping the key file; a set-but-invalid key errors instead of silently falling through.proef doctorreports secret store/key health (readable, parseable, private permissions); a corrupt.proef-secrets.jsonno longer brickssecret set— it is moved aside to.corruptand a fresh store begins.
Fixed
- Secret-valued captures never reach
.proef-state.json: asaveAs: globalcapture whose value equals a known secret is refused — the owning step warns with the reason — closing the one sink the redaction invariant (ADR-0005) did not cover. - Secret resolution reads the store and key once per run instead of once
per secret (no torn view against a concurrent
secret set). - Warned steps now print their reason on the console (
↳ …) — a bare ⚠ glyph explained nothing, foroptional:failures too.
[0.3.0] - 2026-07-29 (data-safety blockers, Then visibility, taxonomy)
Fixed (v0.2.1 review — every finding reproduced before fixing)
- Asset copy destroyed user files:
proef artifacts -opointing at the suite truncated referenced assets to 0 bytes, and..references escaped the output directory. Copies now refuse absolute/..references (exit 2), never copy a file onto itself, and surface IO errors (exit 3). - Run rotation deleted arbitrary directories: with
runs-dirshared with user content, rotation could recursively delete user directories — and its own in-flight run. Only uuid-named run records rotate now, never the live run, and rotation happens before the new run dir exists. - Zero-entry payloads passed silently: a comment-only
hurl:block ran nothing while the scenario reported green. Load-time lint rejects it; the engine backstop emits Skipped outcomes for anything that slips through. proef flows … | head(and every other command) tolerates a closed pipe; a non-UTF-8 environment variable no longer aborts any command.- Raw
[Options] retry:/repeat:values are parsed and capped (10000), anddelay:is capped at 1 hour in both typed and raw forms;repeat:now counts toward the batch budget so long repeats aren’t blamed on the environment. - Concurrent
proef secret setcalls no longer lose keys (advisory-locked, atomic 0600 temp+rename store; the key-creation race resolves to the winner’s key).proef fmtandschema --add-towrite atomically. proef fmtkeeps fenced body bytes verbatim (blank lines and trailing whitespace inside ``` fences are the bytes the test sends).- Nested suites now load their packs: pack discovery recurses like feature
discovery (
packs/directories at any depth);proef fmtshares the rule. - Duplicate/empty Examples header columns are a named error instead of a
silent last-value-wins; an empty
.featuregets a plain-language error; a UTF-8 BOM is stripped instead of shifting every diagnostic span.
Changed
- Then steps are visible everywhere:
expect:macros now surface as their own step rows in console, events, JUnit, andexplain, with assert failures attributed to the authoredThenline — the host request no longer inherits its followers’ assert failures. Artifact bytes are unchanged; sidecars gain one row perThen(schema-compatible). - Error taxonomy: mistakes in the test’s own text (undefined
{{var}}, bad JSONPath/regex/URL/options, unreadable body file) exit 2 instead of 3, anchored on hurl’s own assert-context flag. when:guards skip on a literalfalse/0as well as empty — an author writingwhen: ${flag}withflag=falsemeans skip.proef.tomlis no longer gitignored (it is documented, committed project config).- proef-core public API: removed dead surface (
NormalizeReporter, the never-populatedconfigresolution tier,StepOutcome.artifact_span,LoweredStep.retry,StepKeyword, and friends); addedEngineErrorClass::UserInput,StepPayload::MergedAsserts,ScenarioOutcome.artifact_slug,Guard::skips.
[0.2.1] - 2026-07-29 (review P0 + failure UX)
Fixed
[Options]header detection follows hurl’s token grammar — the injection can never land inside XML/JSON/prose bodies (class closed, unit-tested).proef artifactssurvives a closed pipe (exit 0, best-effort writes).--dry-runhonors--scenario/--tagswith the same zero-match exit 2.- Duplicate scenario names dedup feature-wide (
#N): unique artifacts, console buffers, and events — no silent overwrite. .proef-secrets.jsonis created0600, gitignored, and documented.
Changed
- Failure details surface hurl’s computed expected/actual (
fixme) anchored on the error’s own artifact line, not the entry’s first line. - GETTING-STARTED uses
PROEF_BASE_URL, points the reader at a runnable target, and frames sample output honestly.
[0.2.0] - 2026-07-29 (correctness, output contract, author docs)
Fixed (v0.1.0 deep-review follow-up — all three blockers reproduced first)
[Options]injection is body-fence-aware: aretry:/delay:step whose body contains method-looking lines no longer gets options spliced into the body it sends.- Step↔entry correlation is a partition anchored on each entry’s request line: a comment-only step can no longer cause the next request to be sent twice (one authored POST is one POST, asserted via the event stream).
delay:joins the watchdog budget (with saturating duration math throughout), so delayed steps are no longer killed as system errors;retry.countis capped at 10000 by the pack lint.- A panicking scenario thread is contained (
catch_unwind), reported as a System fault under its real identity immediately — never a budget timeout; abandoned scenarios keep their real file/name/line; steps in batches never reached reportSkippedinstead of vanishing from every report.
Changed
- Output contract:
--output jsonowns stdout exclusively (human report on stderr — pipeable intojq);StepFinishedevents carry adetailfailure field (additive);optional:failures reportWarnedeverywhere consistently; engine failure details use hurl’s own error descriptions instead of RustDebug; diagnostics drop ANSI when stderr is not a terminal; a filter selection matching nothing exits 2; failure output prints a ready-to-runreproduce: hurl …line; the artifact replay header names required--secretplaceholders; the undocumented.envautoload was removed.
Added
- Author-facing documentation:
docs/GETTING-STARTED.md(first suite in ten minutes) anddocs/AUTHORING.md(the full pack/feature reference). - Mechanical alignment gates:
xtask docs-check(crates and ADRs must appear in their indexes) runs in PR CI;xtask public-apisnapshotsproef-core’s public API surface (1.4k items) and fails CI on unreviewed changes — the mechanical form of the zero-core-diff invariant.
[0.1.0] - 2026-07-29
Initial release.
Added
- Authoring: Gherkin
.featurefiles in plain business prose; YAML macro packs bind prose to executable steps viamatch:patterns, typed params, defaults,use:composition (cycle-checked),expect:assert-only macros,optional:, finiteretry:,delay:,when:guards, andsaveAs: globalpromotions. - Validation:
proef test --dry-runbinds, lowers, emits, and parse-validates every scenario without touching the network; stable diagnostic codes with source-span rendering; a seeded error corpus pins every code; pack payloads are validated at load by the engine that claims them. - Execution: the hurl engine runs artifacts in-process (exact-pinned
hurl 8.0.1); contiguous same-engine steps batch maximally; variables and cookies chain across batch splits; per-entry[Options]override batch defaults; finite budgets with a watchdog bound every scenario; Ctrl-C cancels gracefully (twice = hard exit); parallel scenarios share a typed World with write-set-only merge-back and a persistent global store. - Artifacts as the contract: every scenario emits canonical
.hurltext that is byte-identical to what the engine executes, plus a sidecar map (entry ↔ feature anchors, explicit batch/step indices),.vars, and any referenced file assets — replayable with stockhurl --test. - Record & reporting: a versioned JSONL event stream is the run record
(live per-entry progress included); console BDD tree; JUnit XML; GitHub job
summaries;
proef explainreplays the record; secrets are encrypted at rest, injected via hurl’s redaction, and value-redacted once at the event sink — never present in artifacts, events, logs, or reports. - Tooling:
proef flows,artifacts,schema(merged JSON Schema with editor modelines),secret set|list,fmt(canonical hurl blocks),doctor,--watch; a libtest-mimic harness exposes one test per scenario to nextest/IDEs;${fake:*}deterministic synthetic data seeded from the run id. - Quality gates: unit + property tests, fuzz targets, insta snapshot corpus (artifacts, diagnostics, events), fixture-server integration suite, assert_cmd CLI/exit-code suite (0/1/2/3 contract), cargo deny/machete/ zizmor in CI, cargo audit nightly, a scheduled canary against the next hurl release, and CI on Linux, macOS, and Windows.
- Distribution: tagged releases build five targets (macOS arm64/x86_64,
Linux arm64/x86_64-gnu, Windows x86_64-msvc) with
cargo auditable, ship a Homebrew tap formula and acargo binstall-compatible layout, and attest SLSA provenance once the repository is public.