Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

proef — Improvement Plan

Status: complete. Round 1 (§1–§11) and Round 2 (§12) are both shipped — every batch (N2·N6a·N9·N8·N7·N4·N6b·N3), with N1 deferred per ADR-0013, N5 rejected, and the OpenAPI generator settled in ADR-0016. This file is the historical feature roadmap and its file:line citations are from 2026-07/08; the live worklist is OPEN-FINDINGS and the running ledger is CLAUDE.md’s Status block. · Date: 2026-07-31, appended 2026-08-02, closed 2026-08-31 · Owner: Emre Companion docs: PRD (scope + the binding non-goals, §3), adr/ (the invariants every item must respect), TECH-SPEC (types/pipeline), IMPLEMENTATION-PLAN (milestones + definition of done).

0. What this is (and isn’t)

A competitive analysis of proef against the BDD and API-testing field, converted into a roadmap of candidate improvements — each validated against the current architecture at file:line, and filtered strictly through proef’s permanent non-goals (PRD §3).

Nothing here is committed work; it is the durable output of the post-M5 review round. Effort grades (S/M/L) and sequencing are advisory. Feature numbers (#1…#15) are stable identifiers used across the tables. File:line citations are as of 2026-07-31 and will drift — treat them as “start reading here”, not addresses.

Every item is in scope by construction: none proposes a second engine, mocking, contract testing, load testing, a dashboard/server, OpenTelemetry, or importing hand-written hurl — all permanent non-goals (§3 below). The work is almost entirely surfacing data the sans-IO core already computes, not new engine capability.

1. Headline finding

proef’s engine is already ahead of its cohort; the gaps are in reporting surfaces and authoring/maintenance DX, not in HTTP power. Because proef pins hurl 8.0.1, it already ships hurl 8.0’s full assertion arsenal (RFC 9535 JSONPath with filter functions, type predicates isUuid/isIsoDate/isString/isObject/isList, response-time duration < ms, the filter chain split/count/toDate/base64Decode/daysAfterNow). Pack authors can write Karate-grade assertions today, inside raw hurl: blocks — they are just undocumented and under-surfaced. The roadmap is therefore mostly exposure and tooling, which is cheap, rather than engine work, which is done.

2. Where proef already wins (positioning to defend)

proef strengthCompetitor weakness it beats
One canonical way, raw-hurl-only, no escape-hatch languageKarate’s most-cited flaw: a three-language model (Gherkin + DSL + embedded JS) the moment anything gets non-trivial
Deterministic sans-IO coreKarate/Tavern/Newman are non-deterministic; reproducibility is now a headline selling point
Artifacts = executed bytes (hash-locked, git-diffable, replayable hurl --test)Postman’s opaque JSON collections (unreviewable merges) — the reason Bruno is displacing it
Finite retries + budgets + watchdog (ADR-0007)hurl itself has no cancellation and unbounded retries — proef fixes its own engine’s biggest gap
Property-tested secret masking + typed exit codes (ADR-0009)Most tools treat masking loosely and lack a stable exit-code contract
Dev-maintained macro packs = an enforced “step dictionary”The AI-authoring trend is groping toward exactly this; proef has it structurally
No cloud, no account, plain textPostman’s 2026 pricing exodus is driving the whole git-native wave

3. Scope guardrails — what this plan will NOT propose

Permanent non-goals (PRD §3) — never revisit under “competitor parity”: further engines (browser/gRPC/etc.), API mocking, contract testing (OpenAPI drift / Pact / Schemathesis), load testing, a desktop dashboard or server mode, OpenTelemetry export, dynamic plugin loading, importing hand-written hurl into Gherkin (artifacts flow outward only), static musl/Windows binaries.

Named anti-patterns (from the 2025–2026 trend research) to avoid: silent retries / green-on-attempt-2; retry ceilings > 3 or unbounded (proef’s finite cap is already correct); permanent quarantine without owner+expiry; treating masking as a security boundary; LLM self-healing / non-deterministic test mutation; config sprawl.

A Karate-style marker DSL (#uuid, ##optional) is explicitly rejected — it would be a second assertion mechanism competing with raw hurl predicates. Achieve the same readability by surfacing hurl’s native predicates (item #2), not by inventing a layer.

The overriding design gate is one-canonical-way (see §6): four items must replace or augment an existing mechanism, never add a parallel knob.

4. What code validation changed about the roadmap

Three deep code-validation passes (reporting, CLI/tags, authoring) reshaped the outside-in list. Six meta-findings:

  1. “Surface, don’t build” is confirmed at file:line. The sans-IO core already computes and records the data behind most items: attempts/duration_ms/detail on StepFinished (proef-core/src/event.rs), the winning macro_name per bound step (bind.rs:20), deterministic seeded fakes, and hurl’s own curl_cmd (already returned by EntryResult, currently discarded in engine-hurl session.rs).

  2. Two items are already ~90% built — validation downgraded them.

    • #14 seeded fakes: fakes are already deterministic (hand-rolled SplitMix64, no rand crate, fake.rs) and already seeded by the injected run_id, which is already recorded in RunStarted. The feature collapses to “let test pin run_id the way artifacts --run-id already can.”
    • #9 stub-gen: the “did you mean” fuzzy suggestion already exists (bind.rs:214, levenshtein); only the paste-ready stub template is missing.
  3. One item’s premise is broken — #13 impacted-only re-run. There is no stored content hash anywhere; ADR-0010 is enforced as a byte-identity assert_eq! on outputs (crates/proef-cli/tests/execute.rs:188), not a reusable digest. The only honest impact fingerprint — the emitted .hurl — is deliberately not stable run-to-run (run_id lives in runtime globals). #13 is gated behind a determinism prerequisite.

  4. Two plumbing gaps gate a cluster. Scenario tags stop at the CLI edge — they never reach the runner (ScenarioSpec/ScenarioOutcome in runner.rs carry no tags) or the event stream — so #15 (quarantine) and part of #8 (rerun) need new plumbing. And #4 (boolean tags) is hard-blocked by value_delimiter=',' on --tags (main.rs:57); it is a contract-changing replace, not an add.

  5. The recurring architectural rule is one-canonical-way. Four items (#4, #9, #14, #15) each risk a second mechanism; each must replace or augment an existing one.

  6. The exit-code contract (ADR-0009) is a live wire for #15. Quarantine changes which scenarios feed RunSummary::exit_code (fine — the fold stays pure in core) but must extend the pinned assert_cmd tests and never let a quarantined system fault mask exit 3.

5. The validated roadmap (master table)

Status is what exists in the tree, re-verified against main on 2026-09-08 — --help for flags, source for the rest. Verdict is the 2026-07-31 judgement of whether the idea fits the architecture; it never meant “done”, and reading it that way is why this table looked like a backlog when, by 2026-08-10, 13 of its 16 items had shipped (14 today; one partial, one gated).

Status: shipped · partial · open · gated (premise rejected). Verdict legend: ✅ FITS · ⚠️ NEEDS-ADAPTATION · 🚫 premise broken. Effort: S ≤ ~1 day · M ~days · L ~weeks.

#ItemStatusVerdictLives inEffortArchitectural truth (as of 2026-07-31)
2Assertion cookbook (surface hurl-8.0 predicates/filters)shipped✅ docs-onlydocs/SLowering copies all but ${…} verbatim (resolve.rs:163); load runs the real hurl_core parser (engine-hurl/src/lib.rs:95). Predicates already work. Caveat: grammar validation, not JSONPath semantics.
1GitHub ::error file=,line=,title= annotationsshipped✅proef-cli ci_reports.rsSfile+line+detail already flow to write_github_summary (ci_reports.rs:98). Gate stdout vs --output json; percent-encode multiline detail. Line-only (no byte-span at runtime).
5--curl export per requestpartial✅engine-hurl session.rsShurl’s EntryResult.curl_cmd is already returned, just discarded (session.rs:346). Must redact (holds resolved secrets, ADR-0005). Fold into the existing reproduce: block (exec.rs:230). Partial as of 2026-08-10: the curl line is surfaced on a failing step (exec.rs), which covers the debugging case; a per-request export flag for passing steps is still open.
3a“passed on attempt N” badge (JUnit/summary)shipped✅proef-cli ci_reports.rsSattempts:u32 already on StepFinished (event.rs:72) + StepOutcome; JUnit ignores it today (ci_reports.rs:43).
9Stub-gen for unbound stepsshipped⚠️proef-core bind.rsSAugment the existing did-you-mean help (bind.rs:89), zero-match arm only — not a new command. Derive {param} from quoted tokens (matcher already sheds quotes, matcher.rs:85).
10SARIF export of --dry-run diagnosticsshipped✅proef-cli new sarif.rsS–MDiag (diag.rs:56) → SARIF result ~1:1: code→ruleId, byte span→region.byteOffset. Pre-populate rules[] from the closed diagnostic-code set. A parallel serializer to render.rs.
14--seed (reproducible fakes)shipped (as run-id)⚠️proef-cli main.rs/exec.rsSThread into front::run’s existing run_id param (artifacts already exposes --run-id, main.rs:101). Caveats: arbitrary seed breaks JUnit’s UUID parse (ci_reports.rs:22); occurrence is per-scenario, not per-run — identical ${fake:X} at the same position in two different scenarios still draws the same value (a known limitation; see OPEN-FINDINGS). Note: the per-step reset this row originally cited (Refs::default() on every lower() call) was fixed in 0.6.0 — the counter now threads through lower with a high-water mark. Only the cross-scenario half remains. Shipped as the one knob §7 demanded (no separate --seed): test --run-id pins the fakes, and --shuffle seeds its permutation from the same id.
7Dead-macro / usage reportshipped✅proef-cli new macros --usageS–MBoundStep.macro_name (bind.rs:20) vs packs.macros. Count use:-only macros (pattern:None, pack/mod.rs:165) as reachable via the use: graph. Report the whole corpus, not a --tags subset.
6Self-contained HTML reportshipped✅core render_html(&[Event]) + cli writeMPost-hoc proef report <run-id> replaying events.jsonl like explain (explain.rs:12) — the command is the HTML report, so it took no --html flag; -o picks the file. Bodies live in artifacts/ — deep-link, don’t inline. Derived view, never a second record (ADR-0008).
12proef diff between two runsshipped⚠️proef-cli diff.rsMIdentity (file,scenario) (why ADR-0008 added file, event.rs:86); key step diffs on text not line (lines shift on edit). attempts+duration_ms → free flakiness/perf-regression detector. Pre-file records replay file="".
4Boolean tag expressions (@a and not @b)shipped⚠️ (replace)grammar in core, apply in cli front.rsMvalue_delimiter=',' (main.rs:57) actively breaks and/or/not; must replace the CSV/OR contract (front.rs:388), keep empty-match=exit-2 (front.rs:382). Grammar/evaluator is deterministic → proptest/fuzz-shaped, belongs in core.
8Rerun-only-failures (--rerun)shipped⚠️proef-cli, reuse explain replayMexplain already reads the latest record + failed (file,name) (explain.rs:100, event.rs:83). Needs a multi-identity predicate (today’s scenario/scenario-file filters are single-valued, exec.rs:301,321). Factor a shared record::failed_scenarios.
15@quarantine non-gating tagshipped⚠️ (contract)thread gating:bool core+cliMTags must first reach ScenarioSpec/ScenarioOutcome (P1). exit_code() stays pure in core (runner.rs:89) and skips non-gating outcomes. Events still emit the scenario → not hidden. Extend the pinned assert_cmd tests; never mask a Fault::System (exit 3).
3bTrue <flakyFailure> with earlier-attempt detailshipped⚠️schema + engine-hurlMNeeds an additive attempt_details field (ADR-0008 additive-only) + engine-hurl collecting per-retry bodies before the final one. Bigger than 3a.
11proef lsp (feature/pack language server)shipped✅new proef-lsp crateLAll diagnostic substrate is headless/sans-IO already (bind, pack::load, resolve Probe mode, matcher). New: a sync lsp-server (tokio ban forbids async), a byte-offset→token API (not exposed), and a partial-results wrapper (bind/load are all-or-nothing today). Karate notably lacks good IDE support → differentiator.
13Impacted-only re-run (content-hash)gated🚫— (gated)LNo input hash exists; raw-input hashing is unsound (shared packs, use: nesting, config vars fan out). Honest fingerprint = per-scenario emitted .hurl, but it is not run-to-run stable (run_id in globals). Needs a determinism prerequisite first; silent-green risk. 2026-09-07: a suite-level input hash now exists (inputs.json, for flaky windows) — not per-scenario; still gated.

6. Prerequisites that unlock clusters

  • P1 — carry scenario tags + a gating flag past the CLI edge into ScenarioSpec / ScenarioOutcome (runner.rs:30,121) and, additively, the event stream. Done (RF waves): tags ride ScenarioSpec/ScenarioOutcome and scenario_finished.tags; gating became the reserved-tag instruction plus the non-gating list (ADR-0019) rather than a bool. #15 and #8 both shipped on top of it.
  • P2 — a shared record::failed_scenarios(run_id) + a multi-identity scenario predicate. Reused by explain, --rerun (#8), and proef diff (#12).
  • P3 — a deterministic emitted-.hurl fingerprint (stable run-to-run despite run_id). Prerequisite for #13; do not attempt #13 without it. 2026-09-07: proef_core::fingerprint now hashes a run’s whole input set (feature sources, loaded macros and fragments, the resolved [url]/[vars] scope) into inputs.json — proef flaky’s equivalence class. Suite-level and deliberately coarse (any edit ends a window), so it is not the per-scenario, run_id-independent fingerprint #13 needs; P3 stands.

7. The one-canonical-way watch-list

Each of these must fold into an existing mechanism, never ship beside it:

ItemMust replace / augment (not duplicate)
#4 boolean tagsReplace the CSV/OR --tags semantics — no second tag syntax
#9 stub-genAugment the existing did-you-mean Diag.help — no separate proef stub command
#14 --seedFold into the existing run_id determinism knob — no parallel seed unless it replaces run_id-keyed fakes
#15 quarantineExactly one non-gating tag name; must not spawn a second “skip” concept
#5 --curlAttach to the single reproduce: mechanism, not a parallel debug path
  1. Batch A — free / small, all FITS, all reuse existing data. #2 cookbook (docs) → #1 annotations → #5 --curl → #3a attempt badge → #9 stub.
  2. Batch B — small, high-leverage. #10 SARIF · #7 dead-macro · #14 --seed.
  3. Prereqs → Batch C — medium, now unblocked. Build P1+P2, then #6 HTML · #12 diff · #8 rerun · #15 quarantine · #4 boolean tags · #3b flaky-detail.
  4. Batch D — strategic. #11 LSP. And #13 only after committing to P3.

9. Per-item detail & competitor provenance

Each entry: what it borrows from whom → the validated architectural note. Numbers cross- reference §5.

#2 Assertion cookbook — from Karate’s fuzzy markers + hurl’s own docs. Confirmed a docs task: resolve() leaves everything but ${…} byte-for-byte (resolve.rs:163-208, test runtime_tier_passes_through), and pack load validates the full grammar via hurl_core::parser::parse_hurl_file (engine-hurl/src/lib.rs:95), re-checked on the emitted artifact (front.rs:159). Extension: ship a tests/features/ reference feature exercising each predicate so the cookbook is snapshot-locked against hurl upgrades (the canary catches drift).

#1 GitHub annotations — from the 2025–2026 CI-reporting shift (annotations displace log-diving). The failures loop already prints `{file}:{line}` — {detail} (ci_reports.rs:98). Emit ::error workflow commands as a sibling; title = scenario + step text. Risk: stdout is owned by --output json (exec.rs:114) — gate it.

#5 --curl export — from hurl’s loved --curl; Bruno/Postman “copy as curl”. hurl hands us EntryResult.curl_cmd already (session iterates result.entries at session.rs:346 but reads only captures/errors/duration). Must pass Redactions (session.rs:396) before any sink — the curl line contains resolved secrets. Cannot be derived pre-execution (needs runtime {{…}}).

#3a “passed on attempt N” — from the flaky-test-honesty consensus (never hide a retry). attempts is first-class (event.rs:72, step.rs:149) and already printed on the console (report.rs:220); JUnit simply drops it. Count-based badge is S.

#9 Stub-gen — from Cucumber/Behave snippet generation. The matcher already computes the nearest macro via closest_pattern/levenshtein (bind.rs:214, matcher.rs:248); add a match:+hurl: | skeleton to the help text for the zero-match arm only (an ambiguous step, bind.rs:99, must not get a stub).

#10 SARIF — from SARIF’s rise for static/validation findings inline in PRs. Dry-run diags are a structured Vec<Diag> before miette (diag.rs:123, front.rs:64). Diag maps ~1:1 to a SARIF result; the closed code set (one per tests/errors/ dir) pre-fills rules[]. cli-edge serializer, no core change.

#14 --seed — from seeded-faker reproducibility. Fakes already deterministic (fake.rs:12 SplitMix64/FNV, seeded fnv1a(run_id) ^ …), seed already recorded in RunStarted{run_id} (event.rs:26). Design fork: alias run_id (zero core change, but must stay uuid-parseable for JUnit) vs a dedicated recorded seed field (cleaner, but a second knob — resolve per one-canonical-way). Known limit: the occurrence counter is scoped per scenario (lower.rs:69-86, Refs::fakes, threaded through resolve()’s fakes: &mut usize parameter, not reset per call) → cross-step uniqueness within a scenario now holds, but cross-scenario uniqueness still does not: two different scenarios each resolving ${fake:X} at the same position in their own step order get the same value. Document before advertising “unique fakes” — it means per-scenario, not per-run or per-entity.

#7 Dead-macro report — from Cucumber’s usage formatter marking UNUSED. Binding records BoundStep.macro_name (bind.rs:192); iterate front.features[].scenarios[].bound.steps[].macro_name vs packs.macros.keys() (pack/mod.rs:116). use:-only macros (pattern:None) need reachability via the use: graph (pack/validate.rs:628) to avoid false “unused”. Report the whole corpus.

#6 HTML report — from Cucumber/Karate/hurl HTML reports; the industry’s convergence on the Cucumber-Messages/JSONL stream. Core render_html(&[Event]) -> String, cli writes; best as post-hoc over events.jsonl (explain.rs already replays it) so historical runs render. Events are pre-redacted at the sink (report.rs:124).

#12 proef diff — from Allure history / test-observability-without-OTel. Identity is (file, scenario) (report.rs:147 ScenarioKey; ADR-0008 added file for exactly this). Key step diffs on text, not the volatile line. attempts+duration_ms make it a flakiness/perf-regression detector. run_id is uuid-v7 → chronology recoverable.

#4 Boolean tags — from Cucumber tag expressions (and/or/not/()). Single filter fn tag_selected (front.rs:388), three callers. value_delimiter=',' (main.rs:57) blocks the operator syntax → drop it, take one expression string, replace the CSV contract. Grammar/evaluator → core (deterministic, fuzz-shaped). Preserve empty-match=exit-2.

#8 Rerun-only-failures — from Cucumber’s rerun formatter (@rerun.txt). explain already discovers + replays the latest record and extracts failed identities (explain.rs:63). Add a multi-identity predicate reusing build_specs (exec.rs:289). Empty failure set → reuse no_scenarios_matched (exit 2), never silent-pass.

#15 Quarantine — from the flaky-quarantine-with-owner+expiry consensus. Thread gating:bool from CLI (which sees scenario.lowered.tags, exec.rs:318) into ScenarioSpec/ScenarioOutcome; exit_code() (runner.rs:89) skips non-gating outcomes and stays pure in core. Events unchanged → scenario still reported. Extend the cli.rs/execute.rs exit-code assertions; a quarantined Fault::System still exits 3.

#3b Flaky-failure detail — from JUnit <flakyFailure> / Allure retries. Needs an additive attempt_details on StepFinished (ADR-0008 additive-only) and engine-hurl collecting per-retry messages (hurl retry is per-entry internal — verify the adapter isn’t already discarding earlier bodies).

#11 proef lsp — from Cucumber’s language server (unbound-step diagnostics, go-to-def, completion); a gap Karate never closed. Reuses feature::parse, bind, pack::load, resolve Probe mode, matcher — all headless, all with stable codes + byte-offset spans that already map to editor ranges. New work: sync lsp-server (tokio banned), byte→token API, and a “collect diags, don’t early-return” wrapper (bind/load are all-or-nothing today).

#13 Impacted-only re-run — from selective/affected-test re-run. Premise broken: no reusable input hash (ADR-0010 is a byte-identity assert_eq! on outputs, execute.rs:188), and the honest fingerprint (emitted .hurl) is not run-to-run stable because runtime globals include run_id (execute.rs:304). Gated on P3; a hash miss must never skip a scenario that would fail (needs --force/first-run fallback).

10. Karate feature ledger — considered / adopted / rejected

Karate (github.com/karatelabs/karate) was the closest competitor and the deepest research stream. This ledger makes the “considered → decision” trail explicit, so each Karate idea is an intentional call rather than an omission.

Adopted — drove a plan item or the positioning:

Karate featureproef outcome
Inline fuzzy markers (match response == { id: '#uuid', age: '#number' })#2 assertion cookbook — the same readability via hurl 8.0’s native predicates (isUuid, isIsoDate, …) surfaced in docs, not a new marker DSL.
Weak/immature IDE support (Karate’s own gap)#11 LSP — reframed as a differentiator proef can win, since Karate never closed it.
HTML report with a timeline view#6 self-contained HTML report (with Cucumber’s and hurl’s).
Three-language cognitive load (Gherkin + DSL + embedded JS)proef’s headline positioning (§2): one-canonical-way, raw-hurl-only, sans-IO — the inverse of Karate’s most-cited flaw.
call / callonce cross-feature reuseAlready covered by proef’s use: / with: macro composition (ADR-0004) — no new work.

Rejected — with the reason (so it stays rejected):

Karate featureWhy not
A marker DSL (#uuid, ##optional, #? _ > 0)A second assertion mechanism competing with raw hurl predicates — violates one-canonical-way (§3). Readability comes from #2 instead.
Embedded JavaScript escape hatchConflicts with the sans-IO deterministic core and one-canonical-way — it is the thing proef exists to avoid.
Soft assertions (configure continueOnStepFailure)Conflicts with proef’s deliberate stop-at-first-failed-step model (the ∅ cascade); a failed step’s downstream is intentionally not run.
Service mocking · karate-gatling perf · UI automationPermanent non-goals (PRD §3 and §3 above).

Deferred — genuine candidates, not yet planned:

Karate featureNote
Dynamic data-driven Examples (rows from a read('data.json') array)Table-driven coverage from an external data file. Plausible, but in scope-tension with product-neutrality and the sans-IO/determinism line (an external read at lower time). Revisit if a real need appears.
match each / schema-as-a-value reuseAchievable today via the #2 cookbook’s hurl predicates and reusable expect: macros — no new engine feature needed.

11. Sources (competitive research, 2026-07-31)

12. Round 2 — post-execution competitive re-review (2026-08-02)

Round 1 (§1–§11) is largely shipped (across the releases that followed). Round 2 re-ran the Karate + Cucumber + adjacent-landscape survey against the post-execution codebase, then put every surviving candidate through a four-stream deep code-validation pass (matcher/binding · reporting/events · lifecycle/i18n/snapshot · scope-boundary), mirroring Round 1’s method. Identifiers N1–N9 are stable and doc-local — this registry is their only sanctioned home; they never appear in code comments (per the no-task-ids-in-source rule). File:line citations validated 2026-08-02 and will drift.

12.1 What the re-survey confirmed is already shipped (positioning to defend)

The external agents, blind to the just-landed work, flagged many “gaps” that Round 1 already closed — recording them so the omission-vs-decision trail stays explicit:

Re-flagged “gap”Already shipped as
Cucumber snippet/stub suggestion on unbound step#9 stub-gen (unbound-step diagnostic prints a paste-ready macro)
Rerun-only-failures#8 --rerun (keyed on (file, name))
Soft-fail / allow-failure tag#15 @quarantine (non-gating, still reported)
Boolean tag expressions#4 proef_core::tags (fuzzed grammar)
“retry until assert passes” (Karate retry until)macro retry: → hurl [Options] retry (retries until asserts pass or budget ends)
Data-table → step argumentsbind.rs merges | key | value | rows into macro args
One trial per Examples rowoutline expansion → ScenarioDef per row → one harness Trial
Explicit skipped/pending statusStatus::Skipped (post-failure steps) + Warned (optional)
Whole-run JSONL event record; git-native plain text; JUnit/SARIF/HTMLADR-0008 event spine; .feature+YAML+proef.toml; the reporter family

Convergent-evolution note (validates the architecture, nothing to adopt): Karate v2’s karate-events.jsonl (2025) is proef’s ADR-0008 event stream re-invented; Bruno’s plain-text git-native rise is proef’s text model; both confirm the design is industry-aligned.

12.2 The validated Round-2 roadmap (master table)

Verdict legend as §5 (✅ FITS · ⚠️ NEEDS-ADAPTATION · 🚫 rejected/premise-broken). Effort: S ≤ ~1 day · M ~days · L ~weeks.

#ItemVerdictLives inEffortArchitectural truth (validated 2026-08-02)
N2Run-level SLA thresholds (p95/max(duration) gate)✅proef.toml [sla] + cli exec.rsS–MStrongest — zero schema change. duration_ms already on StepFinished (event.rs:74) and in the record; a pure CLI fold. Config as [sla] (env-overridable, ADR-0012), not a flag. Breach = TestFailure (exit 1) folded before exec.rs:361; malformed table = exit 2; no new exit code. Must be opt-in by presence of [sla] — absent, behaviour is byte-identical, so pinned exit-0 tests + reference snapshot are untouched. Distinct from hurl per-request duration < (aggregate vs per-entry) → keep SLA aggregate-only, one home.
N6aHTML per-scenario timing waterfall✅core html.rsSZero schema change. Step start-offset = cumulative sum of prior duration_ms in the scenario, width = own duration_ms; new render in html.rs:176. Cannot show cross-worker occupancy (no clock/worker id) — intra-scenario only. Ship this first.
N8i18n # language: — verify, fix, test, keep the claim⚠️core feature.rs + testsSClaim asserted twice (PRD.md:104, TECH-SPEC.md:110) but unverified. gherkin-0.16 honours the header transparently; proef strips no keywords. One English-only bug: feature.rs:167 detects outlines via keyword.contains("Outline"/"Template") — false under any dialect (fr Plan du scénario, de Szenariogrundriss). Blast radius small (only a no-Examples malformed outline degrades). Fix: detect via !examples.is_empty(); add a localized fixture + byte-span test. Keep the docs claim — fix, don’t retract.
N9Curated expect: shape-macro library (expectUuid, expectIsoDate, expectNonEmptyList…)✅helpers/*.yaml + docsSNew — surfaced by validation. Augments the existing expect:/MergedAsserts mechanism (step.rs:60, emitted emit.rs:192); zero engine/core change; product-neutral (generic shapes only). Same lever as the §12.5-B ergonomic uplift. Narrows the deep-equality ergonomic gap — not the semantic one (§12.4).
N7Near-duplicate macro lint (extend proef macros)⚠️core sim-fn + cli commands.rsS–MAbsent; reuse literal_skeleton (matcher.rs:236) + levenshtein (matcher.rs:260). Extend the shipped dead-macro report (commands.rs:270), not a load pass (those are hard errors). Tight heuristic — skeleton-equal-modulo-captures — or it false-positives on the shipped corpus (boardShows* family; activateChannel “…and ready”). Advisory JSON field beside unused (commands.rs:306), exit 0, never a gate. Drop the conjunction + “organize-by-domain” sub-lints (false-positive on shipped prose; no machine model of “domain”).
N4TAP reporter⚠️cli new tap.rs via --output tapMValid, but the “surface hurl’s native TAP” rationale is wrong — proef calls run_entries in-process, never shells out; TAP must derive from the event spine (scenario = test point), like every reporter. Live Reporter (report.rs:120) → inherits sink redaction. Plan count from exec.rs:194. @quarantine → # TODO needs the non_gating set injected (it’s computed at exec.rs:313, not in the stream). One surface: --output tap (reuses the stdout-ownership machinery), never also a proef tap replay.
N1Typed parameter types in the matcher ({int}/{uuid}/custom, bind-time)⚠️matcher/bind in coreMGenuinely absent (captures are untyped strings, matcher.rs:221; params: Vec<String>). Must be declaration-site, not inline {name:type}: 3 of 4 arg sources aren’t captures (data-table, defaults, with:, and use:-only macros have no pattern), and inline typing is invisible to proef schema. One-canonical forces a single params spelling (params: {q: uuid}, bare = any) via custom Deserialize → breaking pack migration → needs an ADR. Model on the fake::GENERATORS typed registry. Two-tier caveat: skip validation when the raw arg contains ${/{{ (resolves later) → a best-effort literal-args lint, not a type system. Diagnostic proef::bind::param_type_mismatch at bind.rs:193 (+ defaults at validate.rs:57, with: at validate.rs:353).
N3Suite-level setup/teardown (once-before / once-after)⚠️proef.toml [run] + cli exec.rsMReal gap (only per-scenario Background; teardown is engine-internal session.finish()). Premise correction: tags never reach the core runner — ScenarioSpec/ScenarioOutcome carry no tags/gating (runner.rs:30,131); quarantine is a CLI-edge non_gating set (exec.rs:313). So use a proef.toml [run] setup/teardown construct, not a tag (a tag would entangle with --tags/--rerun/flows/dedup + undefined ordering). Orchestrate in execute() around runner::run (exec.rs:220); state crosses only via saveAs: global, which must merge before the parallel pool snapshots the store (runner.rs:439). Explicit failure short-circuit required — an assert-failed setup is fault:None→exit 1 and would not abort the pool (cascading failures on un-seeded state); refuse to launch and surface the fault. Excluded from build_specs so it never double-runs; --dry-run unaffected (never calls runner::run).
N5Golden response snapshots🚫—LRejected. Response bodies exist but are discarded (HurlResult…calls[].response.body; session reads only captures/errors). It is a second assertion mechanism competing with hurl body-asserts + expect:, and whole-body regression is already proef diff’s job. No normalization machinery exists (sink redaction masks only known injected values, report.rs:29). Secret-leak risk: backend-minted tokens/PII would be committed unredacted — against ADR-0005 (session.rs:346 already refuses to persist a capture equal to a secret). Do not build. If ever needed: diff-time over run records, never committed goldens.

12.3 Verdict-change ledger (validation overturned the first sketch)

The deep pass is on the record because it changed conclusions — the point of validating:

ItemFirst sketchAfter validationWhy
N5 golden snapshots⚠️ candidate🚫 rejectedDuplicates hurl asserts + diff; no normalization; leaks backend-minted secrets
N4 TAP rationale“surface hurl’s native TAP”corrected: derive from the event spineproef never shells out; hurl --report-tap is unreachable + per-file, not per-scenario
N3 selector@setup/@teardown tagproef.toml [run] constructtags never reach the core runner; a tag overloads “filter” with “phase” and races the pool
N6 timelineone “timeline” itemsplit N6a (zero-schema, now) / N6b (injected timestamp_ms, later)true cross-worker occupancy needs an injected clock/worker id
N1 typed params“small matcher tweak”M + ADR + breaking migrationdeclaration-site forced by schema coherence; single spelling forced by one-canonical; literal-args-only forced by two-tier vars

12.4 The named architectural ceiling (accepted, not a defect)

Karate’s match response == { id:'#uuid', items:'#[]' } — order-insensitive whole-body deep-equality with type-holes, exhaustive-key checking, and one readable structural diff — cannot be assembled under hurl-only (asserts are path-at-a-time). Reusable expect: macros over hurl jsonpath cover per-path type/value/shape, collection membership (contains), cardinality (count), optional keys (exists/not exists), and RFC-9535 filtered queries (AUTHORING.md:100) — the ergonomic gap, narrowed further by N9. The semantic gap (single order-insensitive whole-body diff + exhaustiveness) stays open by design. The Round-1 marker-DSL rejection stands (§3, ledger §10 — a second assertion mechanism). This is a deliberate ceiling of the hurl-only bet, stated honestly, not engineered away.

12.5 Non-goal-adjacent — explicit scope decisions (keep excluded absent an ADR)

The two biggest capabilities a market reviewer would name are on/over the PRD §3 line. The governing boundary the validation extracted: CLI-edge IO that injects values into the sans-IO core is sanctioned (ADR-0012); IO that re-shapes the corpus or acts as a recurring oracle is contract testing (out).

  • A — OpenAPI → scenario generator (proef generate). Verdict: needs an ADR; default = deferred/out-of-scope. Strict generate-then-freeze clears sans-IO/determinism (ADR-0012 precedent) and echoes #9 stub-gen at suite granularity — but the bright line is “the spec may be a one-shot seed; it may never become a recurring oracle.” It sits one --check flag from OpenAPI-drift (§3 non-goal), introduces a new inward generation direction, and pressures one-canonical-way on regeneration (a second maintenance path). Only an ADR that bans the oracle/drift mode and accepts the OpenAPI dependency can green-light even the narrow scaffolder. → Now settled in ADR-0016 (Proposed): the oracle/drift mode is permanently rejected; the narrow one-shot scaffolder is deferred (output-quality + dependency cost) but buildable later under the bright line.
  • B — JSON-Schema conformance assert. Verdict: shape/type conformance is already-achievable today via expect: + hurl type predicates (the #2 cookbook — zero new features); the ergonomic uplift is N9 (curated shape macros). Full external .schema.json whole-body validation is out — hurl has no jsonschema predicate, and using the API’s canonical schema as oracle is drift-detection.

12.6 Prerequisites & one-canonical-way watch-list (round 2)

  • P4 — injected per-event timestamp_ms/worker (Option, skip_serializing_if), stamped by a CLI sink-wrapper on the worker thread (the run_id injection pattern), left None by the sans-IO core; kept off RunStarted (exact-bytes pin event.rs:157). Note: old records still parse (additive holds), but the new-run reference snapshot changes and needs a new insta filter + deliberate review. Unlocks N6b.
  • ADR needed: N1 (params-shape migration + literal-args-only semantics); Tier-3-A OpenAPI (bans the oracle mode).

One-canonical watch-list — each must fold into an existing mechanism, never ship beside it:

ItemMust replace / augment (not duplicate)
N1 typed paramsOne params spelling (name→type map); no second inline {name:type} form
N3 setup/teardownExactly one [run] setup + one teardown; not a tag, not a second Background
N4 TAPOne surface (--output tap); no parallel proef tap replay
N9 shape macrosAugment the existing expect: mechanism; never a schema/marker DSL
N2 SLAAggregate run/scenario budget only; per-request latency stays hurl duration <
  1. Batch E — free / small / zero-schema — SHIPPED 2026-08-02: N2 SLA (opt-in) · N6a waterfall · N9 expect: library · N8 i18n verify+harden.
  2. Batch F — small–medium — SHIPPED 2026-08-02: N7 near-duplicate lint · N4 TAP (--output tap).
  3. Batch G — medium, design/ADR call first — mostly SHIPPED 2026-08-02: N6b full timeline (ADR-0015, P4 delivered) · N3 setup/teardown (ADR-0014, [run] setup/teardown). N1 typed params — deferred (ADR-0013): the research-grounded call given proef’s deferred-heavy, string-ish corpus.
  4. Blocked pending an ADR: Tier-3-A OpenAPI generator. Rejected: N5 golden snapshots.

12.8 Sources (round 2)

Competitor sources unchanged from §11 (Karate/Cucumber/Hurl/Bruno/Schemathesis/Pact/k6/ Playwright). Round-2 findings are code-internal — every verdict is anchored to a file:line validated 2026-08-02, not to an external claim.