proef — Improvement Plan
Status: complete. Round 1 (§1–§11) and Round 2 (§12) are both shipped — every
batch (N2·N6a·N9·N8·N7·N4·N6b·N3), with N1 deferred per ADR-0013, N5 rejected, and the
OpenAPI generator settled in ADR-0016. This file is the historical feature roadmap
and its file:line citations are from 2026-07/08; the live worklist is
OPEN-FINDINGS and the running ledger is CLAUDE.md’s Status block.
· Date: 2026-07-31, appended 2026-08-02, closed 2026-08-31 · Owner: Emre
Companion docs: PRD (scope + the binding non-goals, §3), adr/ (the
invariants every item must respect), TECH-SPEC (types/pipeline),
IMPLEMENTATION-PLAN (milestones + definition of done).
0. What this is (and isn’t)
A competitive analysis of proef against the BDD and API-testing field, converted into a roadmap of candidate improvements — each validated against the current architecture at file:line, and filtered strictly through proef’s permanent non-goals (PRD §3).
Nothing here is committed work; it is the durable output of the post-M5 review round. Effort grades (S/M/L) and sequencing are advisory. Feature numbers (#1…#15) are stable identifiers used across the tables. File:line citations are as of 2026-07-31 and will drift — treat them as “start reading here”, not addresses.
Every item is in scope by construction: none proposes a second engine, mocking, contract testing, load testing, a dashboard/server, OpenTelemetry, or importing hand-written hurl — all permanent non-goals (§3 below). The work is almost entirely surfacing data the sans-IO core already computes, not new engine capability.
1. Headline finding
proef’s engine is already ahead of its cohort; the gaps are in reporting surfaces and
authoring/maintenance DX, not in HTTP power. Because proef pins hurl 8.0.1, it
already ships hurl 8.0’s full assertion arsenal (RFC 9535 JSONPath with filter functions,
type predicates isUuid/isIsoDate/isString/isObject/isList, response-time
duration < ms, the filter chain split/count/toDate/base64Decode/daysAfterNow).
Pack authors can write Karate-grade assertions today, inside raw hurl: blocks — they
are just undocumented and under-surfaced. The roadmap is therefore mostly exposure and
tooling, which is cheap, rather than engine work, which is done.
2. Where proef already wins (positioning to defend)
| proef strength | Competitor weakness it beats |
|---|---|
| One canonical way, raw-hurl-only, no escape-hatch language | Karate’s most-cited flaw: a three-language model (Gherkin + DSL + embedded JS) the moment anything gets non-trivial |
| Deterministic sans-IO core | Karate/Tavern/Newman are non-deterministic; reproducibility is now a headline selling point |
Artifacts = executed bytes (hash-locked, git-diffable, replayable hurl --test) | Postman’s opaque JSON collections (unreviewable merges) — the reason Bruno is displacing it |
| Finite retries + budgets + watchdog (ADR-0007) | hurl itself has no cancellation and unbounded retries — proef fixes its own engine’s biggest gap |
| Property-tested secret masking + typed exit codes (ADR-0009) | Most tools treat masking loosely and lack a stable exit-code contract |
| Dev-maintained macro packs = an enforced “step dictionary” | The AI-authoring trend is groping toward exactly this; proef has it structurally |
| No cloud, no account, plain text | Postman’s 2026 pricing exodus is driving the whole git-native wave |
3. Scope guardrails — what this plan will NOT propose
Permanent non-goals (PRD §3) — never revisit under “competitor parity”: further engines (browser/gRPC/etc.), API mocking, contract testing (OpenAPI drift / Pact / Schemathesis), load testing, a desktop dashboard or server mode, OpenTelemetry export, dynamic plugin loading, importing hand-written hurl into Gherkin (artifacts flow outward only), static musl/Windows binaries.
Named anti-patterns (from the 2025–2026 trend research) to avoid: silent retries / green-on-attempt-2; retry ceilings > 3 or unbounded (proef’s finite cap is already correct); permanent quarantine without owner+expiry; treating masking as a security boundary; LLM self-healing / non-deterministic test mutation; config sprawl.
A Karate-style marker DSL (#uuid, ##optional) is explicitly rejected — it would be
a second assertion mechanism competing with raw hurl predicates. Achieve the same
readability by surfacing hurl’s native predicates (item #2), not by inventing a layer.
The overriding design gate is one-canonical-way (see §6): four items must replace or augment an existing mechanism, never add a parallel knob.
4. What code validation changed about the roadmap
Three deep code-validation passes (reporting, CLI/tags, authoring) reshaped the outside-in list. Six meta-findings:
-
“Surface, don’t build” is confirmed at file:line. The sans-IO core already computes and records the data behind most items:
attempts/duration_ms/detailonStepFinished(proef-core/src/event.rs), the winningmacro_nameper bound step (bind.rs:20), deterministic seeded fakes, and hurl’s owncurl_cmd(already returned byEntryResult, currently discarded in engine-hurlsession.rs). -
Two items are already ~90% built — validation downgraded them.
- #14 seeded fakes: fakes are already deterministic (hand-rolled SplitMix64, no
randcrate,fake.rs) and already seeded by the injectedrun_id, which is already recorded inRunStarted. The feature collapses to “lettestpinrun_idthe wayartifacts --run-idalready can.” - #9 stub-gen: the “did you mean” fuzzy suggestion already exists (
bind.rs:214, levenshtein); only the paste-ready stub template is missing.
- #14 seeded fakes: fakes are already deterministic (hand-rolled SplitMix64, no
-
One item’s premise is broken — #13 impacted-only re-run. There is no stored content hash anywhere; ADR-0010 is enforced as a byte-identity
assert_eq!on outputs (crates/proef-cli/tests/execute.rs:188), not a reusable digest. The only honest impact fingerprint — the emitted.hurl— is deliberately not stable run-to-run (run_idlives in runtime globals). #13 is gated behind a determinism prerequisite. -
Two plumbing gaps gate a cluster. Scenario tags stop at the CLI edge — they never reach the runner (
ScenarioSpec/ScenarioOutcomeinrunner.rscarry no tags) or the event stream — so #15 (quarantine) and part of #8 (rerun) need new plumbing. And #4 (boolean tags) is hard-blocked byvalue_delimiter=','on--tags(main.rs:57); it is a contract-changing replace, not an add. -
The recurring architectural rule is one-canonical-way. Four items (#4, #9, #14, #15) each risk a second mechanism; each must replace or augment an existing one.
-
The exit-code contract (ADR-0009) is a live wire for #15. Quarantine changes which scenarios feed
RunSummary::exit_code(fine — the fold stays pure in core) but must extend the pinned assert_cmd tests and never let a quarantined system fault mask exit 3.
5. The validated roadmap (master table)
Status is what exists in the tree, re-verified against main on 2026-09-08 —
--help for flags, source for the rest. Verdict is the 2026-07-31 judgement of
whether the idea fits the architecture; it never meant “done”, and reading it that
way is why this table looked like a backlog when, by 2026-08-10, 13 of its 16 items
had shipped (14 today; one partial, one gated).
Status: shipped · partial · open · gated (premise rejected). Verdict legend: ✅ FITS · ⚠️ NEEDS-ADAPTATION · 🚫 premise broken. Effort: S ≤ ~1 day · M ~days · L ~weeks.
| # | Item | Status | Verdict | Lives in | Effort | Architectural truth (as of 2026-07-31) |
|---|---|---|---|---|---|---|
| 2 | Assertion cookbook (surface hurl-8.0 predicates/filters) | shipped | ✅ docs-only | docs/ | S | Lowering copies all but ${…} verbatim (resolve.rs:163); load runs the real hurl_core parser (engine-hurl/src/lib.rs:95). Predicates already work. Caveat: grammar validation, not JSONPath semantics. |
| 1 | GitHub ::error file=,line=,title= annotations | shipped | ✅ | proef-cli ci_reports.rs | S | file+line+detail already flow to write_github_summary (ci_reports.rs:98). Gate stdout vs --output json; percent-encode multiline detail. Line-only (no byte-span at runtime). |
| 5 | --curl export per request | partial | ✅ | engine-hurl session.rs | S | hurl’s EntryResult.curl_cmd is already returned, just discarded (session.rs:346). Must redact (holds resolved secrets, ADR-0005). Fold into the existing reproduce: block (exec.rs:230). Partial as of 2026-08-10: the curl line is surfaced on a failing step (exec.rs), which covers the debugging case; a per-request export flag for passing steps is still open. |
| 3a | “passed on attempt N” badge (JUnit/summary) | shipped | ✅ | proef-cli ci_reports.rs | S | attempts:u32 already on StepFinished (event.rs:72) + StepOutcome; JUnit ignores it today (ci_reports.rs:43). |
| 9 | Stub-gen for unbound steps | shipped | ⚠️ | proef-core bind.rs | S | Augment the existing did-you-mean help (bind.rs:89), zero-match arm only — not a new command. Derive {param} from quoted tokens (matcher already sheds quotes, matcher.rs:85). |
| 10 | SARIF export of --dry-run diagnostics | shipped | ✅ | proef-cli new sarif.rs | S–M | Diag (diag.rs:56) → SARIF result ~1:1: code→ruleId, byte span→region.byteOffset. Pre-populate rules[] from the closed diagnostic-code set. A parallel serializer to render.rs. |
| 14 | --seed (reproducible fakes) | shipped (as run-id) | ⚠️ | proef-cli main.rs/exec.rs | S | Thread into front::run’s existing run_id param (artifacts already exposes --run-id, main.rs:101). Caveats: arbitrary seed breaks JUnit’s UUID parse (ci_reports.rs:22); occurrence is per-scenario, not per-run — identical ${fake:X} at the same position in two different scenarios still draws the same value (a known limitation; see OPEN-FINDINGS). Note: the per-step reset this row originally cited (Refs::default() on every lower() call) was fixed in 0.6.0 — the counter now threads through lower with a high-water mark. Only the cross-scenario half remains. Shipped as the one knob §7 demanded (no separate --seed): test --run-id pins the fakes, and --shuffle seeds its permutation from the same id. |
| 7 | Dead-macro / usage report | shipped | ✅ | proef-cli new macros --usage | S–M | BoundStep.macro_name (bind.rs:20) vs packs.macros. Count use:-only macros (pattern:None, pack/mod.rs:165) as reachable via the use: graph. Report the whole corpus, not a --tags subset. |
| 6 | Self-contained HTML report | shipped | ✅ | core render_html(&[Event]) + cli write | M | Post-hoc proef report <run-id> replaying events.jsonl like explain (explain.rs:12) — the command is the HTML report, so it took no --html flag; -o picks the file. Bodies live in artifacts/ — deep-link, don’t inline. Derived view, never a second record (ADR-0008). |
| 12 | proef diff between two runs | shipped | ⚠️ | proef-cli diff.rs | M | Identity (file,scenario) (why ADR-0008 added file, event.rs:86); key step diffs on text not line (lines shift on edit). attempts+duration_ms → free flakiness/perf-regression detector. Pre-file records replay file="". |
| 4 | Boolean tag expressions (@a and not @b) | shipped | ⚠️ (replace) | grammar in core, apply in cli front.rs | M | value_delimiter=',' (main.rs:57) actively breaks and/or/not; must replace the CSV/OR contract (front.rs:388), keep empty-match=exit-2 (front.rs:382). Grammar/evaluator is deterministic → proptest/fuzz-shaped, belongs in core. |
| 8 | Rerun-only-failures (--rerun) | shipped | ⚠️ | proef-cli, reuse explain replay | M | explain already reads the latest record + failed (file,name) (explain.rs:100, event.rs:83). Needs a multi-identity predicate (today’s scenario/scenario-file filters are single-valued, exec.rs:301,321). Factor a shared record::failed_scenarios. |
| 15 | @quarantine non-gating tag | shipped | ⚠️ (contract) | thread gating:bool core+cli | M | Tags must first reach ScenarioSpec/ScenarioOutcome (P1). exit_code() stays pure in core (runner.rs:89) and skips non-gating outcomes. Events still emit the scenario → not hidden. Extend the pinned assert_cmd tests; never mask a Fault::System (exit 3). |
| 3b | True <flakyFailure> with earlier-attempt detail | shipped | ⚠️ | schema + engine-hurl | M | Needs an additive attempt_details field (ADR-0008 additive-only) + engine-hurl collecting per-retry bodies before the final one. Bigger than 3a. |
| 11 | proef lsp (feature/pack language server) | shipped | ✅ | new proef-lsp crate | L | All diagnostic substrate is headless/sans-IO already (bind, pack::load, resolve Probe mode, matcher). New: a sync lsp-server (tokio ban forbids async), a byte-offset→token API (not exposed), and a partial-results wrapper (bind/load are all-or-nothing today). Karate notably lacks good IDE support → differentiator. |
| 13 | Impacted-only re-run (content-hash) | gated | 🚫 | — (gated) | L | No input hash exists; raw-input hashing is unsound (shared packs, use: nesting, config vars fan out). Honest fingerprint = per-scenario emitted .hurl, but it is not run-to-run stable (run_id in globals). Needs a determinism prerequisite first; silent-green risk. 2026-09-07: a suite-level input hash now exists (inputs.json, for flaky windows) — not per-scenario; still gated. |
6. Prerequisites that unlock clusters
- P1 — carry scenario tags + a
gatingflag past the CLI edge intoScenarioSpec/ScenarioOutcome(runner.rs:30,121) and, additively, the event stream. Done (RF waves): tags rideScenarioSpec/ScenarioOutcomeandscenario_finished.tags; gating became the reserved-tag instruction plus the non-gating list (ADR-0019) rather than a bool. #15 and #8 both shipped on top of it. - P2 — a shared
record::failed_scenarios(run_id)+ a multi-identity scenario predicate. Reused byexplain,--rerun(#8), andproef diff(#12). - P3 — a deterministic emitted-
.hurlfingerprint (stable run-to-run despiterun_id). Prerequisite for #13; do not attempt #13 without it. 2026-09-07:proef_core::fingerprintnow hashes a run’s whole input set (feature sources, loaded macros and fragments, the resolved[url]/[vars]scope) intoinputs.json—proef flaky’s equivalence class. Suite-level and deliberately coarse (any edit ends a window), so it is not the per-scenario,run_id-independent fingerprint #13 needs; P3 stands.
7. The one-canonical-way watch-list
Each of these must fold into an existing mechanism, never ship beside it:
| Item | Must replace / augment (not duplicate) |
|---|---|
| #4 boolean tags | Replace the CSV/OR --tags semantics — no second tag syntax |
| #9 stub-gen | Augment the existing did-you-mean Diag.help — no separate proef stub command |
#14 --seed | Fold into the existing run_id determinism knob — no parallel seed unless it replaces run_id-keyed fakes |
| #15 quarantine | Exactly one non-gating tag name; must not spawn a second “skip” concept |
#5 --curl | Attach to the single reproduce: mechanism, not a parallel debug path |
8. Recommended sequencing
- Batch A — free / small, all FITS, all reuse existing data.
#2 cookbook (docs) → #1 annotations → #5
--curl→ #3a attempt badge → #9 stub. - Batch B — small, high-leverage. #10 SARIF · #7 dead-macro · #14
--seed. - Prereqs → Batch C — medium, now unblocked. Build P1+P2, then #6 HTML · #12 diff · #8 rerun · #15 quarantine · #4 boolean tags · #3b flaky-detail.
- Batch D — strategic. #11 LSP. And #13 only after committing to P3.
9. Per-item detail & competitor provenance
Each entry: what it borrows from whom → the validated architectural note. Numbers cross- reference §5.
#2 Assertion cookbook — from Karate’s fuzzy markers + hurl’s own docs. Confirmed a
docs task: resolve() leaves everything but ${…} byte-for-byte (resolve.rs:163-208,
test runtime_tier_passes_through), and pack load validates the full grammar via
hurl_core::parser::parse_hurl_file (engine-hurl/src/lib.rs:95), re-checked on the
emitted artifact (front.rs:159). Extension: ship a tests/features/ reference feature
exercising each predicate so the cookbook is snapshot-locked against hurl upgrades (the
canary catches drift).
#1 GitHub annotations — from the 2025–2026 CI-reporting shift (annotations displace
log-diving). The failures loop already prints `{file}:{line}` — {detail}
(ci_reports.rs:98). Emit ::error workflow commands as a sibling; title = scenario +
step text. Risk: stdout is owned by --output json (exec.rs:114) — gate it.
#5 --curl export — from hurl’s loved --curl; Bruno/Postman “copy as curl”. hurl
hands us EntryResult.curl_cmd already (session iterates result.entries at
session.rs:346 but reads only captures/errors/duration). Must pass Redactions
(session.rs:396) before any sink — the curl line contains resolved secrets. Cannot be
derived pre-execution (needs runtime {{…}}).
#3a “passed on attempt N” — from the flaky-test-honesty consensus (never hide a
retry). attempts is first-class (event.rs:72, step.rs:149) and already printed on
the console (report.rs:220); JUnit simply drops it. Count-based badge is S.
#9 Stub-gen — from Cucumber/Behave snippet generation. The matcher already computes
the nearest macro via closest_pattern/levenshtein (bind.rs:214, matcher.rs:248);
add a match:+hurl: | skeleton to the help text for the zero-match arm only (an
ambiguous step, bind.rs:99, must not get a stub).
#10 SARIF — from SARIF’s rise for static/validation findings inline in PRs. Dry-run
diags are a structured Vec<Diag> before miette (diag.rs:123, front.rs:64). Diag
maps ~1:1 to a SARIF result; the closed code set (one per tests/errors/ dir) pre-fills
rules[]. cli-edge serializer, no core change.
#14 --seed — from seeded-faker reproducibility. Fakes already deterministic
(fake.rs:12 SplitMix64/FNV, seeded fnv1a(run_id) ^ …), seed already recorded in
RunStarted{run_id} (event.rs:26). Design fork: alias run_id (zero core change, but
must stay uuid-parseable for JUnit) vs a dedicated recorded seed field (cleaner, but
a second knob — resolve per one-canonical-way). Known limit: the occurrence counter is
scoped per scenario (lower.rs:69-86, Refs::fakes, threaded through resolve()’s
fakes: &mut usize parameter, not reset per call) → cross-step uniqueness within a
scenario now holds, but cross-scenario uniqueness still does not: two different scenarios
each resolving ${fake:X} at the same position in their own step order get the same value.
Document before advertising “unique fakes” — it means per-scenario, not per-run or
per-entity.
#7 Dead-macro report — from Cucumber’s usage formatter marking UNUSED. Binding
records BoundStep.macro_name (bind.rs:192); iterate
front.features[].scenarios[].bound.steps[].macro_name vs packs.macros.keys()
(pack/mod.rs:116). use:-only macros (pattern:None) need reachability via the use:
graph (pack/validate.rs:628) to avoid false “unused”. Report the whole corpus.
#6 HTML report — from Cucumber/Karate/hurl HTML reports; the industry’s convergence on
the Cucumber-Messages/JSONL stream. Core render_html(&[Event]) -> String, cli writes;
best as post-hoc over events.jsonl (explain.rs already replays it) so historical runs
render. Events are pre-redacted at the sink (report.rs:124).
#12 proef diff — from Allure history / test-observability-without-OTel. Identity is
(file, scenario) (report.rs:147 ScenarioKey; ADR-0008 added file for exactly this).
Key step diffs on text, not the volatile line. attempts+duration_ms make it a
flakiness/perf-regression detector. run_id is uuid-v7 → chronology recoverable.
#4 Boolean tags — from Cucumber tag expressions (and/or/not/()). Single filter fn
tag_selected (front.rs:388), three callers. value_delimiter=',' (main.rs:57) blocks
the operator syntax → drop it, take one expression string, replace the CSV contract.
Grammar/evaluator → core (deterministic, fuzz-shaped). Preserve empty-match=exit-2.
#8 Rerun-only-failures — from Cucumber’s rerun formatter (@rerun.txt). explain
already discovers + replays the latest record and extracts failed identities
(explain.rs:63). Add a multi-identity predicate reusing build_specs (exec.rs:289).
Empty failure set → reuse no_scenarios_matched (exit 2), never silent-pass.
#15 Quarantine — from the flaky-quarantine-with-owner+expiry consensus. Thread
gating:bool from CLI (which sees scenario.lowered.tags, exec.rs:318) into
ScenarioSpec/ScenarioOutcome; exit_code() (runner.rs:89) skips non-gating outcomes
and stays pure in core. Events unchanged → scenario still reported. Extend the
cli.rs/execute.rs exit-code assertions; a quarantined Fault::System still exits 3.
#3b Flaky-failure detail — from JUnit <flakyFailure> / Allure retries. Needs an
additive attempt_details on StepFinished (ADR-0008 additive-only) and engine-hurl
collecting per-retry messages (hurl retry is per-entry internal — verify the adapter isn’t
already discarding earlier bodies).
#11 proef lsp — from Cucumber’s language server (unbound-step diagnostics,
go-to-def, completion); a gap Karate never closed. Reuses feature::parse, bind,
pack::load, resolve Probe mode, matcher — all headless, all with stable codes +
byte-offset spans that already map to editor ranges. New work: sync lsp-server (tokio
banned), byte→token API, and a “collect diags, don’t early-return” wrapper (bind/load are
all-or-nothing today).
#13 Impacted-only re-run — from selective/affected-test re-run. Premise broken: no
reusable input hash (ADR-0010 is a byte-identity assert_eq! on outputs,
execute.rs:188), and the honest fingerprint (emitted .hurl) is not run-to-run stable
because runtime globals include run_id (execute.rs:304). Gated on P3; a hash miss must
never skip a scenario that would fail (needs --force/first-run fallback).
10. Karate feature ledger — considered / adopted / rejected
Karate (github.com/karatelabs/karate) was the closest competitor and the deepest research stream. This ledger makes the “considered → decision” trail explicit, so each Karate idea is an intentional call rather than an omission.
Adopted — drove a plan item or the positioning:
| Karate feature | proef outcome |
|---|---|
Inline fuzzy markers (match response == { id: '#uuid', age: '#number' }) | #2 assertion cookbook — the same readability via hurl 8.0’s native predicates (isUuid, isIsoDate, …) surfaced in docs, not a new marker DSL. |
| Weak/immature IDE support (Karate’s own gap) | #11 LSP — reframed as a differentiator proef can win, since Karate never closed it. |
| HTML report with a timeline view | #6 self-contained HTML report (with Cucumber’s and hurl’s). |
| Three-language cognitive load (Gherkin + DSL + embedded JS) | proef’s headline positioning (§2): one-canonical-way, raw-hurl-only, sans-IO — the inverse of Karate’s most-cited flaw. |
call / callonce cross-feature reuse | Already covered by proef’s use: / with: macro composition (ADR-0004) — no new work. |
Rejected — with the reason (so it stays rejected):
| Karate feature | Why not |
|---|---|
A marker DSL (#uuid, ##optional, #? _ > 0) | A second assertion mechanism competing with raw hurl predicates — violates one-canonical-way (§3). Readability comes from #2 instead. |
| Embedded JavaScript escape hatch | Conflicts with the sans-IO deterministic core and one-canonical-way — it is the thing proef exists to avoid. |
Soft assertions (configure continueOnStepFailure) | Conflicts with proef’s deliberate stop-at-first-failed-step model (the ∅ cascade); a failed step’s downstream is intentionally not run. |
Service mocking · karate-gatling perf · UI automation | Permanent non-goals (PRD §3 and §3 above). |
Deferred — genuine candidates, not yet planned:
| Karate feature | Note |
|---|---|
Dynamic data-driven Examples (rows from a read('data.json') array) | Table-driven coverage from an external data file. Plausible, but in scope-tension with product-neutrality and the sans-IO/determinism line (an external read at lower time). Revisit if a real need appears. |
match each / schema-as-a-value reuse | Achievable today via the #2 cookbook’s hurl predicates and reusable expect: macros — no new engine feature needed. |
11. Sources (competitive research, 2026-07-31)
- Karate — match keyword / fuzzy markers / reuse: https://docs.karatelabs.io/assertions/match-keyword/, https://docs.karatelabs.io/reusability/calling-features/
- Cucumber-JS formatters (usage / rerun / snippets / html): https://github.com/cucumber/cucumber-js/blob/main/docs/formatters.md
- Reqnroll HTML report + Cucumber Messages; SpecFlow EOL: https://reqnroll.net/news/2025/06/roadmap-update-html-report/, https://reqnroll.net/news/2025/01/specflow-end-of-life-has-been-announced/
- Hurl 8.0 (RFC 9535 JSONPath,
--curl, TAP, secrets redaction) + release history: https://hurl.dev/blog/2026/04/27/announcing-hurl-8.0.0.html, https://hurl.dev/docs/filters.html - Bruno vs Postman (git-native trajectory): https://www.usebruno.com/compare/bruno-vs-postman
- GitHub Actions workflow commands (annotations + job summaries): https://docs.github.com/en/actions/reference/workflows-and-actions/workflow-commands
- Flaky-test quarantine/hardening; GitHub Actions OIDC/masking limits; Allure history: https://pie.inc/blog/flaky-tests-cicd/, https://www.stepsecurity.io/blog/github-actions-security-best-practices, https://allurereport.org/docs/how-it-works-history-files/
- Cucumber Language Server: https://github.com/cucumber/language-server
12. Round 2 — post-execution competitive re-review (2026-08-02)
Round 1 (§1–§11) is largely shipped (across the releases that followed). Round 2 re-ran the Karate + Cucumber + adjacent-landscape survey against the post-execution codebase, then put every surviving candidate through a four-stream deep code-validation pass (matcher/binding · reporting/events · lifecycle/i18n/snapshot · scope-boundary), mirroring Round 1’s method. Identifiers N1–N9 are stable and doc-local — this registry is their only sanctioned home; they never appear in code comments (per the no-task-ids-in-source rule). File:line citations validated 2026-08-02 and will drift.
12.1 What the re-survey confirmed is already shipped (positioning to defend)
The external agents, blind to the just-landed work, flagged many “gaps” that Round 1 already closed — recording them so the omission-vs-decision trail stays explicit:
| Re-flagged “gap” | Already shipped as |
|---|---|
| Cucumber snippet/stub suggestion on unbound step | #9 stub-gen (unbound-step diagnostic prints a paste-ready macro) |
| Rerun-only-failures | #8 --rerun (keyed on (file, name)) |
| Soft-fail / allow-failure tag | #15 @quarantine (non-gating, still reported) |
| Boolean tag expressions | #4 proef_core::tags (fuzzed grammar) |
“retry until assert passes” (Karate retry until) | macro retry: → hurl [Options] retry (retries until asserts pass or budget ends) |
| Data-table → step arguments | bind.rs merges | key | value | rows into macro args |
| One trial per Examples row | outline expansion → ScenarioDef per row → one harness Trial |
| Explicit skipped/pending status | Status::Skipped (post-failure steps) + Warned (optional) |
| Whole-run JSONL event record; git-native plain text; JUnit/SARIF/HTML | ADR-0008 event spine; .feature+YAML+proef.toml; the reporter family |
Convergent-evolution note (validates the architecture, nothing to adopt): Karate v2’s
karate-events.jsonl (2025) is proef’s ADR-0008 event stream re-invented; Bruno’s plain-text
git-native rise is proef’s text model; both confirm the design is industry-aligned.
12.2 The validated Round-2 roadmap (master table)
Verdict legend as §5 (✅ FITS · ⚠️ NEEDS-ADAPTATION · 🚫 rejected/premise-broken). Effort: S ≤ ~1 day · M ~days · L ~weeks.
| # | Item | Verdict | Lives in | Effort | Architectural truth (validated 2026-08-02) |
|---|---|---|---|---|---|
| N2 | Run-level SLA thresholds (p95/max(duration) gate) | ✅ | proef.toml [sla] + cli exec.rs | S–M | Strongest — zero schema change. duration_ms already on StepFinished (event.rs:74) and in the record; a pure CLI fold. Config as [sla] (env-overridable, ADR-0012), not a flag. Breach = TestFailure (exit 1) folded before exec.rs:361; malformed table = exit 2; no new exit code. Must be opt-in by presence of [sla] — absent, behaviour is byte-identical, so pinned exit-0 tests + reference snapshot are untouched. Distinct from hurl per-request duration < (aggregate vs per-entry) → keep SLA aggregate-only, one home. |
| N6a | HTML per-scenario timing waterfall | ✅ | core html.rs | S | Zero schema change. Step start-offset = cumulative sum of prior duration_ms in the scenario, width = own duration_ms; new render in html.rs:176. Cannot show cross-worker occupancy (no clock/worker id) — intra-scenario only. Ship this first. |
| N8 | i18n # language: — verify, fix, test, keep the claim | ⚠️ | core feature.rs + tests | S | Claim asserted twice (PRD.md:104, TECH-SPEC.md:110) but unverified. gherkin-0.16 honours the header transparently; proef strips no keywords. One English-only bug: feature.rs:167 detects outlines via keyword.contains("Outline"/"Template") — false under any dialect (fr Plan du scénario, de Szenariogrundriss). Blast radius small (only a no-Examples malformed outline degrades). Fix: detect via !examples.is_empty(); add a localized fixture + byte-span test. Keep the docs claim — fix, don’t retract. |
| N9 | Curated expect: shape-macro library (expectUuid, expectIsoDate, expectNonEmptyList…) | ✅ | helpers/*.yaml + docs | S | New — surfaced by validation. Augments the existing expect:/MergedAsserts mechanism (step.rs:60, emitted emit.rs:192); zero engine/core change; product-neutral (generic shapes only). Same lever as the §12.5-B ergonomic uplift. Narrows the deep-equality ergonomic gap — not the semantic one (§12.4). |
| N7 | Near-duplicate macro lint (extend proef macros) | ⚠️ | core sim-fn + cli commands.rs | S–M | Absent; reuse literal_skeleton (matcher.rs:236) + levenshtein (matcher.rs:260). Extend the shipped dead-macro report (commands.rs:270), not a load pass (those are hard errors). Tight heuristic — skeleton-equal-modulo-captures — or it false-positives on the shipped corpus (boardShows* family; activateChannel “…and ready”). Advisory JSON field beside unused (commands.rs:306), exit 0, never a gate. Drop the conjunction + “organize-by-domain” sub-lints (false-positive on shipped prose; no machine model of “domain”). |
| N4 | TAP reporter | ⚠️ | cli new tap.rs via --output tap | M | Valid, but the “surface hurl’s native TAP” rationale is wrong — proef calls run_entries in-process, never shells out; TAP must derive from the event spine (scenario = test point), like every reporter. Live Reporter (report.rs:120) → inherits sink redaction. Plan count from exec.rs:194. @quarantine → # TODO needs the non_gating set injected (it’s computed at exec.rs:313, not in the stream). One surface: --output tap (reuses the stdout-ownership machinery), never also a proef tap replay. |
| N1 | Typed parameter types in the matcher ({int}/{uuid}/custom, bind-time) | ⚠️ | matcher/bind in core | M | Genuinely absent (captures are untyped strings, matcher.rs:221; params: Vec<String>). Must be declaration-site, not inline {name:type}: 3 of 4 arg sources aren’t captures (data-table, defaults, with:, and use:-only macros have no pattern), and inline typing is invisible to proef schema. One-canonical forces a single params spelling (params: {q: uuid}, bare = any) via custom Deserialize → breaking pack migration → needs an ADR. Model on the fake::GENERATORS typed registry. Two-tier caveat: skip validation when the raw arg contains ${/{{ (resolves later) → a best-effort literal-args lint, not a type system. Diagnostic proef::bind::param_type_mismatch at bind.rs:193 (+ defaults at validate.rs:57, with: at validate.rs:353). |
| N3 | Suite-level setup/teardown (once-before / once-after) | ⚠️ | proef.toml [run] + cli exec.rs | M | Real gap (only per-scenario Background; teardown is engine-internal session.finish()). Premise correction: tags never reach the core runner — ScenarioSpec/ScenarioOutcome carry no tags/gating (runner.rs:30,131); quarantine is a CLI-edge non_gating set (exec.rs:313). So use a proef.toml [run] setup/teardown construct, not a tag (a tag would entangle with --tags/--rerun/flows/dedup + undefined ordering). Orchestrate in execute() around runner::run (exec.rs:220); state crosses only via saveAs: global, which must merge before the parallel pool snapshots the store (runner.rs:439). Explicit failure short-circuit required — an assert-failed setup is fault:None→exit 1 and would not abort the pool (cascading failures on un-seeded state); refuse to launch and surface the fault. Excluded from build_specs so it never double-runs; --dry-run unaffected (never calls runner::run). |
| N5 | Golden response snapshots | 🚫 | — | L | Rejected. Response bodies exist but are discarded (HurlResult…calls[].response.body; session reads only captures/errors). It is a second assertion mechanism competing with hurl body-asserts + expect:, and whole-body regression is already proef diff’s job. No normalization machinery exists (sink redaction masks only known injected values, report.rs:29). Secret-leak risk: backend-minted tokens/PII would be committed unredacted — against ADR-0005 (session.rs:346 already refuses to persist a capture equal to a secret). Do not build. If ever needed: diff-time over run records, never committed goldens. |
12.3 Verdict-change ledger (validation overturned the first sketch)
The deep pass is on the record because it changed conclusions — the point of validating:
| Item | First sketch | After validation | Why |
|---|---|---|---|
| N5 golden snapshots | ⚠️ candidate | 🚫 rejected | Duplicates hurl asserts + diff; no normalization; leaks backend-minted secrets |
| N4 TAP rationale | “surface hurl’s native TAP” | corrected: derive from the event spine | proef never shells out; hurl --report-tap is unreachable + per-file, not per-scenario |
| N3 selector | @setup/@teardown tag | proef.toml [run] construct | tags never reach the core runner; a tag overloads “filter” with “phase” and races the pool |
| N6 timeline | one “timeline” item | split N6a (zero-schema, now) / N6b (injected timestamp_ms, later) | true cross-worker occupancy needs an injected clock/worker id |
| N1 typed params | “small matcher tweak” | M + ADR + breaking migration | declaration-site forced by schema coherence; single spelling forced by one-canonical; literal-args-only forced by two-tier vars |
12.4 The named architectural ceiling (accepted, not a defect)
Karate’s match response == { id:'#uuid', items:'#[]' } — order-insensitive whole-body
deep-equality with type-holes, exhaustive-key checking, and one readable structural diff —
cannot be assembled under hurl-only (asserts are path-at-a-time). Reusable expect:
macros over hurl jsonpath cover per-path type/value/shape, collection membership
(contains), cardinality (count), optional keys (exists/not exists), and RFC-9535
filtered queries (AUTHORING.md:100) — the ergonomic gap, narrowed further by N9.
The semantic gap (single order-insensitive whole-body diff + exhaustiveness) stays open by
design. The Round-1 marker-DSL rejection stands (§3, ledger §10 — a second assertion
mechanism). This is a deliberate ceiling of the hurl-only bet, stated honestly, not
engineered away.
12.5 Non-goal-adjacent — explicit scope decisions (keep excluded absent an ADR)
The two biggest capabilities a market reviewer would name are on/over the PRD §3 line. The governing boundary the validation extracted: CLI-edge IO that injects values into the sans-IO core is sanctioned (ADR-0012); IO that re-shapes the corpus or acts as a recurring oracle is contract testing (out).
- A — OpenAPI → scenario generator (
proef generate). Verdict: needs an ADR; default = deferred/out-of-scope. Strict generate-then-freeze clears sans-IO/determinism (ADR-0012 precedent) and echoes #9 stub-gen at suite granularity — but the bright line is “the spec may be a one-shot seed; it may never become a recurring oracle.” It sits one--checkflag from OpenAPI-drift (§3non-goal), introduces a new inward generation direction, and pressures one-canonical-way on regeneration (a second maintenance path). Only an ADR that bans the oracle/drift mode and accepts the OpenAPI dependency can green-light even the narrow scaffolder. → Now settled in ADR-0016 (Proposed): the oracle/drift mode is permanently rejected; the narrow one-shot scaffolder is deferred (output-quality + dependency cost) but buildable later under the bright line. - B — JSON-Schema conformance assert. Verdict: shape/type conformance is
already-achievable today via
expect:+ hurl type predicates (the #2 cookbook — zero new features); the ergonomic uplift is N9 (curated shape macros). Full external.schema.jsonwhole-body validation is out — hurl has nojsonschemapredicate, and using the API’s canonical schema as oracle is drift-detection.
12.6 Prerequisites & one-canonical-way watch-list (round 2)
- P4 — injected per-event
timestamp_ms/worker(Option,skip_serializing_if), stamped by a CLI sink-wrapper on the worker thread (therun_idinjection pattern), leftNoneby the sans-IO core; kept offRunStarted(exact-bytes pinevent.rs:157). Note: old records still parse (additive holds), but the new-run reference snapshot changes and needs a new insta filter + deliberate review. Unlocks N6b. - ADR needed: N1 (params-shape migration + literal-args-only semantics); Tier-3-A OpenAPI (bans the oracle mode).
One-canonical watch-list — each must fold into an existing mechanism, never ship beside it:
| Item | Must replace / augment (not duplicate) |
|---|---|
| N1 typed params | One params spelling (name→type map); no second inline {name:type} form |
| N3 setup/teardown | Exactly one [run] setup + one teardown; not a tag, not a second Background |
| N4 TAP | One surface (--output tap); no parallel proef tap replay |
| N9 shape macros | Augment the existing expect: mechanism; never a schema/marker DSL |
| N2 SLA | Aggregate run/scenario budget only; per-request latency stays hurl duration < |
12.7 Recommended sequencing (round 2)
- Batch E — free / small / zero-schema — SHIPPED 2026-08-02: N2 SLA (opt-in) ·
N6a waterfall · N9
expect:library · N8 i18n verify+harden. - Batch F — small–medium — SHIPPED 2026-08-02: N7 near-duplicate lint ·
N4 TAP (
--output tap). - Batch G — medium, design/ADR call first — mostly SHIPPED 2026-08-02:
N6b full timeline (ADR-0015, P4 delivered) · N3 setup/teardown (ADR-0014,
[run] setup/teardown). N1 typed params — deferred (ADR-0013): the research-grounded call given proef’s deferred-heavy, string-ish corpus. - Blocked pending an ADR: Tier-3-A OpenAPI generator. Rejected: N5 golden snapshots.
12.8 Sources (round 2)
Competitor sources unchanged from §11 (Karate/Cucumber/Hurl/Bruno/Schemathesis/Pact/k6/ Playwright). Round-2 findings are code-internal — every verdict is anchored to a file:line validated 2026-08-02, not to an external claim.