Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

ADR-0007 — Cancellation: cooperative at batch boundaries, with budgets

Status: Accepted · Date: 2026-07-28

Context

Source-verified: hurl has no cancellation mechanism anywhere — no signal handling in the workspace, no abort check in the entry loop; delay/retry-interval are uninterruptible thread::sleeps; retry/repeat accept Count::Infinite; the only bounds are libcurl’s per-request timeouts (default 300 s), which do not bound total entry time under retries. The standard here is structural cancellation (CancellationToken threaded everywhere). Verified: tokio_util::sync::CancellationToken is runtime-agnostic (tokio sync primitives are documented runtime-independent; default-features = false — no tokio runtime enters the tree); is_cancelled() polling and child_token() work from plain threads.

Decision

Cancellation is cooperative at batch boundaries: the orchestrator checks a per-run CancellationToken (child token per scenario) before opening sessions and before each batch; EngineSession::run_batch receives the token so engines may honor it at finer grain when they can (a future non-hurl engine might; engine-hurl cannot mid-run_entries). Stuck-batch policy, layered: (1) the pack lint rejects infinite retries (retry: must carry a finite count) and unbounded repeat; (2) engine-hurl clamps per-request timeouts and computes a batch budget = Σ(entry timeout × (retries+1)) + retry intervals + margin; (3) a watchdog marks a scenario thread abandoned when its budget expires — the runner records a System failure with full context and detaches the thread (process exit reaps it) rather than blocking the run on an unjoinable thread. Ctrl-C: first signal cancels the token (graceful: finish current batches, run teardowns, write reports); second signal hard-exits.

Consequences

Bounded, explainable runs; no dependency on hurl gaining cancellation; a clean seam for engines that can do better. Costs: a cancel can wait out one in-flight batch (bounded by its budget); abandoned threads leak until process exit (accepted: the process is short-lived by design). If interrupt support ever lands upstream, adopting it is an engine-internal change (candidate for the ADR-0003 patch pipeline).

One consequence reaches a later feature, so it is recorded here rather than discovered again. [run] exclusive-tags promises a matching scenario the pool to itself, and the dispatcher enforces that against its own active set. An abandoned scenario leaves that set the moment the watchdog fires, while its detached thread keeps issuing requests until the entry in flight returns — the whole reason abandonment exists. So an exclusive scenario can start while an abandoned neighbour is still talking to the target: exactly the interference the key promises away, in the one window this ADR knowingly leaves open. It takes a budget blowout in the same moment, and a run that never trips the watchdog never meets it, so this is documented rather than engineered away — closing it means a bounded grace on the exclusive fill gate while detached threads report alive, which buys a rare guarantee with a per-exclusive-scenario delay every run pays. Revisit if the isolation guarantee ever has to be absolute.

Amendment (2026-09-07) — the budget family is closed over its inputs, and bounded as a product

The 0.18 survey found three holes in the value-cap regime, one of them the exact shape this ADR exists to prevent:

  • max-time: was read by the budget calculator and invisible to the lint. entry_timeout has always taken a literal max-time: as the entry’s timeout, while the option recogniser did not know the key — so [Options] max-time: 100000h was lint-clean and produced a multi-year batch budget the watchdog dutifully honoured. It now carries the duration cap like every budget input, and a test pins the rule the hole broke: every option the budget reads must be one the lint can see (every_budget_input_carries_a_value_rule).
  • retry-interval: multiplied into the budget with no value cap. It was in the recogniser for double-declaration purposes only; it now carries the duration cap.
  • Individually capped values compose into an unbounded product. retry: 10_000 (at the count cap) times a 30 s timeout is ~83 hours, lint-clean; saturated arithmetic reaches Duration::MAX, whose Instant + budget addition panics — contained by the dispatcher’s catch_unwind, but reported as a phantom “scenario thread panicked” system fault from a user-authored value. Two closures: the computed batch budget clamps to an absolute ceiling of four hours (MAX_BATCH_BUDGET — generous for any batch of API calls with finite retries, and a truthful watchdog abandonment for a runaway product), and the dispatcher’s deadline arithmetic uses checked_add with a far-future fallback, so an engine that ever hands core an unclamped budget degrades to “no deadline” rather than a panic.

In the same family: [http] timeout-ms = 0 was accepted and means no timeout to libcurl — the exact unbounded hang the default defends against, opted into by a value that reads like “immediately”. Refused as a user error now.

Alternatives considered

Killing scenario threads — unsound in Rust (no safe thread kill). Running each batch in a subprocess for killability — reintroduces the subprocess architecture ADR-0001 rejected, per-batch. Relying on timeouts alone — unbounded under retry loops (verified), and no graceful-report path on Ctrl-C.