diff --git a/.changeset/humble-webs-peel.md b/.changeset/humble-webs-peel.md new file mode 100644 index 000000000..02fa5a685 --- /dev/null +++ b/.changeset/humble-webs-peel.md @@ -0,0 +1,5 @@ +--- +"@hashintel/brunch": minor +--- + +Add compare greenfield execution diff --git a/memory/PLAN.md b/memory/PLAN.md index 3b61a28c1..a58636947 100644 --- a/memory/PLAN.md +++ b/memory/PLAN.md @@ -114,6 +114,7 @@ Older completion history (incl. FE-1196, FE-1195, FE-1190, and FE-1180): [`docs/ Everything executor/orchestrator-shaped or Execute-mode-owned belongs to Kostandin's stream and is **outside the LN quarantine**. Cross-stream touchpoints: FE-1187 rows O7/O8/O9 (live D120-L Execute workflows) — coordinate before building those rows. +- `execution-comparison-tracer` ([FE-1230](https://linear.app/hash/issue/FE-1230/greenfield-execution-comparison-tracer)) — **active on `ka/fe-1230-greenfield-execution-comparison-tracer`:** the Opus 4.8 Brunch elicitation run produced a human-approved, build-ready minimal browser Petri-net editor specification with whole-app mount, complete drag-lifecycle, Petri semantics, persistence, and malformed-import criteria. Next: freeze the durable case packet and controller-only oracle contract, then smoke Brunch Execute before any matched lane. - **Carved from FE-1167 (2026-07-13):** the Execute-mode evidence sub-list — Execute entry beats on thin vs rich seeds (assessment honesty: Ask on thin, Proceed on rich) and the FE-1107/KA residue (close-or-narrow, demo/walkthrough session via `TESTING_PLAN.md`, post-KA plan pass). The former sticky-posture question is no longer KA residue: FE-1187 `remediation-4` owns the persistent Specify elicitation-style audit/SPEC revision, and its Continue lexical audit owns the old `continue` ambiguity. Full context in the archived FE-1167 definition (`docs/archive/PLAN_HISTORY.md`). - `planning-process-model` — **moved to the KA stream 2026-07-13; reshaped by D126-L**: the durable scope handoff is settled, so this item now owns only plan projection and epistemic-horizon questions beyond committed scopes. Definition below. - `executor-slice-attempt-lifecycle` ([FE-1192](https://linear.app/hash/issue/FE-1192/executor-slice-attempt-lifecycle)) — **active in the KA lane, picked up 2026-07-13** on `ka/fe-1192-executor-slice-attempt-lifecycle`. First member of the Petri sequence: first-class slice attempts (identity, bounded in-run retry, honest failed-attempt facts). Shape settled at pickup: attempt facts first (topology unchanged), constant retry bound with `ceiling:`, agent step only. Definition below. @@ -386,6 +387,24 @@ Instrumentation experiments and far-horizon items. Each re-enters only via re-qu - **Traceability:** D126-L (settled scope handoff), D103-L (durable slice retirement), D100-L (`project` seam), D87-L (`unknown` = horizon on the intent plane), D99-L (advisory/settled); KA owns executor/orchestration consequences. SPEC §Future Direction "Planning persistence evolution". +### execution-comparison-tracer + +- **Name:** Greenfield execution comparison tracer +- **Linear:** [FE-1230](https://linear.app/hash/issue/FE-1230/greenfield-execution-comparison-tracer), child of [FE-1211](https://linear.app/hash/issue/FE-1211/brunch-testing-execution-side-evaluation-of-outputs). +- **Branch:** `ka/fe-1230-greenfield-execution-comparison-tracer` (off `next`). +- **Kind:** bounded evaluation tracer — frozen execution input, isolated adapters, controller-only oracles, and immutable evidence; no product operator command yet. +- **Certainty:** proving. +- **Status:** active. The saved `minimal-petri-net-editor` mission was rerun through Brunch Specify with `anthropic/claude-opus-4-8`; the exported specification was enriched after review and explicitly approved as the common execution input. The direct run used the documented headless TUI fallback after the approachable `/compare-specs` controller stalled before harness launch; this input-generation run is not comparison evidence. +- **Objective:** compare Brunch Execute with Claude Code from identical empty repositories using one human-approved browser-app specification and the same pinned Opus 4.8 model, while separating mechanical outcome evidence from unblinded process evidence and stopping Brunch at `promotion_prepared`. +- **Lights up:** approved elicitation output → frozen lane-neutral execution packet → isolated Brunch/Claude implementations → unchanged hidden browser/Petri oracles → masked outcome + unblinded process adjudication. +- **Stabilizes:** the first execution-case artifact contract, validity/intervention ledger, no-landing boundary, and reproducibility split between generated-plan/Petri procedure and structural code output. +- **Acceptance:** the approved Petri-net editor spec and controller-only oracle pack freeze before a valid lane; one Brunch-only empty-dir smoke reaches `promotion_prepared` without host landing; one matched Brunch-vs-Claude run uses identical public input and fresh repositories; hidden checks run unchanged against both outputs; three valid Brunch runs report canonicalized generated-plan/transition/action similarity separately from code-structure similarity; failed/invalid attempts remain retained; only a reviewed portable bundle promotes. +- **Verification:** inner — hashes/schema checks for the public packet, case manifest, validity ledger, and hidden-oracle installation; middle — real isolated adapter smoke plus unchanged build/test/headless-browser/Petri checks against each lane; outer — named human review of identity-masked outcome and unblinded process packets. Permission/intervention metrics count observed events only; unavailable signals are `not assessable`. +- **Cross-cutting obligations:** reuse FE-1210's split-judgment, failure-retention, masking, and promotion discipline without mixing execution cases into the private elicitation-mission namespace; record exact provider/model/harness versions; never expose controller-only oracles to a lane; never invoke `/brunch:land`. +- **Explicitly out:** Pi Campaign Machine/Clay, brownfield Brunch/Petrinaut cases, automatic landing, Cursor/Codex lanes, production `/compare-execution`, and broad reliability/cost/speed claims from one mission. +- **Traceability:** D40-L, D120-L, I62-L; FE-1210/FE-1215 comparison evidence discipline; [`testing/execution-comparisons/cases/minimal-petri-net-editor/spec.md`](../testing/execution-comparisons/cases/minimal-petri-net-editor/spec.md); origin mission [`testing/comparisons/missions/minimal-petri-net-editor.md`](../testing/comparisons/missions/minimal-petri-net-editor.md); `docs/praxis/comparison-runs.md`; `src/executor/TOPOLOGY.md`. +- **Current execution pointer:** [`memory/cards/execution-comparison-tracer--brunch-oracle-smoke.md`](cards/execution-comparison-tracer--brunch-oracle-smoke.md). + ### executor-slice-attempt-lifecycle - **Name:** Slice attempt lifecycle — first-class attempts in the executor net @@ -557,6 +576,11 @@ group-4 (cleanups): rides group-1 stack | named-inline-extension-identity (P1) KA stream: carved FE-1167 Execute beats + FE-1107 residue planning-process-model (moved 2026-07-13) + execution-comparison-tracer (FE-1230) + status: active; approved Opus 4.8 Petri-editor spec ready to freeze as a durable case + reuses: FE-1210 split judgment | failure retention | masked outcome discipline + lights_up: frozen spec -> Brunch/Claude isolated lanes -> hidden browser/Petri oracles -> adjudication + excludes: host landing | product operator command | Cursor/Codex | broad benchmark claims # petri-interpreter-port (FE-1183) and petrinaut-live-run-stream (FE-1190) merged; # run.json remains lifecycle truth, Petri artifacts remain projection/evidence/resume hints. executor-slice-attempt-lifecycle (FE-1192) diff --git a/memory/SPEC.md b/memory/SPEC.md index 721eec576..0ae5b3126 100644 --- a/memory/SPEC.md +++ b/memory/SPEC.md @@ -686,6 +686,8 @@ For agent-as-user evaluation, the primary behavioral claim is **consequential-fa **FE-1187 A48-L readiness-preflight assessment (2026-07-17).** Observability is **partial** until the tracer lands: menu output and elapsed time are visible, but semantic comparison needs a structured per-attempt report carrying scenario/key, deterministic and model flags, elapsed/cache/fallback outcome, evaluator stamp/rationale, and transcript/graph/session before/after deltas. Reproducibility is **partial**: a tracked human-approved catalog fixes seven exact graph-state labels plus one injected reconciliation-blocker rival, while authenticated latency still varies. Controllability is **partial**: schema, timeout, fallback, cache, and side-effect rivals are autonomous; the configured soft recommended evaluator model and network remain external. Admission uses three uncached feasibility calls before a ten-run campaign; ≥8/10 must complete within 3 seconds, every completed result must match exact flags, the model must correct at least one named fallback miss, and it may add no false-positive move. +**FE-1230 greenfield execution-comparison assessment (2026-07-20).** Observability is **partial**: build/test output, git trees, browser DOM/accessibility state, console errors, JSON downloads, Brunch run/Petri artifacts, and target-visible interactions are text-native, while ordinary visual hierarchy and drag feel remain human judgments; product-private reasoning and diagnostics are excluded from common evidence. Reproducibility is **partial**: the approved Petri-editor specification, empty-repository base, public automation contract, Opus 4.8 model, budgets, and hidden oracle bytes freeze before the first valid lane, but generated implementation shape and model conduct vary; three Brunch repetitions detect gross instability rather than reliability tails. Controllability is **partial**: the controller can install and run unchanged build, browser, reference-model, round-trip, and negative-space checks against fresh lane outputs, while provider availability and product-native planning remain external. A minimal public accessibility contract gives both lanes stable roles/names without revealing hidden scenarios or expected results. + **Standalone-web tracer assessment (2026-07-14).** Observability is **partial**: JSONL, target-addressed RPC, semantic event frames, and accessible DOM/text states make structural session behavior observable; stream cadence, visual hierarchy, and error feel remain manual. Reproducibility is **high** for this tracer: paired temporary production boots use the deterministic faux provider and an existing JSONL session, not a live model or a static transcript golden. Controllability is **partial**: the middle loop controls the browser journey and cache/overlay loss, while a workbench walkthrough owns the bounded UX verdict. The normalizer may remove only declared nondeterministic ids/timestamps; it must not mask Brunch binding/runtime/exchange differences. **Combined trajectory/evaluation assessment.** The completed interactive TUI driver raises presentation-path controllability but does not by itself close causal attribution, real-provider reproducibility, or semantic ground truth. The first evaluation frontier should address those gaps with a minimal text-native legibility envelope over existing Pi lifecycle events, provider-payload introspection, Pi JSONL, TUI observations, and graph readback — not with OTel adoption. Its run/scenario/report contract must be product-neutral enough for later Claude Code or Cursor adapters, while Brunch-only trajectory enrichment remains optional diagnostic evidence. Campaigns are automated through a controlled user actor; deterministic checks own structural facts, and a small human-labeled set calibrates and audits semantic judgments. @@ -760,6 +762,11 @@ Dev-loop artifacts route to gitignored `.fixtures/scratch///`, res | Middle | **Subagent-reconciliation oracle battery (`subagent-reconciliation`)** | Four deterministic faux-substrate oracles for the foreground/background agent reconciliation. … | | Outer | Manual walkthrough with checklist | UX/presentation life: TUI chrome, spec/session picker, web shell feel, coherence visibility, elicitation usefulness. Adds: ambient-affordance rendering from establishment-offer structured-exchange facets; … | | Outer | **Saved-mission comparison witness** (`saved-mission-comparison-witness`) | After FE-1215's D134-L remediation, a real stock-Pi `/compare-specs` Brunch + Claude run proves one top-level simulated-user actor can drive one direct, normal-width harness shell at a time; plain-text choices/approvals work without `ask_user_question`; private-mission isolation, target-visible evidence, cleanup, aggregate notification, and report usefulness remain visible. Revising and rerunning the saved mission then proves prior run snapshots remain byte-stable. No surrogate or simulated gate substitutes for these entry-point witnesses. | +| Inner | **FE-1230 case-contract validation** | Public specification, accessibility/serve contract, lane manifest, budgets, validity rules, and controller-only oracle manifest are schema-valid, content-addressed before lane launch, and path-separated so hidden material cannot enter a target cwd. | +| Middle | **FE-1230 black-box browser + accessibility journey** | The built `dist/` entry mounts without uncaught errors; public role/name contracts locate controls in both implementations; creation, full pointer drag release, rename/delete, arc creation/rejection, value validation, reset, reload, JSON round-trip, malformed import, and new/clear pass under one unchanged controller-owned suite. | +| Middle | **FE-1230 reference Petri model differential + metamorphic checks** | A tiny controller-owned P/T reference model independently computes enablement and weighted firing; the rendered marking/enablement agrees before and after firing, reset is idempotent, reload returns to the initial marking, export/import preserves structure, and disabled firing plus invalid operations leave state unchanged. | +| Middle | **FE-1230 lane/process evidence contract** | Fresh repository identity, pinned model/product versions, budgets, start/end state, final git range, build/test/browser output, retries, interventions, terminal status, and cleanup are retained for every valid, failed, and invalid attempt. Only signals shared across lanes enter comparison packets; unavailable permission/cost/private-tool evidence is `not assessable`. | +| Outer | **FE-1230 split execution judgment** | Identity-masked final trees/diffs plus common mechanical results receive criterion-level code/outcome review; a separate unblinded packet judges visible planning/execution conduct. Visual review mechanically rejects only catastrophic unusability (app absent, controls unreachable, interactions impossible); ordinary hierarchy, clarity, and feel remain qualitative and cannot override a mechanical failure. | | Middle | **FE-1187 controlled provider conduct gate — paused** | On explicit re-entry, reconcile the extractor/oracle against the landed mixed-settlement contract, then run three fresh normalized-ingest samples. All must use free-text digest feedback, one bounded questionnaire when several questions exist, no combinatorial options, an honestly assigned per-node/per-edge review proposal when review is used, exact atomic settlement preservation, and no post-review mutation that completes or rewrites the approved proposal. Direct advisory mutation without review remains valid. Reports retain provider/model stamps and deterministic conduct markers; the stopped 2026-07-17 run is diagnostic and counts 0/3. | | Inner | **FE-1187 Impact Ledger golden + word-wrap-tolerant render-honesty** | Golden/inline snapshots at narrow/normal/wide widths lock populated-section-only canonical order, absence of empty heading/`None` pairs, per-node/per-edge settlement visibility, elision, `refs:` row shape, and the `obligation` fallback label for the borderless Impact Ledger renderer (D27-L/D131-L). `missingRenderedDetailsLeaves` is extended to reassemble `table`'s word-wrapped physical lines back into logical cells before leaf-presence checking, so a value silently split across wrapped rows cannot pass as "rendered" by accident. | | Middle | **FE-1187 Impact Ledger differential reference extractor** | A deliberately naive reference extractor (flat node/edge/term inventory, no styling or grouping) is compared against the ledger's code/connection inventory over the witnessed fixture plus hand-authored edge fixtures (empty group, single-node group, mixed-settlement group, term-only group, max-refs group). Proves item/status inventory completeness and empty-group omission independent of the "real" renderer's own logic. | @@ -834,6 +841,7 @@ The first required probe is M0: after manual TUI interaction, a checker proves ` ### Design Notes - **Operator-led cross-product comparisons (FE-1215; D134-L remediation before later `saved-mission-comparison-witness` evidence).** The PM-facing comparison door is deliberately distinct from rigorous frozen-packet evaluation. One project Pi prompt conversationally creates or revises a rich private **agent-as-user mission** for a simulated user: their objective, context, priorities, preferences, constraints, knowledge, uncertainty, decision latitude, and conversational posture. The invoking top-level Pi agent receives that mission and directly performs the user's side of each interaction while driving exactly one comparison-harness subshell at a time; it must not spawn a Pi actor that opens another interactive shell. Each comparison harness receives only minimal visible framing plus the opening user message and subsequent mission-grounded answers. Harness selection and framing are run setup, not mission content. Ordinary conversational text is the portable baseline for operator choices and approvals; environment-specific structured-question tools are optional presentation only. Setup checks cover actual selected-harness prerequisites without synthetic actor/provider turns on every run. Editable missions live outside `.fixtures/` under `testing/comparisons/missions/`; each run snapshots the private mission, separately identified target-visible setup/interactions, and outputs under `.fixtures/runs/agent-as-user-comparison/` while temporary lane work stays in scratch. The readable operator report may expose the full private mission so elicitation can be compared against what each harness actually learned, but it keeps that baseline separate from target-visible evidence and declares no automatic winner or prescribed rubric. Because the top-level actor context spans sequential harnesses, this approachable workflow discloses order and does not claim the per-lane actor-process isolation required by rigorous frozen-packet studies; frozen reveal policies, matched budgets, blinding, fresh-per-lane actor sessions, structured adjudication, multi-run statistics, and scripted judges remain separate tools for focused improvement/regression claims. +- **FE-1230 execution-comparison oracle boundary.** Execution cases are distinct from private elicitation missions: the human-approved specification and a minimal public runtime/accessibility contract are visible to every lane, while exact browser journeys, reference-model states, expected results, label mapping, and adversarial fixtures remain controller-only and outside every lane cwd. The public contract requires a static production build at `dist/`, `npm run build`, `npm test`, and stable accessible roles/names for the canvas and named controls; it does not prescribe framework, source topology, implementation decomposition, test library, or internal state model. One controller-owned suite and byte-identical oracle pack run after each lane, never executor-authored tests alone. Outcome judgment sees identity-masked trees/diffs and common gate results; process judgment is separately unblinded and uses only normalized visible evidence. Brunch stops at `promotion_prepared` and never lands. The first tracer pins `anthropic/claude-opus-4-8` in both products, but same model does not imply equivalent hidden prompting or thinking controls. - **Shared-session-host convergence oracle (planned; `shared-session-host-tracer` → `shared-session-host-cutover`).** FE-1200's standalone proof is necessary but not sufficient for architectural replacement. The proving oracle must launch the production host as the sole owner of one writable sealed Pi runtime, attach a real TUI presentation and React observer/driver to the same durable target, exercise an ordinary turn plus one extension-owned structured ask and one TUI-only product interaction, detach/restart a client without ending the hosted runtime, and reject a duplicate writable open or conflicting driver. Durable settlement must converge through the same JSONL presentation, and browser traffic must remain Brunch semantic RPC/events rather than raw Pi RPC/events. Once A47-L retires, the cutover becomes a closed coverage sweep: every required TUI/web lifecycle, command/UI, exchange, transcript, graph-update, model/auth, and shutdown row has one host-owned path and a closure oracle; only then may `SessionEventRelay`, `brunch.sessionEvent`, `/rpc/driver`, and their sidecar harnesses be deleted. Human outer evidence must confirm that the TUI remains useful rather than becoming a thin degraded shell. - **Standalone-web compound oracle (2026-07-14; coverage completed 2026-07-15, `standalone-web-session-host`).** The initial five complementary oracles prove the host tracer: (1) inner RPC/host negative-space contracts require `(specId, sessionId)` on every lifecycle/driver/ask/event path and reject targetless fallback, duplicate writable opens, second drivers, and mismatched targets; (2) projection shape/malformed-detail tests keep JSONL and raw Pi/detail shapes behind the named semantic presentation; (3) a production-wired standalone-web browser journey, controlled only by a deterministic faux provider, asserts accessible hydration, streamed text, one `ask`, answer, and `agent_settled` states; (4) paired temporary web/TUI runs use a narrow declared normalizer to prove both live→settled→fresh JSONL hydration and equivalent Brunch binding/runtime/exchange semantics—no static JSONL golden; (5) a workbench manual checklist judges only stream cadence, busy/settled truthfulness, ask interaction, reload, and error feel. Browser cache/overlay loss is part of the middle-loop journey. The follow-on production-host concurrency differential retires A43-L with overlapping graph writes, asks, failure/recovery, target-local event sequences, reconnect, and separate JSONL readback; shared graph changes appear only through canonical `worldUpdate` continuity. The completed coverage pass adds projection no-loss/malformed tests for every required persisted ask terminal shape (including questionnaire read-back), React render/answer tests for free text and listed single/multi choices, headless schema-envelope questionnaire answering coverage (with no dedicated React questionnaire form), distinct candidate/review-set/digest production settlement/reconnect witnesses, concurrency/target isolation, and receipt-bearing review settlement. Live-provider conduct and process-restart survival remain explicitly deferred to their named later work; no second truth plane or raw Pi browser contract is introduced. - **Trace → eval → score → regression flywheel.** The first combined trajectory/evaluation proof is a controlled Brunch A/B over the existing consequential-fact claim, not a generic tracing platform or omnibus architecture score. One realistic non-inferable scenario carries a human-authored hidden-fact ledger, forbidden rivals, and a controlled reveal policy. The only intervention is an eval/dev ablation of the warrant-before-commit directive at the real prompt-composition seam; all other run conditions remain fixed. Three real-provider TUI-driven runs per arm produce a joined legibility envelope: stamped run configuration, directive inventory and content hashes, advertised/read/provider-visible directive state, ordered model/tool/exchange/TUI/graph effects, atomic evaluator judgments with evidence and rubric identity, and a replayable report. Completion requires promoted evidence that discriminates the full directive from the ablated rival; the report is an instrument, not the completion claim. This proves bounded evaluator discrimination, not competitor superiority or broad prompt quality. After the tracer, mine one real walkthrough failure into the same corpus. Keep OTel, broad subagent span joining, provider matrices, interaction-quality scoring, and competitor campaigns trigger-gated until a named claim requires them. @@ -863,6 +871,11 @@ The first required probe is M0: after manual TUI interaction, a checker proves ` | Blind spot | Reason | Mitigation | | ---------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| FE-1230 public automation-contract steering | Stable accessible roles/names and a fixed `dist/` build output remove selector/runtime ambiguity but slightly constrain both implementations. | Keep the contract minimal and product-neutral; hide scenarios, fixtures, expected results, and reference-model states; judge framework, architecture, decomposition, and ordinary UX as unconstrained output. | +| FE-1230 source-style leakage under masking | Product vocabulary, code organization, or generated prose may reveal a lane even after names and metadata are removed. | Treat labels as identity masking, not style anonymity; record suspected source cues as uncertainty and exclude them from criterion evidence. | +| FE-1230 incomparable private process signals | Token accounting, permission prompts, hidden reasoning, and private tool traces differ by product or may be unavailable. | Compare only target-visible/common evidence; retain Brunch-only diagnostics outside judgment packets; mark unavailable cost/permission/tool metrics `not assessable` rather than fabricating parity. | +| FE-1230 three-run reliability ceiling | Three Brunch repetitions can reveal catastrophic instability but cannot estimate tail reliability or generalize beyond one browser-app case. | State the ceiling in every report; make no broad speed/cost/reliability claim; add cases or repetitions only after the tracer proves the artifact/oracle contract and a named decision requires stronger evidence. | +| FE-1230 same-model non-equivalence | Pinning Opus 4.8 does not equalize product-owned system prompts, tool policy, context management, thinking controls, or provider wrappers. | Frame the result as a same-base-model workflow/product comparison, record exact product/provider/model/configuration versions, and keep process interpretation separate from mechanical outcome correctness. | | Full TUI automation | Cost exceeds value before the product state seams are proven, but startup-switcher regressions need a stronger visual signal than store-only checks. | Manual checklist plus artifact/query probe oracle; for FE-744 startup, add pty/ANSI-stripped capture assertions for the pre-Pi decision surface and absence of stale transcript bef … | | LLM elicitation quality and interaction flow | No stable deterministic ground truth for “good interview” early in the POC, and retired M1 scripted exchanges encoded only a thin obsolete exchange model. | Transcript-backed probe runs, human-reviewed probe reports, adversarial probe scenarios, expected structural coverage, … | | Subscription reconnect/resume | POC can prove initial state payload + live update without hardening network recovery yet. | Contract tests for initial state payload and ordered update sequence; **(2026-06-15)** reconnect/resume promoted to a `web-driver-streaming` battery claim — a turn-cut-point prope … | diff --git a/memory/cards/execution-comparison-tracer--brunch-oracle-smoke.md b/memory/cards/execution-comparison-tracer--brunch-oracle-smoke.md new file mode 100644 index 000000000..4bc9e6cfa --- /dev/null +++ b/memory/cards/execution-comparison-tracer--brunch-oracle-smoke.md @@ -0,0 +1,179 @@ +# Brunch execution-oracle smoke + +Frontier: execution-comparison-tracer +Status: active +Mode: single +Created: 2026-07-20 + +## Orientation + +- Containing seam: controller-owned execution evaluation over Brunch's existing `empty_dir` plan/run/Petri/promotion pipeline; the comparison harness is dev/evaluation tooling, not product runtime. +- Frontier: `execution-comparison-tracer` (FE-1230), child of FE-1211 and independent of the private elicitation-mission namespace. +- Volatile state: the Opus 4.8 Brunch Specify run produced a human-approved build-ready Petri-editor specification, now frozen at `testing/execution-comparisons/cases/minimal-petri-net-editor/spec.md`; the initial `/compare-specs` controller stalled before harness launch, so that input-generation attempt is diagnostic and not comparison evidence. +- Main open risk: an implementation-neutral hidden browser suite cannot drive independently generated UIs without leaking exact tests; the approved resolution is a minimal public runtime/accessibility contract plus hidden scenarios and reference-model expectations. + +Posture: proving (inherited from `execution-comparison-tracer`). + +## Target Behavior + +The approved Petri-editor case drives one isolated Brunch Execute run to `promotion_prepared` and evaluates its output with the frozen controller-only oracle pack without host landing. + +## Cold-start reads + +- `memory/SPEC.md` — D40-L, D120-L, I62-L; Verification Design sections for FE-1230 diagnostic, loop-tier strategy, design boundary, and blind spots +- `memory/PLAN.md` — frontier: `execution-comparison-tracer` +- `testing/execution-comparisons/cases/minimal-petri-net-editor/spec.md` — approved lane-neutral product specification +- `testing/comparisons/missions/minimal-petri-net-editor.md` — private elicitation origin only; never copy into an execution lane +- `docs/praxis/comparison-runs.md`, `comparison-runs/mission-packet.md`, and `comparison-runs/judgment-prompt-pack.md` — validity, retention, masking, split-judgment, and promotion discipline +- `src/dev/TOPOLOGY.md` — dev/evaluation harness ownership +- `src/executor/TOPOLOGY.md` — `empty_dir`, frozen plan, Petri journal, promotion, and no-landing boundaries +- `src/app/TOPOLOGY.md` — real planner/agent/test/promotion port composition +- `docs/praxis/manual-testing.md` — browser/TUI control and cleanup rules + +## Public case contract + +Only the following material enters the lane: + +```text +case: + id: minimal-petri-net-editor-v1 + specification: spec.md + model: anthropic/claude-opus-4-8 + repository: fresh empty git repository at a frozen empty base commit + elapsed_budget_minutes: 90 + mechanical_intervention_budget: 2 + substantive_human_interventions: 0 + +delivery: + npm_test: required + npm_build: required + static_output: dist/ + runtime_network: forbidden + +accessible_names: + application: Petri net editor + canvas: Petri net canvas + controls: + - Add place + - Add transition + - New net + - Reset marking + - Export JSON + - Import JSON + dynamic: + - "Place: