Skip to content

[QA·dogfood] Judgment-quality dogfood — does a real consider answer now carry a supported insight? #610

Description

@atomchung

Status

Ready — owner walk with agent support. The first live test of the judgment-quality layer that landed 2026-07-30.

#488's ruling deferred the next formal acceptance campaign "until the agent side has enough judgment quality to be worth re-walking." This issue is that test, deliberately lighter than a formal campaign: one real decision, scored honestly. It does not replace #488's future exact-SHA re-walk and must not be cited as it.

What changed since the last dogfood

The 2026-07-30 dogfood verdict on a live consider answer was: for-side stated the obvious, against-side was five boilerplate lines, 「整體來看,回答毫無意義」. Since then, on main@94b7e5f:

This walk answers: does a real answer now clear the bar those changes were built for?

The walk

  1. One real decision. The owner brings one trade they are genuinely considering (or most recently considered). The agent runs the full current flow: consider with --decision-context, L0 position packet, L1 lookup if triggered (with the motive-confirmation question), decision-first answer. Receipted via the card-free consider route where the environment allows (qa discipline per docs/qa-runbook.md if run as a formal QA campaign; a private real-book walk records the generic verdict only).
  2. Score the answer — the four checks (from the [design·M1] Decision-first TradeEvaluation answer — one supported judgment, not a disclosure dump #579 ruling; each pass/fail plus one sentence why):
    • Cover the limitation sentences — does the answer still stand on its own?
    • Does it say at least one thing the owner did not already know, or had not connected?
    • Does the counter-case directly engage the lead's strongest support (a real rebuttal, not a relabeled disclaimer)?
    • Is the owner's next step more concrete than before reading?
  3. Record live wall-clock segments ([design·perf] Response-time budget by route — separate real-user latency from QA/test overhead #603's first baseline — no backfill, per [QA·process] ux receipt trace was backfilled in one burst: all 10 events stamped within 1s, defeating latency instrumentation #278): the consider call, each lookup, total time to the delivered answer.
  4. Walk at least two market-lookup scenes from tests/agent/market-lookup-scenes.md as they naturally occur (scene 1 or 6 for the L1 path, scene 4 or 7 for the L0 boundary), and record pass/fail against the scene's own criteria.

What the results feed

Success condition

Not four green checks. Success = an honest per-check verdict on a real decision, timings recorded live, and every failure converted into a named next cut. A dishonest pass would poison #579's evidence-first loop.

Privacy

The real trade, book, motive, and answer stay local. Only the generic pass/fail structure, timings, and de-identified failure shapes are posted here (tools/privacy_lint.py before posting, per the QA runbook).

Refs #579, #601, #603, #488 (deferred formal acceptance), #475 (Phase 1 usage gate), #609 (unlocks on that gate).

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions