Skip to content

0.4 — Author capability-error-explanation scenarios + run Eval 4 baseline #4

Description

@bb-connor

Phase 0 task 0.4 (PLAN.md, Phase 0 task 0.4).

Build the capability-error-explanation outcome eval (Eval 4, subjective).

Deliverables

  • Author 10 scenarios in chio-pack/eval/fixtures/cap-error-explanation/ per the example template.
  • Identify 3 raters + 1 alternate; record handles → pseudonyms in RATERS.md.
  • Run a Run-0 inter-rater calibration round; record scores in RATERS.md and per-rater adjustments in vault/_meta/dashboards/rater-calibration.md.
  • Implement chio-pack/chio_pack/eval/runners/cap_error_explanation.py per PHASE-0.md Eval 4.
  • Run baseline for the raw augmentation only (the others may improve in Phase 1+); commit to ADR-0002.

Done when

  • make kb-eval-outcomes reports capability-error-explanation as ok.
  • RATERS.md Run-0 calibration row is filled.
  • ADR-0002's row has the raw-augmentation baseline + sample size + date.

Reference

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions