Skip to content

Add baseline evaluation script for verification system - #10

Open
ashritha-03102004 wants to merge 1 commit into
interviewstreet:mainfrom
ashritha-03102004:patch-2
Open

Add baseline evaluation script for verification system#10
ashritha-03102004 wants to merge 1 commit into
interviewstreet:mainfrom
ashritha-03102004:patch-2

Conversation

@ashritha-03102004

Copy link
Copy Markdown

This script initiates the verification system baseline evaluation run, executes the evaluation pipeline, and saves the evaluation report.

This script initiates the verification system baseline evaluation run, executes the evaluation pipeline, and saves the evaluation report.
cbeaulieu-gt referenced this pull request in cbeaulieu-gt/hackathon-claims-agent Jul 6, 2026
Verified findings from the approach-critic review (#10):
- Validator: disagreement KEEPS the primary verdict + manual_review_required
  (was: force not_enough_information). Every user_history_risk sample row is a
  confident supported/contradicted, not NEI; forcing NEI corrupted 3 coupled
  fields (claim_status + evidence_standard_met + severity).
- manual_review_required: flag-class rule (high-stakes flags escalate; pure
  capture-quality flags do not). Counterexamples case_006, case_007.
- Call 1 observation made comprehensive (enumerate parts + note absent) so
  Call 2 can detect a claimed part is not shown (case_006).
- Zoom escalation cut (YAGNI; hosted APIs downscale-then-tile, so cropping
  upsamples blur, not detail).
- Eval adoption reframed to per-case error segmentation (n=20 ~plus/minus 20pp CI).

Persist full critique: docs/research/2026-06-19-approach-critique.md.
History reconciliation (FLAW 5) deferred for separate discussion.

Closes #11

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

2 participants