docs(skills): latency-forensics skill + ledger/model/flag/finish capture (last-few-days skillify) - #578
Conversation
… skill + ledger/model/flag/finish fixes Skillify backlog (5 HIGH items, from the wf_ed9a8ad2 discovery workflow, gap-checked): 1. NEW skill .claude/skills/worldos-latency-forensics — SDK duration_api_ms measurement method (true tool-exec ~1-4%, beats are generation-bound), quality-cost lever taxonomy (neutral: alwaysLoad/streaming/scene_context/prose-trim; trading: effort), REFUTED hypotheses (Haiku-helper, --tools, --fast). 2. worldos-dev SKILL.md: re-add the ⚠ Support-VM-lane callout (heavy sweeps, not local); ledger drift fix (canonical is scores_db.py -> scores_ledger.md, SCORECARD legacy); palette now 9-tool with finish() (finish=satisfied/gave_up=false vs give_up=blocked, #574); + DISCIPLINE: 'QA must exercise the flag' (run_duo ignored CLAWDND_LEAN_BEATS -> confounded A/Bs) and 'DM latency is reasoning not the GUI/harness'. 3. docs/MODEL-TIERING-STRATEGY.md: supersede the stale proposal with the MEASURED state (generation-bound; effort is the lever; ONE model/campaign, never switch mid-campaign=cache; Opus story 4.4-4.5 vs Sonnet 4.2). Model choice left OPEN (owner's call; models are options, not bolted-on); Haiku-helper REFUTED. 4. Ledger-drift banner on qa/SCORECARD.md (LEGACY -> scores_db.py) + qa/SCORING.md + qa/QA_TOOLS.md refs. Doc-only. The 5th HIGH item (verify-subagent-factual-claims) lands in the user-global verification-before-completion skill, separately.
|
Caution Review failedThe pull request is closed. ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (6)
📝 WalkthroughWalkthroughThis PR consolidates WorldOS QA operational discipline, scoring infrastructure, and measured DM performance strategy across six documentation files. It redirects the scoring source-of-truth from legacy narrative to SQLite-backed ledger, introduces latency measurement methodology, updates model-tiering strategy based on measured findings, and refines the GUI playtester harness semantics for satisfaction-based gating. ChangesQA Operations & Scoring Infrastructure
DM Latency Forensics & Model Strategy
Estimated code review effort🎯 2 (Simple) | ⏱️ ~10 minutes Possibly related PRs
Comment |
…rule (#13/#6) + REGRESSION-FORENSICS cross-link (#9) (#584) Adversarial re-review (post-VM-sweep, 3h later) graduated 3 backlog items to HIGH on new evidence: - #13 (LOW→HIGH): the scorer's DERIVED sat = 8-friction (crit=1⇒6) STRUCTURALLY can't clear G3 (≥7); the 2026-06-02 VM sweep had gave_up=false on all 5 but 4/5 DERIVED 5-6 → the low sat is a self-report-COVERAGE artifact, not a quality failure. Added the reading rule to worldos-dev (verify satisfaction_source; derived G3-miss = inconclusive) + soft caveat (finish() fires in the minority; 'self-reported' can be the Satisfaction:N/10 verdict line, not only finish()). - #6 (finish-it): new WorldOS-GUI-RUNBOOK 'Reading the sweep — honest satisfaction' section (finish vs give_up; budget=story-beats; spinner=progress; every persona MUST finish(satisfaction); PARTIAL≠product-RED). - #9 (MED→HIGH): cross-link qa/REGRESSION-FORENSICS.md from QA_TOOLS Evidence-Reading-Order (the 4.x-vs-2.x 'regression' is a surface+rubric artifact; was orphaned, no nav path). - Also fixed the residual SCORECARD ledger-drift in the GUI-RUNBOOK gate-sweep line. Doc-only. The 5 prior HIGH items (PR #578) were adversarially re-verified against a40beaf + the VM sweep: all sound, no contradictions (#579-582 were additive render/ + server.py SSE, no protected files touched). Co-authored-by: Eva <arncalso@gmail.com>
Captures the last-few-days workflows so a post-compaction agent finds them. 1 new skill (worldos-latency-forensics) + worldos-dev callout/ledger/palette/discipline updates + MODEL-TIERING measured-state rewrite (model choice left OPEN) + SCORECARD/SCORING/QA_TOOLS ledger-drift fixes. Doc-only; from the gap-checked skillify-discovery workflow.
Summary by CodeRabbit