Skip to content

docs(skills): latency-forensics skill + ledger/model/flag/finish capture (last-few-days skillify) - #578

Merged
100yenadmin merged 1 commit into
mainfrom
docs/skillify-5high
Jun 2, 2026
Merged

docs(skills): latency-forensics skill + ledger/model/flag/finish capture (last-few-days skillify)#578
100yenadmin merged 1 commit into
mainfrom
docs/skillify-5high

Conversation

@100yenadmin

@100yenadmin 100yenadmin commented Jun 2, 2026

Copy link
Copy Markdown
Member

Captures the last-few-days workflows so a post-compaction agent finds them. 1 new skill (worldos-latency-forensics) + worldos-dev callout/ledger/palette/discipline updates + MODEL-TIERING measured-state rewrite (model choice left OPEN) + SCORECARD/SCORING/QA_TOOLS ledger-drift fixes. Doc-only; from the gap-checked skillify-discovery workflow.

Summary by CodeRabbit

  • Documentation
    • Updated QA operational guidance with enhanced HEAVY QA SWEEPS procedures and refined scoring/ledger specifications.
    • Added latency measurement methodology for WorldOS performance analysis.
    • Updated model tiering strategy to clarify campaign consistency rules, effort mechanics, and validation approach.
    • Restructured scoring ledger system from legacy narrative format to SQLite-backed database with rendered output.

… skill + ledger/model/flag/finish fixes

Skillify backlog (5 HIGH items, from the wf_ed9a8ad2 discovery workflow, gap-checked):
1. NEW skill .claude/skills/worldos-latency-forensics — SDK duration_api_ms measurement method (true tool-exec ~1-4%, beats are generation-bound), quality-cost lever taxonomy (neutral: alwaysLoad/streaming/scene_context/prose-trim; trading: effort), REFUTED hypotheses (Haiku-helper, --tools, --fast).
2. worldos-dev SKILL.md: re-add the ⚠ Support-VM-lane callout (heavy sweeps, not local); ledger drift fix (canonical is scores_db.py -> scores_ledger.md, SCORECARD legacy); palette now 9-tool with finish() (finish=satisfied/gave_up=false vs give_up=blocked, #574); + DISCIPLINE: 'QA must exercise the flag' (run_duo ignored CLAWDND_LEAN_BEATS -> confounded A/Bs) and 'DM latency is reasoning not the GUI/harness'.
3. docs/MODEL-TIERING-STRATEGY.md: supersede the stale proposal with the MEASURED state (generation-bound; effort is the lever; ONE model/campaign, never switch mid-campaign=cache; Opus story 4.4-4.5 vs Sonnet 4.2). Model choice left OPEN (owner's call; models are options, not bolted-on); Haiku-helper REFUTED.
4. Ledger-drift banner on qa/SCORECARD.md (LEGACY -> scores_db.py) + qa/SCORING.md + qa/QA_TOOLS.md refs.

Doc-only. The 5th HIGH item (verify-subagent-factual-claims) lands in the user-global verification-before-completion skill, separately.
@100yenadmin
100yenadmin merged commit 3e531cc into main Jun 2, 2026
3 of 4 checks passed
@coderabbitai

coderabbitai Bot commented Jun 2, 2026

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 610f8b5e-bd2e-40dc-917f-83089bf0fec7

📥 Commits

Reviewing files that changed from the base of the PR and between c867f31 and ce6bba7.

📒 Files selected for processing (6)
  • .claude/skills/worldos-dev/SKILL.md
  • .claude/skills/worldos-latency-forensics/SKILL.md
  • docs/MODEL-TIERING-STRATEGY.md
  • qa/QA_TOOLS.md
  • qa/SCORECARD.md
  • qa/SCORING.md

📝 Walkthrough

Walkthrough

This PR consolidates WorldOS QA operational discipline, scoring infrastructure, and measured DM performance strategy across six documentation files. It redirects the scoring source-of-truth from legacy narrative to SQLite-backed ledger, introduces latency measurement methodology, updates model-tiering strategy based on measured findings, and refines the GUI playtester harness semantics for satisfaction-based gating.

Changes

QA Operations & Scoring Infrastructure

Layer / File(s) Summary
Scoring ledger infrastructure redirection
qa/QA_TOOLS.md, qa/SCORECARD.md, qa/SCORING.md
Canonical scoring ledger migrated from legacy qa/SCORECARD.md to SQLite qa/scores_db.py rendered as qa/scores_ledger.md via add_run(...). SCORECARD.md marked narrative history; manual ledger edits deprecated.
QA operational discipline & heavy sweep guidance
.claude/skills/worldos-dev/SKILL.md
Heavy 5-persona QA sweep guidance added for support VM execution; scoring contamination markers clarified; discipline requirements for feature flag verification and DM beat latency interpretation inserted.
GUI playtester harness 9-tool palette & satisfaction semantics
.claude/skills/worldos-dev/SKILL.md
Playtester harness upgraded from 8-tool to 9-tool palette; finish(satisfaction, verdict) vs give_up semantics introduced; goal-gate now requires satisfaction ≥ 7 and gave_up=false.

DM Latency Forensics & Model Strategy

Layer / File(s) Summary
DM latency forensics measurement methodology
.claude/skills/worldos-latency-forensics/SKILL.md
New authoritative guide for measuring per-beat DM latency; duration_api_ms defined as generation-time baseline; tool+orchestration overhead computed as duration_ms - duration_api_ms; transcript gap misattribution pitfalls documented.
Latency mitigation levers & refuted approaches
.claude/skills/worldos-latency-forensics/SKILL.md
Lever taxonomy (neutral mitigations: alwaysLoad, compact digest, prose trimming; quality-tradeoff: --effort); refuted hypotheses (research-packet prefetch, non-existent --fast, tool-routing misconceptions); ordered experimental sequence from 2-beat probe through 5-persona gate.
Model-tiering strategy: measured findings & invariants
docs/MODEL-TIERING-STRATEGY.md
Reframed from proposal to measured status (dated 2026-06-02); DM latency dominated by generation/thinking and input mass, not tool execution; --effort tiering controls wall-clock time; mid-campaign model switches forbidden due to prompt-cache invalidation.
Refuted model ideas & validation ladder
docs/MODEL-TIERING-STRATEGY.md
Research-packet prefetch helper and headless --fast mode explicitly refuted; prior test sequence replaced with validation ladder (digest correctness → cache stability → flag wiring → short A/B → persona .app gate); open items updated (Opus model id, lower-effort quality threshold, scorer re-baselining).

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

  • electricsheephq/WorldOS#325: Playtester harness implementation and tooling surface for the 9-tool palette, give_up semantics, and satisfaction gating that this PR documents.
  • electricsheephq/WorldOS#333: Prior expansions of worldos-dev/SKILL.md harness discipline sections and heavy QA operational guidance that this PR refines.
  • electricsheephq/WorldOS#509: QA documentation updates to qa/QA_TOOLS.md and scorecard navigation, overlapping the ledger/source-of-truth redirection in this PR.

🐰 The scores are tallied, the tactics refined,
Nine tools in hand, one satisfied mind,
Swift beats measured, models aligned,
QA discipline crystallized, latency unwound.


Comment @coderabbitai help to get the list of available commands and usage tips.

100yenadmin added a commit that referenced this pull request Jun 2, 2026
…rule (#13/#6) + REGRESSION-FORENSICS cross-link (#9) (#584)

Adversarial re-review (post-VM-sweep, 3h later) graduated 3 backlog items to HIGH on new evidence:
- #13 (LOW→HIGH): the scorer's DERIVED sat = 8-friction (crit=1⇒6) STRUCTURALLY can't clear G3 (≥7); the
  2026-06-02 VM sweep had gave_up=false on all 5 but 4/5 DERIVED 5-6 → the low sat is a self-report-COVERAGE
  artifact, not a quality failure. Added the reading rule to worldos-dev (verify satisfaction_source;
  derived G3-miss = inconclusive) + soft caveat (finish() fires in the minority; 'self-reported' can be the
  Satisfaction:N/10 verdict line, not only finish()).
- #6 (finish-it): new WorldOS-GUI-RUNBOOK 'Reading the sweep — honest satisfaction' section (finish vs give_up;
  budget=story-beats; spinner=progress; every persona MUST finish(satisfaction); PARTIAL≠product-RED).
- #9 (MED→HIGH): cross-link qa/REGRESSION-FORENSICS.md from QA_TOOLS Evidence-Reading-Order (the 4.x-vs-2.x
  'regression' is a surface+rubric artifact; was orphaned, no nav path).
- Also fixed the residual SCORECARD ledger-drift in the GUI-RUNBOOK gate-sweep line.

Doc-only. The 5 prior HIGH items (PR #578) were adversarially re-verified against a40beaf + the VM sweep:
all sound, no contradictions (#579-582 were additive render/ + server.py SSE, no protected files touched).

Co-authored-by: Eva <arncalso@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant