Skip to content

qa(scores): RRI-25e55fa row (5.5/10, post-viewer-fix, +1.0 same-ruler) - #877

Merged
100yenadmin merged 1 commit into
mainfrom
data/rri-25e55fa-row
Jun 14, 2026
Merged

qa(scores): RRI-25e55fa row (5.5/10, post-viewer-fix, +1.0 same-ruler)#877
100yenadmin merged 1 commit into
mainfrom
data/rri-25e55fa-row

Conversation

@100yenadmin

@100yenadmin 100yenadmin commented Jun 14, 2026

Copy link
Copy Markdown
Member

Data-only. RRI 5.5/10 (6/11) vs RRI-5e98e6f 4.5 — same ruler, direct compare. zero_critical flipped to PASS (optimizer item-inspector crit gone). story 4.2 (within 0.1 of 4.3). Remaining FAILs: native_gate+image_render (measurement-lane, part-A deferred), cross_persona_sat 6.8, mech 2.6 (single-duo variance).

Summary by CodeRabbit

  • Documentation
    • Updated quality assurance scores ledger with a new run entry containing performance metrics and comparison analysis against the previous test run.

…, +1.0 vs RRI-5e98e6f)

zero_critical FLIPPED to PASS (optimizer item-inspector crit resolved in live play via
#872). All 5 personas zero-crit, 0 gave up, arc 5/5, behavioral GREEN, sat 6.6->6.8,
story 3.9->4.2 (within 0.1 of gate). mech 2.6 = single-duo angry-dm variance (smoke 3.9).
native_gate + image_render still FAIL = measurement-lane (Mac part-A deferred). Same ruler
as RRI-5e98e6f = direct compare. RRI arc this session: rc2 2.7 -> 4.5 -> 5.5.
@100yenadmin 100yenadmin added this to the v1.0.5 milestone Jun 14, 2026
@coderabbitai

coderabbitai Bot commented Jun 14, 2026

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 51771ae3-50a4-4aa7-b79d-fa008d9a6ace

📥 Commits

Reviewing files that changed from the base of the PR and between a002342 and e3343de.

⛔ Files ignored due to path filters (1)
  • qa/scores.db is excluded by !**/*.db
📒 Files selected for processing (1)
  • qa/scores_ledger.md

📝 Walkthrough

Walkthrough

One new scored run row (RRI-25e55fa) is prepended to the table in qa/scores_ledger.md. The ledger row count increments from 67 to 68, the rendered timestamp is updated, and the new row includes run metadata and comparison notes against the prior RRI-5e98e6f run.

QA Scores Ledger Row Addition

Layer / File(s) Summary
Add RRI-25e55fa ledger entry
qa/scores_ledger.md
Row count incremented (67 → 68), rendered timestamp updated, and new top-of-table row added with surface/model/scorer/ruler refs, RRI and lens values, pass/fail status, and notes comparing against RRI-5e98e6f.

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~2 minutes

Possibly related PRs

  • electricsheephq/WorldOS#747: Also adds a new top scored run row to qa/scores_ledger.md with updated row count and rendered timestamp — same ledger append workflow.
  • electricsheephq/WorldOS#724: Inserts a new auto-scored run row into qa/scores_ledger.md with updated metadata, directly the same pattern as this PR.
  • electricsheephq/WorldOS#413: Introduced the RRI scoring and scorecard-row tooling that generates the ledger records being appended here.

Poem

🐇 A new row hops in, fresh and bright,
RRI-25e55fa joins the ledger tonight.
Pass or fail, the scores are clear,
Sixty-eight entries now appear!
The rabbit stamps the timestamp true —
Another scored run, good as new. ✨

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description is incomplete. It lacks required sections: a 'Summary' section that describes what changed and why, and a 'Licensing/CLA' checklist. Only technical details are provided without proper structure. Add a 'Summary' section explaining the change clearly, and complete the 'Licensing/CLA' checklist with all required boxes and confirmations.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the main change: adding a new QA scores ledger row for RRI-25e55fa with specific metrics and comparison details.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands and usage tips.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e3343de0cb

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread qa/scores_ledger.md

| Run | When | SHA | Build date | Surface | DM model | Actor model | Scorer | Ruler | Lens ruler | RC | Methodology | Story | Mech | AngryDM | Behav | Sat | RRI | Crit | Img% | s/beat | cold-open s | turns/beat | Pass | Source | Notes |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| RRI-25e55fa | 2026-06-14T18:45:14Z | 25e55fa | 2026-06-14 | GUI-built-app | opus | | | sc_15c2436b4267 | | | gate-partA-skipped | 4.2 | 2.6 | | | 6.8 | 5.5 | 0 | 0% | 135 | 156.6 | 8.2 | FAIL | VM 5-persona part-B @25e55fa | Post-viewer-fix re-sweep (3 viewer PRs #872/#873/#875 on top of the 31 engine PRs). RRI 5.5/10 (6/11) vs run RRI-5e98e6f (4.5) — SAME ruler, DIRECT compare. zero_critical FLIPPED to PASS: optimizer's item-inspector crit (Studded Leather no-AC) resolved in live play (#872 renders the engine Item stats). ALL 5 personas zero-crit, 0 gave up, arc 5/5, behavioral GREEN, sat 6.6->6.8 (newbie 8/veteran 7/adversarial 6/narrative 8/optimizer 5). story 3.9->4.2 (within 0.1 of the 4.3 gate). mech 3.3->2.6 = single-duo angry-dm variance (the audit's own finding: one 6-beat duo is noise; smoke duo this SHA was 3.9) — needs a multi-duo median. STILL-FAIL: native_gate + image_render = measurement-lane gaps (Mac part-A handoff DEFERRED, Mac too disk/mem-constrained); cross_persona_sat 6.8 (<7, optimizer self-reports 5 surfacing the NEXT min-maxer depth layer: subclass-option completeness, class-feature rules inspector, weapon range field, hit-dice-spend UI). cold-open held 156.6s (latency win persists). |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Populate the behavioral result for this RRI row

The new row’s Behav cell is empty even though the same row’s notes state behavioral GREEN; the inserted qa/scores.db row also has behavioral set to NULL. When readers or qa/scores_db.py --compare --compare-rc-surface use the structured columns instead of the prose notes, this run incorrectly looks like it lacks behavioral-gate evidence, which undermines the 6/11 gate accounting and future trend comparisons.

Useful? React with 👍 / 👎.

@100yenadmin
100yenadmin merged commit f081e71 into main Jun 14, 2026
16 checks passed
@100yenadmin
100yenadmin deleted the data/rri-25e55fa-row branch June 14, 2026 18:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant