Skip to content

qa(scores): mech median 3.8 (5 combat-sprints) — refutes single-duo 2.6 variance - #881

Merged
100yenadmin merged 1 commit into
mainfrom
data/mech-median-row
Jun 14, 2026
Merged

qa(scores): mech median 3.8 (5 combat-sprints) — refutes single-duo 2.6 variance#881
100yenadmin merged 1 commit into
mainfrom
data/mech-median-row

Conversation

@100yenadmin

@100yenadmin 100yenadmin commented Jun 14, 2026

Copy link
Copy Markdown
Member

Data-only. The single social duo in RRI-25e55fa scored angry-dm 2.6 (low combat coverage = variance). 5 combat-focused sprints give the honest median: [3.1,3.2,3.8,3.9,4.0] → 3.8, all GREEN. Mechanical is ~3.8 (still under the 4.5 gate), not 2.6.

Summary by CodeRabbit

  • Chores
    • Updated internal quality assurance records with the latest test execution results and performance metrics.

…uo 2.6 variance

5 pre-seeded combat-sprints @25e55fa, all behavioral GREEN: [3.1,3.2,3.8,3.9,4.0] median 3.8.
The RRI-25e55fa sweep's lone social duo scored angry-dm 2.6 (low combat coverage = variance,
per the audit's own finding). The honest mechanical score is ~3.8 (still <4.5 gate).
@100yenadmin 100yenadmin added this to the v1.0.5 milestone Jun 14, 2026
@coderabbitai

coderabbitai Bot commented Jun 14, 2026

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: c9263bfa-697e-429a-90ca-55ed1aee4dc3

📥 Commits

Reviewing files that changed from the base of the PR and between 4b6dcd4 and 5c5d51a.

⛔ Files ignored due to path filters (1)
  • qa/scores.db is excluded by !**/*.db
📒 Files selected for processing (1)
  • qa/scores_ledger.md

📝 Walkthrough

Walkthrough

A single new row (mech-median-25e55fa-n5) is appended to qa/scores_ledger.md for an engine-duo combat-sprint median run on SHA 25e55fa, ruled by sc_15c2436b4267, with a Mech score of 3.8 and a FAIL outcome. The ledger row count is incremented from 68 to 69 and the render timestamp is refreshed.

Changes

WorldOS Scores Ledger Entry

Layer / File(s) Summary
New scored run row and header metadata
qa/scores_ledger.md
Row count updated from 68 to 69, rendered timestamp refreshed, and run mech-median-25e55fa-n5 appended with engine-duo surface, SHA 25e55fa, ruler sc_15c2436b4267, Mech 3.8, FAIL outcome, and notes covering median-combat-sprint rationale and variance vs prior lone-duo AngryDM finding.

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~2 minutes

Possibly related PRs

  • electricsheephq/WorldOS#747: Same file (qa/scores_ledger.md), same pattern of appending a new ruler-stamped scored run row while updating row count and rendered timestamp.
  • electricsheephq/WorldOS#877: Also adds a ledger entry for the same SHA 25e55fa, making the two PRs directly related as companion scored-run additions for the same build.

Poem

🐇 A new row appears in the ledger so bright,
SHA 25e55fa joins the fight!
Mech scores three-point-eight, the verdict: FAIL,
But behavioral checks sailed without a hail.
The rabbit stamps the count to sixty-nine,
And hops away — the timestamp's right on time! 🕐

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The PR description explains what changed and why (data-only update with rationale), but it omits required Licensing/CLA section and Validation checklist from the template. Add the Licensing/CLA section with required checkboxes and a Validation section listing the checks performed.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title accurately summarizes the main change: updating mechanical median score to 3.8 based on 5 combat-sprints, with a note about refuting the prior 2.6 variance.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands and usage tips.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 5c5d51a26b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread qa/scores_ledger.md

| Run | When | SHA | Build date | Surface | DM model | Actor model | Scorer | Ruler | Lens ruler | RC | Methodology | Story | Mech | AngryDM | Behav | Sat | RRI | Crit | Img% | s/beat | cold-open s | turns/beat | Pass | Source | Notes |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| mech-median-25e55fa-n5 | 2026-06-14T20:23:52Z | 25e55fa | 2026-06-14 | engine-duo | opus | | | sc_15c2436b4267 | | | combat-sprint-median-n5 | | 3.8 | | | | | | | | | | FAIL | VM combat-sprints csmed-1..5 @25e55fa | Multi-sprint mech MEDIAN to kill the single-duo variance (RRI-25e55fa's lone social duo scored angry-dm 2.6; the audit's own finding: one 6-beat duo is variance-not-signal). 5 combat-sprints (qa/run_combat_sprint.sh, pre-seeded combat = the right mechanical coverage), all behavioral GREEN: [3.1, 3.2, 3.8, 3.9, 4.0] -> MEDIAN 3.8. So the honest mechanical score is ~3.8 (still < the 4.5 gate, but +1.2 over the noisy 2.6). Combat-sprint surface = engine-duo (no GUI). Same ruler sc_15c2436b4267 as the RRI rows. |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Stamp the GREEN behavioral result on the median row

This row says the five combat sprints were "all behavioral GREEN", but the Behav cell (and the DB row's behavioral field) is empty. The QA tooling treats that column as the machine-readable gate signal (for example qa/scores_db.py --compare prints a blank beh for this newest current-ruler row), so consumers cannot distinguish this median from an un-gated score even though the evidence says it passed; stamp behavioral=GREEN in qa/scores.db and re-render.

Useful? React with 👍 / 👎.

@100yenadmin
100yenadmin merged commit a25e91d into main Jun 14, 2026
16 checks passed
@100yenadmin
100yenadmin deleted the data/mech-median-row branch June 14, 2026 20:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant