qa(scores): mech median 3.8 (5 combat-sprints) — refutes single-duo 2.6 variance - #881
Conversation
…uo 2.6 variance 5 pre-seeded combat-sprints @25e55fa, all behavioral GREEN: [3.1,3.2,3.8,3.9,4.0] median 3.8. The RRI-25e55fa sweep's lone social duo scored angry-dm 2.6 (low combat coverage = variance, per the audit's own finding). The honest mechanical score is ~3.8 (still <4.5 gate).
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Plus Run ID: ⛔ Files ignored due to path filters (1)
📒 Files selected for processing (1)
📝 WalkthroughWalkthroughA single new row ( ChangesWorldOS Scores Ledger Entry
Estimated code review effort🎯 1 (Trivial) | ⏱️ ~2 minutes Possibly related PRs
Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Comment |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 5c5d51a26b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
|
||
| | Run | When | SHA | Build date | Surface | DM model | Actor model | Scorer | Ruler | Lens ruler | RC | Methodology | Story | Mech | AngryDM | Behav | Sat | RRI | Crit | Img% | s/beat | cold-open s | turns/beat | Pass | Source | Notes | | ||
| |---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---| | ||
| | mech-median-25e55fa-n5 | 2026-06-14T20:23:52Z | 25e55fa | 2026-06-14 | engine-duo | opus | | | sc_15c2436b4267 | | | combat-sprint-median-n5 | | 3.8 | | | | | | | | | | FAIL | VM combat-sprints csmed-1..5 @25e55fa | Multi-sprint mech MEDIAN to kill the single-duo variance (RRI-25e55fa's lone social duo scored angry-dm 2.6; the audit's own finding: one 6-beat duo is variance-not-signal). 5 combat-sprints (qa/run_combat_sprint.sh, pre-seeded combat = the right mechanical coverage), all behavioral GREEN: [3.1, 3.2, 3.8, 3.9, 4.0] -> MEDIAN 3.8. So the honest mechanical score is ~3.8 (still < the 4.5 gate, but +1.2 over the noisy 2.6). Combat-sprint surface = engine-duo (no GUI). Same ruler sc_15c2436b4267 as the RRI rows. | |
There was a problem hiding this comment.
Stamp the GREEN behavioral result on the median row
This row says the five combat sprints were "all behavioral GREEN", but the Behav cell (and the DB row's behavioral field) is empty. The QA tooling treats that column as the machine-readable gate signal (for example qa/scores_db.py --compare prints a blank beh for this newest current-ruler row), so consumers cannot distinguish this median from an un-gated score even though the evidence says it passed; stamp behavioral=GREEN in qa/scores.db and re-render.
Useful? React with 👍 / 👎.
Data-only. The single social duo in RRI-25e55fa scored angry-dm 2.6 (low combat coverage = variance). 5 combat-focused sprints give the honest median: [3.1,3.2,3.8,3.9,4.0] → 3.8, all GREEN. Mechanical is ~3.8 (still under the 4.5 gate), not 2.6.
Summary by CodeRabbit