Skip to content

Reconcile results across README / docs/RESULTS.md / site: make RESULTS.md the superset, publish the 87-task runs everywhere, remove 8 unsubstantiated numbers - #241

Open
OsherElhadad wants to merge 2 commits into
mainfrom
fix/issue-100-reconcile-results
Open

Reconcile results across README / docs/RESULTS.md / site: make RESULTS.md the superset, publish the 87-task runs everywhere, remove 8 unsubstantiated numbers#241
OsherElhadad wants to merge 2 commits into
mainfrom
fix/issue-100-reconcile-results

Conversation

@OsherElhadad

Copy link
Copy Markdown
Collaborator

Closes #100

The problem, as it stands today

Issue #100 reported that the source-of-truth relationship was inverted: site/results.html published two Qwen 2.5 14B runs (#qwen-tools, #qwen-all) that docs/RESULTS.md did not contain, while site/README.md told readers to "cross-check against docs/RESULTS.md" — an instruction that was false.

Since the issue was filed the Qwen runs landed in docs/RESULTS.md, and merged PR #236 added a SkillsBench — full 87-task optimization section to docs/RESULTS.md that the site never got. The divergence did not close, it swapped direction. So this PR fixes it structurally rather than patching the specific rows the issue named.

Full inventory — every result claim on every surface

Split column: fit = train == val == test (no holdout) · sealed test = ids the optimizer never saw. Provenance is what actually exists on disk in this repo.

Before

# Claim Surfaces Number Split Model Trials Artifact substantiates?
1 toy_calc seed → optimized README, RESULTS, site/results, site/index 0.0 → 1.0 sealed test none (deterministic) 1 ✅ re-derived by core/tests/test_e2e_slice.py
2 τ² airline baseline val README, RESULTS:47, OUTREACH:47, site/results:126, site/index:153 0.536 fit (50) gpt-oss-120b 10 run_full/ui/data/runs_run_full.jsonbaseline_val: 0.536
3 τ² airline best val cand_0007 README, RESULTS:48, site/results:127, site/index:158 0.712 fit (50) gpt-oss-120b 10 ✅ same file → best_val: 0.712, best_id: cand_0007
4 τ² airline sealed test pass@1 RESULTS:49, OUTREACH:47, site/results:128, site/index:164 0.694 fit — test == val, NOT held out gpt-oss-120b 10 run_full/final.jsontest.reward: 0.694
5 τ² airline pass² RESULTS:49, site/results:128 0.584 fit (50) gpt-oss-120b 10 final.jsontest.pass_k["2"]: 0.5844
6 τ² airline per-iteration stair (5 of 10 accepted) RESULTS:52-53, site/results:134-136 0.582/0.634/0.670/0.684/0.712 fit (50) gpt-oss-120b 10 runs_run_full.jsonevaluations[] + per_iteration[] statuses
7 τ² airline held-out val RESULTS:69, site/results:172 56.7 → 70.0 (doc) / 0.567 → 0.700 (site) val (30) no artifact at any commit (#182 reviewer: 0.475 in zero artifact files) — units also disagreed between surfaces
8 τ² airline held-out sealed test README:110, RESULTS:70, COMPARISON:61, OUTREACH:46, site/results:173, site/index:260, presentation:764 30.0 → 47.5 (+58.3%) sealed test (20) no artifact#99/#182's headline
9 τ² agent-mode single-trial RESULTS:101-102, site/results:212-213 val 0.500→0.633, test 0.400→0.550 val fit (30) + sealed test (20) aws/gpt-oss-120b 1 ❌ no artifact (labeled ⚠️ by #182)
10 τ² agent-mode n=3 re-eval RESULTS:108-109 val 0.544→0.644, test 0.467→0.400 val fit + sealed test aws/gpt-oss-120b 3 ❌ no artifact (doc-only; not on site)
11 τ² agent-mode head-to-head deterministic RESULTS:113, site/results:224 val 0.567, sealed test 0.35 val fit + sealed test aws/gpt-oss-120b 1 ❌ no artifact
12 Qwen 14B tools-only RESULTS:144-145, site/results:256-257 val 0.200→0.387, test 0.170→0.240 (+41.2%) val (30) + sealed test (20) Qwen2.5-14B 5 ❌ no artifact; git log --all finds no Qwen file ever added
13 Qwen 14B all-capabilities RESULTS:166-167, site/results:289-290 val 0.273→0.520, test 0.120→0.270 (+125.0%) val (30) + sealed test (20) Qwen2.5-14B 5 ❌ no artifact
14 SkillsBench-87 baselines (opus + gptoss) docs/runs/ only — on NO published surface 0.281(23/87) · 0.0396(2/87) fit (87) claude-opus-4-6 / aws/gpt-oss-120b 1 ⚠️ run records committed; run dirs on recording host
15 SkillsBench-87 optimization docs/RESULTS.md:177-238 only — MISSING from the site val 0.281→0.357 (+27.2%), pass@1 23→28 fit (87) — test == val claude-opus-4-6 1 ⚠️ run record committed; run dir gitignored
16 SkillsBench-87 iterations RESULTS:207-209 0.357 / 0.325 / 0.170 val fit (87) claude-opus-4-6 1 ⚠️ run record :36-38 (4-dp there, 3-dp in RESULTS)
17 SkillsBench-87 task-level delta RESULTS:216-226 8 newly passing, 3 regressed, net +5 val fit (87) claude-opus-4-6 1 ⚠️ derived exactly from the two run records' passing-task lists
18 SkillsBench-87 optimizer spend RESULTS:212 ~$32 total, $400 cap no cost field in the run record at all
19 SkillsBench-87 EvoSkills comparison RESULTS:235 71.1% no citation in sources.bib; string appears nowhere else in the tree
20 SkillsBench held-out val RESULTS:255-256, site/results:333-334 0.333 → 0.714 (+114%) val fit (7) claude-sonnet-4-6 3 skillsbench/run_full/baseline.jsonval.reward: 0.3333
21 SkillsBench held-out sealed test README:111, RESULTS:257-258, OUTREACH:49, site/results:335-336, site/index:266 0.556 → 0.667 (+20.0%) sealed test (3) claude-sonnet-4-6 3 run_full/final.jsontest_baseline.reward: 0.5556, test.reward: 0.6667
22 presentation slide-9 "Done" run 1 presentation/index.html:964-966 0.600 → 0.800 (+33%) unstated claude-sonnet-4-6 3 on no other surface; no artifact, no run record, at any commit
23 presentation slide-9 "Done" run 2 presentation/index.html:979-981 0.533 → 0.933 (+75%) unstated claude-sonnet-4-6 3 same
24 presentation slide-9 "Done" run 3 presentation/index.html:994-996 0.830 → 0.900 (+12%) unstated claude-sonnet-4-6 1 same
25 presentation slide-9 "Done" run 4 presentation/index.html:1009-1011 1.000 → 1.000 (0%) unstated claude-sonnet-4-6 3 same
26 COMPARISON external rows (EvoTool ×2, Evolutionary Context Search) COMPARISON:58-60 35.9→39.1, 14.4→15.7, +23.3% external GPT-4.1 / Qwen3-8B external; ECS already self-labels "citation still to be confirmed"
27 site/benchmarks.fixture.json fixture only 0.0→0.5, 0.2→0.4 n/a fake n/a n/a — dev fixture, not deployed data; not a claim. Untouched.
28 site/benchmarks.js no hardcoded results — renders a live feed n/a n/a n/a n/a. Untouched.
29 llms.txt no result numbers at all n/a n/a n/a n/a. Untouched.

After

# Change Result
7 units unified docs/RESULTS.md now shows 0.567 (56.7%) / 0.300 (30.0%), identical to site/results.html, obeying the page's own "label percentages explicitly" rule
14 published new docs/RESULTS.md §SkillsBench — full 87-task baselines, no optimization + matching site/results.html section, both ⚠️-stamped, each row linking its run record. This also fixes a dangling reference: §15 said "the same four shared office-document skills as the baseline section above" — there was no such section
15 published on the site new site/results.html §skillsbench-87 with the same numbers, split labels, splice footnote, iterations, and ⚠️ callout
15, 16 split labels rows relabeled val (87, fit) and test (87, fit metric — test == val by construction, not held out) on both surfaces
18 REMOVED "Optimizer spend: ~$32 total (well below the $400 cap)" → "Optimizer spend for this run is not recorded in the run record, so no cost figure is quoted here."
19 REMOVED "Not directly comparable to EvoSkills' 71.1%" → names the paradigm difference and arXiv:2604.01687v1 (the only id present in the tree), with "No external score is quoted here: none is cited in sources.bib, and an uncited number is not evidence."
22-25 REMOVED all four number triplets → ; speaker note corrected from "the results already committed" to "their baseline/optimized numbers are deliberately NOT shown, because no run artifact and no run record for them exists anywhere in the repo"
superset rule docs/RESULTS.md preamble now states it explicitly; site/README.md's false "cross-check against docs/RESULTS.md and the committed run artifacts" replaced with the actual rule (add the canonical section first)
blanket claim site/benchmarks.html's "hand-verified canonical numbers" → points at the per-section evidence markers
mechanical guard core/tests/test_published_results_consistency.py — 3 tests, the issue's own acceptance criterion

What I removed and why

Four categories, 8 numbers, all removed rather than kept:

  1. 71.1% (EvoSkills) — the string appears in exactly one place in the whole repo (the claim itself). No sources.bib entry, no arXiv id, no URL. An uncited external number in an honesty doc is worse than no number.
  2. ~$32 / $400 cap — the run record has no cost field. docs/runs/local-20260730-skillsbench-opus46-optimize.md records wall-clock but not spend.
  3. presentation slide-9's four "Done" rows (8 figures) — 0.600→0.800, 0.533→0.933, 0.830→0.900, 1.000→1.000. Grepped across every published surface: they appear only in that table. No run_full/, no docs/runs/ record, nothing in git history. They sat under a speaker note asserting they were "already committed" — the most directly false claim I found.
  4. "hand-verified" (site/benchmarks.html) and "cross-check against … the committed run artifacts" (site/README.md) — aggregate claims that are not true and, per fix(honesty): label the τ²-bench held-out headline as reported, artifact not committed #182's reasoning, rot whenever a run is added.

Design note: why the guard checks section identity, not numbers

My first attempt scraped decimals out of the HTML and compared sets. It was unusable: SVG coordinates, font weights (400;500;600;700;800), arXiv fragments (2603.04900) and external papers' figures (35.9, 39.1) all look exactly like rewards, and every one needed an allowlist entry. A guard whose allowlist keeps growing stops guarding.

So the guard pins the invariant that actually broke: every site/results.html <h2 id> maps to a docs/RESULTS.md ## heading. That is verbatim the issue's acceptance criterion ("every site/results.html run anchor maps to a RESULTS.md section"). Mapping is an explicit table, not derived from #182's <a id> markers, so the guard's result does not depend on merge order — and a renamed section fails loudly where a set-difference would silently accept it. Proven by reintroducing the original bug (evidence comment).

Deliberately left to sibling PRs

Line / claim Left to Why
README.md:102 "cross-checked against committed run artifacts" + the whole Results table #99 / #182 #182 rewrites this into a per-row ✅/⚠️ Artifact column. Touching it = guaranteed conflict.
README.md:110 held-out headline 30.0 → 47.5 #99 / #182 its headline; I did not change the number, the split label, or the row.
site/index.html:78 "Every number ships its artifact" pill; :236 "Cross-checked against committed run artifacts"; the 4-row hero table #99 / #182 #182 rewrites all of them (pill → "Artifact-backed by default", table gains an Artifact column). I did not touch site/index.html at all.
site/results.html:7,12,17 meta descriptions "Every number cross-checked…" #99 / #182 (and #196 owns the file region) #182 rewrites all three; under #196 the <head> is generator-owned via chrome:head:* sentinels, so editing it would fail --check. Doubly not mine.
site/results.html:55 lead; the 7 pre-existing <h2> evidence stamps #99 / #182 its convention. I extended it to my 2 new sections using its exact <span class="muted">— ⚠️ reported, artifact not committed</span> form.
presentation/index.html:635, 745, 788 + slide-6 cards #99 / #182 #182 rewrites each. My presentation edit is only slide 9 (lines 964-1011 + its <aside class="notes">) — zero overlap, confirmed by a test merge that auto-merged presentation/index.html with no conflict.
CHANGELOG.md (0.536 → 0.712, 0.694 named separately) #101 / #184 untouched.
OUTREACH.md:46 (val/test conflation) and :47 #101 / #184 and #99 / #182 both edit these lines. Untouched.
Skill / algorithm / optimizer counts #102 / #189 untouched.
Site chrome, nav, footer, <head>, ?v= hashes #123 / #196 all my site edits are inside <main> bodies.
harbor.html / harbor-openshift.html not registered in PAGES #123 / #196 pre-existing on main; proven identical without my changes (evidence comment).
Host labels #126 / #202 untouched.
test_dashboard_launch.py port-7878 flake #200 pre-existing; reproduced on clean origin/main.

Expected merge order

#189 (counts) ─┐
#184 (changelog)─┤  independent
#182 (#99 headline) ──► THIS PR (#100) ──► #202 (#126) / #192 (#120) ──► #196 (#123) LAST

Verification

Full suite — 181 passed (179 baseline + 3 new), 0 unexpected failures

$ PYTHONPATH=/tmp/wt-100b/core /tmp/ce-venv/bin/python -m pytest core/tests -q
FAILED core/tests/test_dashboard_launch.py::test_maybe_launch_spawns_when_available
1 failed, 181 passed in 64.82s (0:01:04)

The single failure is the known port-7878 flake (#200), reproduced on clean origin/main with none of my changes present:

$ cd /tmp/wt-ctl && git reset --hard -q origin/main
$ PYTHONPATH=/tmp/wt-ctl/core pytest core/tests/test_dashboard_launch.py -q
FAILED core/tests/test_dashboard_launch.py::test_maybe_launch_spawns_when_available
1 failed, 6 passed in 0.05s

Deselecting it: 181 passed, 0 failed.

New guard passes, and catches the #100 bug

$ pytest core/tests/test_published_results_consistency.py -q
...                                                                      [100%]
3 passed in 0.01s

Regression-proof — delete the canonical qwen-tools section (the exact original bug) and it fails:

E  AssertionError: site/results.html publishes run section(s) whose docs/RESULTS.md
   heading is gone or renamed: {'qwen-tools': 'Qwen 2.5 14B (tools only'}

python -m compileall core skills — clean

$ /tmp/ce-venv/bin/python -m compileall -q core skills
compileall EXIT=0

Every 87-task number identical on both surfaces, with the same split label

--- 0.281 ---
docs/RESULTS.md:197:| `claude-opus-4-6` | `run_baseline_opus` | **0.281 ± 0.048** (28.1%) | **23** (26.4%) | ...
docs/RESULTS.md:228:| **val** (87, fit) — mean reward | 0.281 ± 0.048 | **0.357 ± 0.050** | **+0.076 (+27.2% relative)** |
docs/RESULTS.md:230:| **test** (87, *fit metric* — `test == val` by construction, **not** held out) | 0.281 * | **0.357** | +0.076 |
site/results.html:327: ...<strong>0.281 ± 0.048</strong> (28.1%)...<strong>23</strong> (26.4%)...
site/results.html:361: ...<strong>val</strong> (87, fit) — mean reward...0.281 ± 0.048...<strong>0.357 ± 0.050</strong>...+0.076 / +27.2% relative
site/results.html:363: ...<strong>test</strong> (87, <em>fit metric</em> — <code>test == val</code> by construction, <strong>not</strong> held out)...
--- 26.4 ---
docs/runs/local-20260729-skillsbench-baseline-opus46.md:22:| pass_at_1 (fully passing) | **23/87** (26.4%) |
docs/RESULTS.md:197  site/results.html:327   docs/RESULTS.md:229   site/results.html:362
--- 0.040 (== 0.0396 in the run record, 3-dp on published surfaces) ---
docs/RESULTS.md:198   site/results.html:328
docs/runs/local-20260729-skillsbench-baseline-gptoss120b.md:21:| val_reward (mean) | **0.0396 ± 0.0191** (4.0%) |

Zero surviving aggregate claims

On the merged #182 + #100 tree:

$ grep -rniE "every number|every result|all results|cross-check|ships its artifact|committed run artifact|hand-verified|every figure" \
    README.md docs/RESULTS.md docs/COMPARISON.md OUTREACH.md llms.txt site/*.html site/README.md presentation/index.html
docs/RESULTS.md:25:**This page is the superset.** Every result claim on any other surface — `README.md`,

One hit, and it is TRUE. It is my own line and is not an artifact claim — it is a normative rule about where numbers may be published, and it is the rule test_published_results_consistency.py mechanically enforces. Every "every number is artifact-backed"-style claim is gone.

Removed numbers are gone from every published surface

71.1       occurrences on published surfaces now: 0
$32        occurrences on published surfaces now: 0
0.933      occurrences on published surfaces now: 0
0.830      occurrences on published surfaces now: 0
0.533      occurrences on published surfaces now: 0
0.800 / 0.900 / 0.600  -> remaining hits are Fira Sans font weights (`400;500;600;700;800`)
                          and arXiv ids (2603.04900). Not results.

Links and anchors resolve on disk (CI docs-links runs with --include-fragments)

docs/RESULTS.md:      16 resolve, 0 broken   (incl. runs/, #skillsbench-87-baselines ×2, sources.bib)
site/results.html:    10/10 internal anchors OK, 11/11 relative hrefs OK

Site pages still parse; fixture untouched

site/results.html:      unclosed=none  mismatched=none
site/benchmarks.html:   unclosed=none  mismatched=none
presentation/index.html:unclosed=none  mismatched=none
fixture ok (untouched)

site/benchmarks.js and site/benchmarks.fixture.json are not modified — the fixture is a dev eyeball fixture, not deployed data, and benchmarks.js hardcodes no results.

#196's generator: --check exits 0 on the merged tree

main + #182 + this PR + #196:

$ scripts/sync-site-chrome.py --check
site chrome in sync (9 pages)
EXIT=0

And my edits cause zero chrome drift — chrome regions are byte-identical with and without them:

IDENTICAL chrome: index.html, getting-started.html, run-end-to-end.html, results.html,
                  benchmarks.html, architecture.html, optimize-your-own.html,
                  adapter-templates.html, agent-orchestration.html

(The harbor.html / harbor-openshift.html registration failure is pre-existing — byte-identical on main + #196 alone. #196's to resolve; full both-trees comparison in the evidence comment.)

Could not substantiate

Nothing survives unsubstantiated. Every number I kept has either a committed artifact (✅) or a committed run record with the artifact's absence stated in-line (⚠️). The 8 numbers with neither were removed.

One honest limitation, unchanged and pre-existing: five τ²-bench runs have no committed artifact anywhere (held-out, agent-mode, and both Qwen runs — the reviewer of #182 confirmed 0.475 appears in zero artifact files at any commit; I independently confirmed no Qwen file was ever added in git history). They are ⚠️-labeled, not deleted, because that is #99/#182's scope and its reviewer's explicit decision. My contribution is that the labeling convention now covers all 9 sections instead of 7, and both surfaces carry all 9.

Files touched

  • docs/RESULTS.md
  • site/results.html
  • site/benchmarks.html
  • site/README.md
  • presentation/index.html
  • core/tests/test_published_results_consistency.py (new)

… runs everywhere with true provenance

Issue #100: the source-of-truth relationship was inverted. site/results.html
published runs docs/RESULTS.md did not have, while site/README.md told readers to
"cross-check against docs/RESULTS.md" — an instruction that was false.

Since the issue was filed the Qwen 14B runs landed in docs/RESULTS.md, and PR #236
added a SkillsBench 87-task optimization section there that the site never got.
The divergence just moved rather than closing, so this fixes it structurally:

- docs/RESULTS.md declares itself the SUPERSET, and a new guard test enforces it
  by section identity (core/tests/test_published_results_consistency.py). The
  guard fails if a site run section has no canonical counterpart — the exact #100
  bug, verified by reintroducing it.
- Adds the missing "SkillsBench — full 87-task baselines" section, which also
  fixes a dangling "the baseline section above" reference in the 87-task
  optimization section (it pointed at a section that was never written).
- Publishes both 87-task sections on site/results.html with split labels, run-record
  links, and the ⚠️ reported marker (#99/#182's per-section convention), plus TOC rows.
- Labels every 87-task figure with its split: "val (87, fit)" and
  "test (87, fit metric — test == val by construction, not held out)".
- docs/RESULTS.md held-out table now shows 0.567/0.300 (56.7%/30.0%) instead of
  bare 56.7/30.0, matching the site and the page's own units rule.

Removes numbers nothing substantiates rather than keeping them:
- "EvoSkills' 71.1%" — no citation in sources.bib, no source anywhere in the tree.
- "Optimizer spend: ~$32 total (well below the $400 cap)" — no cost field in the
  run record; the sentence now says spend is not recorded.
- presentation/ slide 9's four "Done" rows quoted baseline/optimized/Δ figures that
  exist on no other surface and have no artifact or run record at any commit, under
  a speaker note claiming they were "already committed". Numbers blanked to —, note
  corrected to say the write-ups are not published.

Also removes the two blanket-verification claims #182 does not reach:
site/benchmarks.html's "hand-verified canonical numbers" and site/README.md's
"cross-check against docs/RESULTS.md and the committed run artifacts", the latter
replaced with the actual superset rule.
Copilot AI review requested due to automatic review settings July 30, 2026 14:34

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@skillberry-bot

Copy link
Copy Markdown
Contributor

Automatic Labeling Failed

An error occurred while trying to automatically label this pull request. Please check the workflow logs for details and add labels manually.

@OsherElhadad

Copy link
Copy Markdown
Collaborator Author

🔬 Evidence

All commands run in /tmp/wt-100b (worktree of origin/main @ 75139209), control tree /tmp/wt-ctl, Python /tmp/ce-venv/bin/python. Full literal output.

1. Setup

$ git fetch origin && git worktree add /tmp/wt-100b -b fix/issue-100-reconcile-results origin/main
Preparing worktree (new branch 'fix/issue-100-reconcile-results')
branch 'fix/issue-100-reconcile-results' set up to track 'origin/main'.
HEAD is now at 75139209 docs: opus-4.6 optimization run + summarize_run.py tool (#236)

2. Artifacts that DO exist on disk (the ✅ rows)

$ ls examples/tau2_airline/run_full/ examples/skillsbench/run_full/
examples/skillsbench/run_full/:
JOURNAL.md
SKILLSBENCH_COMMIT.txt
baseline.json
events.jsonl
final.json
history.jsonl
rejected.jsonl
report.md

examples/tau2_airline/run_full/:
TAU2_COMMIT.txt
demo.cast
final.json
ui

τ²-bench airline — every ✅ number, from the artifact:

$ python -c "import json; d=json.load(open('examples/tau2_airline/run_full/final.json'))"
final.json  test.reward      = 0.6940000000000001   <- published 0.694
final.json  test.pass_k["1"] = 0.6940000000000001
final.json  test.pass_k["2"] = 0.5844444444444444  <- published pass² 0.584
final.json  best_id          = cand_0007   <- published cand_0007
final.json  n per_task       = 50    <- published 50 tasks
final.json  trials/task      = 10     <- published 10 trials
ui runs.json baseline_val    = 0.536   <- published 0.536
ui runs.json best_val        = 0.7120000000000001   <- published 0.712
ui runs.json delta_abs       = 0.176   <- published +0.176
ui runs.json delta_pct       = 32.8    <- published +32.8%
ui runs.json test_sealed     = True
ui runs.json counts          = {'accepted': 5, 'rejected': 5, 'failed': 0, 'seed': 1, 'total': 11} <- published 5 of 10 accepted
ui runs.json val stair       = [0.536, 0.582, 0.512, 0.634, 0.636, 0.67, 0.684, 0.712, 0.702, 0.69, 0.706]
   (published accepted stair: 0.536 -> 0.582 -> 0.634 -> 0.670 -> 0.684 -> 0.712)

SkillsBench held-out — every ✅ number, from the artifact:

baseline.json val.reward       = 0.33333333333333337  <- published 0.333
baseline.json val n tasks      = 7                   <- published 7 val tasks
final.json test_baseline.reward= 0.5555555555555555  <- published 0.556
final.json test.reward         = 0.6666666666666666  <- published 0.667
final.json test_delta          = 0.111111             <- published +0.111
final.json best_id             = cand_0004           <- published cand_0004
final.json test n tasks        = 3                   <- published 3 sealed test tasks

3. Artifacts that do NOT exist — the ❌ / ⚠️ rows

No Qwen file was ever added in git history:

$ git log --all --oneline --diff-filter=A --name-only | grep -i qwen
(no output — zero files matching /qwen/i added at any commit on any branch)

The held-out / Qwen test values appear in no artifact JSON at any of the last 40 commits (independently reproducing #182's reviewer finding for 0.475):

$ git grep -l -- "0.475" $(git rev-list --all | head -40) -- "*.json" | grep -E "run_full|final|baseline|split_ids"
(no output — zero artifact files)
$ git grep -l -- "0.387" $(git rev-list --all | head -40) -- "*.json" | grep -E "run_full|final|baseline|split_ids"
(no output — zero artifact files)
$ git grep -l -- "0.240" $(git rev-list --all | head -40) -- "*.json" | grep -E "run_full|final|baseline|split_ids"
(no output — zero artifact files)
$ git grep -l -- "0.270" $(git rev-list --all | head -40) -- "*.json" | grep -E "run_full|final|baseline|split_ids"
(no output — zero artifact files)

71.1 appears in exactly ONE place in the whole repo — the claim itself. No sources.bib entry:

$ grep -rn "71\.1" --include=*.md --include=*.html --include=*.bib --include=*.txt . | grep -v reveal/ | grep -v run_full/ui/
docs/RESULTS.md:235 (on main): 235:- **Not directly comparable to EvoSkills' 71.1%.** Different paradigm (shared
$ grep -i -A6 "evoskill" docs/sources.bib
(no output — no EvoSkills/EvoSkill entry in sources.bib)

No cost field in the SkillsBench-87 run record (the ~$32 claim):

$ grep -in "cost\|usd\|spend\|\$" docs/runs/local-20260730-skillsbench-opus46-optimize.md
(no output — the run record records wall-clock but never spend)

presentation slide-9 "Done" figures exist on NO other surface:

$ grep -Fn "0.800" README.md docs/*.md OUTREACH.md llms.txt site/*.html presentation/index.html docs/runs/*.md examples/*/*.md
  (only presentation/index.html — nowhere else)
$ grep -Fn "0.933" README.md docs/*.md OUTREACH.md llms.txt site/*.html presentation/index.html docs/runs/*.md examples/*/*.md
  (only presentation/index.html — nowhere else)
$ grep -Fn "0.830" README.md docs/*.md OUTREACH.md llms.txt site/*.html presentation/index.html docs/runs/*.md examples/*/*.md
  (only presentation/index.html — nowhere else)
$ grep -Fn "0.900" README.md docs/*.md OUTREACH.md llms.txt site/*.html presentation/index.html docs/runs/*.md examples/*/*.md
  (only presentation/index.html — nowhere else)
$ grep -Fn "0.533" README.md docs/*.md OUTREACH.md llms.txt site/*.html presentation/index.html docs/runs/*.md examples/*/*.md
  (only presentation/index.html — nowhere else)

4. Numbers I kept as ⚠️ — derived exactly from the committed run records

The 87-task task-level delta claim (8 newly passing / 3 regressed / net +5) recomputed from the two run records' passing-task lists:

baseline passing        : 23  <- published 23
optimized passing       : 28  <- published 28
net                     : +5  <- published +5
newly passing (derived) : ['bike-rebalance', 'energy-ac-optimal-power-flow', 'energy-market-pricing', 'exceltable-in-ppt', 'grid-dispatch-operator', 'paper-anonymizer', 'paratransit-routing', 'weighted-gdp-calc']
regressed     (derived) : ['citation-check', 'crystallographic-wyckoff-position-analysis', 'pptx-reference-formatting']
published newly == derived  : True
published regressed == derived: True

5. Full test suite

$ PYTHONPATH=/tmp/wt-100b/core python -m pytest core/tests -q
core/tests/test_dashboard_launch.py:56: AssertionError
=========================== short test summary info ============================
FAILED core/tests/test_dashboard_launch.py::test_maybe_launch_spawns_when_available
1 failed, 181 passed in 64.87s (0:01:04)

The one failure is the known #200 flake — reproduced on clean origin/main, none of my changes present:

$ cd /tmp/wt-ctl && git reset --hard -q origin/main
$ PYTHONPATH=/tmp/wt-ctl/core python -m pytest core/tests/test_dashboard_launch.py -q
=========================== short test summary info ============================
FAILED core/tests/test_dashboard_launch.py::test_maybe_launch_spawns_when_available
1 failed, 6 passed in 0.03s

Deselecting it:

$ PYTHONPATH=/tmp/wt-100b/core python -m pytest core/tests -q -k "not maybe_launch_spawns_when_available"
........................................................................ [ 79%]
.....................................                                    [100%]
181 passed, 1 deselected in 63.06s (0:01:03)

6. New guard test — passes, and catches the original #100 bug

$ PYTHONPATH=/tmp/wt-100b/core python -m pytest core/tests/test_published_results_consistency.py -q
...                                                                      [100%]
3 passed in 0.01s

Regression proof — delete the canonical qwen-tools section (verbatim the #100 bug) and the guard fails:

$ # (canonical qwen-tools heading removed)
$ pytest core/tests/test_published_results_consistency.py -q
E       AssertionError: site/results.html publishes run section(s) whose docs/RESULTS.md heading is gone or renamed: {'qwen-tools': 'Qwen 2.5 14B (tools only'}
E         Headings present: ['toy_calc — deterministic, zero-API', 'τ²-Bench airline — no-holdout fit-metric run (reproducible, committed)', 'τ²-Bench airline — held-out 30(=val)/20 run', 'τ²-Bench airline — agent orchestration mode (`agent-optimize`), held-out 30(=val)/20', 'REMOVED-SECTION', 'τ²-Bench airline — Qwen 2.5 14B, all capabilities (held-out)', 'SkillsBench — full 87-task baselines, no optimization — ⚠️ reported, artifact not committed', 'SkillsBench — full 87-task optimization (Opus 4.6) — ⚠️ reported, artifact not committed', 'SkillsBench — skill-package optimization (held-out, committed)']
E       assert not {'qwen-tools': 'Qwen 2.5 14B (tools only'}
=========================== short test summary info ============================
1 failed, 2 passed in 0.02s

$ # (restored)
$ pytest core/tests/test_published_results_consistency.py -q
...                                                                      [100%]
3 passed in 0.01s

7. compileall

$ python -m compileall -q core skills; echo "EXIT=$?"
EXIT=0

8. Cross-surface consistency — same value, same split label, everywhere

### 0.281
  docs/RESULTS.md:197:| `claude-opus-4-6` | `run_baseline_opus` | **0.281 ± 0.048** (28.1%) | **23** (26.4%) | [`runs/local-20260729-skillsbench-baseline-opus46.md`](runs/local-20260729-skillsbench-baseline-opus46.md) |
  docs/RESULTS.md:228:| **val** (87, fit) — mean reward | 0.281 ± 0.048 | **0.357 ± 0.050** | **+0.076 (+27.2% relative)** |
  docs/RESULTS.md:230:| **test** (87, *fit metric* — `test == val` by construction, **not** held out) | 0.281 * | **0.357** | +0.076 |
  site/results.html:327:            <tr><td><code>claude-opus-4-6</code></td><td><code>run_baseline_opus</code></td><td class="num"><strong>0.281 ± 0.048</strong> (28.1%)</td><td class="num"><strong>23</strong> (26.4%)</td><td><a href="https://github.com/skillberry-ai/cap-evolve/blob/main/docs/runs/local-20260729-skillsbench-baseline-opus46.md">run record</a></td></tr>
  site/results.html:361:            <tr><td><strong>val</strong> (87, fit) — mean reward</td><td class="num">0.281 ± 0.048</td><td class="num"><strong>0.357 ± 0.050</strong></td><td class="gain">+0.076 / +27.2% relative</td></tr>
  site/results.html:363:            <tr><td><strong>test</strong> (87, <em>fit metric</em> — <code>test == val</code> by construction, <strong>not</strong> held out)</td><td class="num">0.281 *</td><td class="num"><strong>0.357</strong></td><td class="num">+0.076</td></tr>
### 0.357
  site/results.html:361:            <tr><td><strong>val</strong> (87, fit) — mean reward</td><td class="num">0.281 ± 0.048</td><td class="num"><strong>0.357 ± 0.050</strong></td><td class="gain">+0.076 / +27.2% relative</td></tr>
  site/results.html:363:            <tr><td><strong>test</strong> (87, <em>fit metric</em> — <code>test == val</code> by construction, <strong>not</strong> held out)</td><td class="num">0.281 *</td><td class="num"><strong>0.357</strong></td><td class="num">+0.076</td></tr>
  site/results.html:376:          <strong>Iterations:</strong> <code>cand_0001</code> val <strong>0.357</strong>
  docs/RESULTS.md:228:| **val** (87, fit) — mean reward | 0.281 ± 0.048 | **0.357 ± 0.050** | **+0.076 (+27.2% relative)** |
  docs/RESULTS.md:230:| **test** (87, *fit metric* — `test == val` by construction, **not** held out) | 0.281 * | **0.357** | +0.076 |
  docs/RESULTS.md:244:| 1 | `cand_0001` | `seed` | **0.357** | **+0.077** | ✓ (paired gate: Δ > 0.2·SE) |
### 26.4
  docs/RESULTS.md:197:| `claude-opus-4-6` | `run_baseline_opus` | **0.281 ± 0.048** (28.1%) | **23** (26.4%) | [`runs/local-20260729-skillsbench-baseline-opus46.md`](runs/local-20260729-skillsbench-baseline-opus46.md) |
  docs/RESULTS.md:229:| **val** (87, fit) — pass_at_1 (fully-passing / 87) | 23 (26.4%) | **28 (32.2%)** | **+5 tasks (+22% relative)** |
  site/results.html:327:            <tr><td><code>claude-opus-4-6</code></td><td><code>run_baseline_opus</code></td><td class="num"><strong>0.281 ± 0.048</strong> (28.1%)</td><td class="num"><strong>23</strong> (26.4%)</td><td><a href="https://github.com/skillberry-ai/cap-evolve/blob/main/docs/runs/local-20260729-skillsbench-baseline-opus46.md">run record</a></td></tr>
  site/results.html:362:            <tr><td><strong>val</strong> (87, fit) — pass_at_1 (fully-passing / 87)</td><td class="num">23 (26.4%)</td><td class="num"><strong>28 (32.2%)</strong></td><td class="gain">+5 tasks / +22% relative</td></tr>
  docs/runs/local-20260729-skillsbench-baseline-opus46.md:22:| pass_at_1 (fully passing) | **23/87** (26.4%) |
### 32.2
  docs/RESULTS.md:229:| **val** (87, fit) — pass_at_1 (fully-passing / 87) | 23 (26.4%) | **28 (32.2%)** | **+5 tasks (+22% relative)** |
  site/results.html:362:            <tr><td><strong>val</strong> (87, fit) — pass_at_1 (fully-passing / 87)</td><td class="num">23 (26.4%)</td><td class="num"><strong>28 (32.2%)</strong></td><td class="gain">+5 tasks / +22% relative</td></tr>
### 27.2
  site/results.html:361:            <tr><td><strong>val</strong> (87, fit) — mean reward</td><td class="num">0.281 ± 0.048</td><td class="num"><strong>0.357 ± 0.050</strong></td><td class="gain">+0.076 / +27.2% relative</td></tr>
  docs/RESULTS.md:228:| **val** (87, fit) — mean reward | 0.281 ± 0.048 | **0.357 ± 0.050** | **+0.076 (+27.2% relative)** |
### 0.040
  docs/RESULTS.md:198:| `aws/gpt-oss-120b` | `run_baseline_gptoss` | **0.040 ± 0.019** (4.0%) | **2** (2.3%) | [`runs/local-20260729-skillsbench-baseline-gptoss120b.md`](runs/local-20260729-skillsbench-baseline-gptoss120b.md) |
  site/results.html:328:            <tr><td><code>aws/gpt-oss-120b</code></td><td><code>run_baseline_gptoss</code></td><td class="num"><strong>0.040 ± 0.019</strong> (4.0%)</td><td class="num"><strong>2</strong> (2.3%)</td><td><a href="https://github.com/skillberry-ai/cap-evolve/blob/main/docs/runs/local-20260729-skillsbench-baseline-gptoss120b.md">run record</a></td></tr>
### 0.0396
  docs/runs/local-20260729-skillsbench-baseline-gptoss120b.md:21:| val_reward (mean) | **0.0396 ± 0.0191** (4.0%) |
  docs/runs/local-20260729-skillsbench-baseline-gptoss120b.md:23:| test_reward | **0.0396** |
### 0.325
  docs/RESULTS.md:245:| 2 | `cand_0002` | `cand_0001` | 0.325 | −0.033 | ✗ regressed |
  site/results.html:377:          (Δ +0.077) — accepted by the paired gate; <code>cand_0002</code> 0.325 (−0.033) and
### 23
  docs/runs/local-20260729-skillsbench-baseline-opus46.md:22:| pass_at_1 (fully passing) | **23/87** (26.4%) |
  docs/runs/local-20260729-skillsbench-baseline-opus46.md:31:- Passed (r = 1.0): **23**
  docs/RESULTS.md:62:What changed: deep in-code tool edits (`tools.py` 593 → 832 lines; policy 166 → 233
  docs/RESULTS.md:76:| **val** (30 tasks) | **0.567** (56.7%) | **0.700** (70.0%) | **+13.3 pp / +23.5% relative** |
  docs/RESULTS.md:197:| `claude-opus-4-6` | `run_baseline_opus` | **0.281 ± 0.048** (28.1%) | **23** (26.4%) | [`runs/local-20260729-skillsbench-baseline-opus46.md`](runs/local-20260729-skillsbench-baseline-opus46.md) |
  docs/RESULTS.md:229:| **val** (87, fit) — pass_at_1 (fully-passing / 87) | 23 (26.4%) | **28 (32.2%)** | **+5 tasks (+22% relative)** |
### 28
  docs/RESULTS.md:60:iter 6 `+0.014` (→0.684), iter 7 `+0.028` (→0.712). **5 of 10** iterations accepted.
  docs/RESULTS.md:197:| `claude-opus-4-6` | `run_baseline_opus` | **0.281 ± 0.048** (28.1%) | **23** (26.4%) | [`runs/local-20260729-skillsbench-baseline-opus46.md`](runs/local-20260729-skillsbench-baseline-opus46.md) |
  docs/RESULTS.md:228:| **val** (87, fit) — mean reward | 0.281 ± 0.048 | **0.357 ± 0.050** | **+0.076 (+27.2% relative)** |
  docs/RESULTS.md:229:| **val** (87, fit) — pass_at_1 (fully-passing / 87) | 23 (26.4%) | **28 (32.2%)** | **+5 tasks (+22% relative)** |
  docs/RESULTS.md:230:| **test** (87, *fit metric* — `test == val` by construction, **not** held out) | 0.281 * | **0.357** | +0.076 |
  docs/RESULTS.md:262:Net: **+5 tasks** (23 → 28) on `pass_at_1`. Two of the newly-passing tasks and

9. Zero surviving aggregate claims

On this PR alone (#182 not merged) — every remaining hit is a #182-owned line:

$ grep -rniE "every number|every result|all results|cross-check|ships its artifact|committed run artifact|hand-verified|every figure" README.md docs/RESULTS.md docs/COMPARISON.md OUTREACH.md llms.txt site/*.html site/README.md presentation/index.html
docs/RESULTS.md:3:The canonical results for cap-evolve. Every number here is derived from a committed
docs/RESULTS.md:17:**This page is the superset.** Every result claim on any other surface — `README.md`,
README.md:102:Numbers are cross-checked against committed run artifacts; each is labeled **fit metric**
site/results.html:7:  <meta name="description" content="Full benchmark results for cap-evolve — toy_calc, τ²-bench airline (fit-metric + held-out)
site/results.html:12:  <meta property="og:description" content="Full benchmark results for cap-evolve — toy_calc, τ²-bench airline (fit-metric + h
site/results.html:17:  <meta name="twitter:description" content="Full benchmark results for cap-evolve — toy_calc, τ²-bench airline (fit-metric + 
site/results.html:55:    The canonical results for cap-evolve. Every number here is derived from a committed run
site/index.html:78:      <span class="pill"><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="3" stroke-linecap="round" stroke-
site/index.html:236:      Cross-checked against committed run artifacts. Each is labeled <strong>fit metric</strong>
presentation/index.html:635:          <div class="diff-desc">Sealed test, gate-verified acceptance, artifacts committed with every result.</div>
presentation/index.html:745:      <p class="subtitle" style="font-size:0.65em;">Every number is derived from a committed run artifact.</p>
presentation/index.html:788:        and stopped on a real ceiling. Every result is committed in examples/*/run_full/.

On the merged #182 + #100 tree:

$ git merge origin/fix/issue-99-sealed-headline   # union-resolved: my sections + #182's stamped headings
$ grep -rniE "every number|every result|all results|cross-check|ships its artifact|committed run artifact|hand-verified|every figure" \
    README.md docs/RESULTS.md docs/COMPARISON.md OUTREACH.md llms.txt site/*.html site/README.md presentation/index.html
docs/RESULTS.md:25:**This page is the superset.** Every result claim on any other surface — `README.md`,

One hit, and it is TRUE. It is my own line, and it is not an artifact claim — it is a normative rule about where a number may be published, and it is exactly the rule test_published_results_consistency.py mechanically enforces. Every "every number is artifact-backed"-style claim is gone.

All 9 sections carry an individual evidence marker on the merged tree (7 from #182, 2 from this PR):

$ grep -n "^## " docs/RESULTS.md
36:## toy_calc — deterministic, zero-API — ✅ reproducible in CI
50:## τ²-Bench airline — no-holdout fit-metric run — ✅ artifact-backed
83:## τ²-Bench airline — held-out 30(=val)/20 run — ⚠️ reported, artifact not committed
111:## τ²-Bench airline — agent orchestration mode (`agent-optimize`), held-out 30(=val)/20 — ⚠️ reported, artifact not committed
159:## τ²-Bench airline — Qwen 2.5 14B (tools only, held-out) — ⚠️ reported, artifact not committed
185:## τ²-Bench airline — Qwen 2.5 14B, all capabilities (held-out) — ⚠️ reported, artifact not committed
213:## SkillsBench — full 87-task baselines, no optimization — ⚠️ reported, artifact not committed
235:## SkillsBench — full 87-task optimization (Opus 4.6) — ⚠️ reported, artifact not committed
310:## SkillsBench — skill-package optimization (held-out) — ✅ artifact-backed

$ grep -n "<h2 id=" site/results.html
87:        <h2 id="toy-calc">toy_calc — deterministic, zero-API 
        MARKER: — ✅ reproducible in CI</span></h2>
112:        <h2 id="tau2-fit">τ²-bench airline — no-holdout fit-metric run 
        MARKER: — ✅ artifact-backed</span></h2>
168:        <h2 id="tau2-heldout">τ²-bench airline — held-out 30(=val)/20 run 
        MARKER: — ⚠️ reported, artifact not committed</span></h2>
208:        <h2 id="tau2-agent">τ²-bench airline — agent orchestration mode (held-out 30 (=val) / 20) 
        MARKER: — ⚠️ reported, artifact not committed</span></h2>
245:        <h2 id="qwen-tools">τ²-bench airline — Qwen 2.5 14B 
        MARKER: — ⚠️ reported, artifact not committed</span></h2>
280:        <h2 id="qwen-all">τ²-bench airline — Qwen 2.5 14B, all capabilities 
        MARKER: — ⚠️ reported, artifact not committed</span></h2>
323:        <h2 id="skillsbench-87-baselines">SkillsBench — full 87-task baselines, no optimization 
        MARKER: — ⚠️ reported, artifact not committed</span></h2>
356:        <h2 id="skillsbench-87">SkillsBench — full 87-task optimization (Opus 4.6) 
        MARKER: — ⚠️ reported, artifact not committed</span></h2>
407:        <h2 id="skillsbench">SkillsBench — skill-package optimization (held-out) 
        MARKER: — ✅ artifact-backed</span></h2>
452:        <h2 id="at-a-glance">At a glance — baseline → optimized</h2>

10. Removed numbers gone from every published surface

$ grep -rEc -- "71.1" <published surfaces>  ->  0
$ grep -rEc -- "\$32" <published surfaces>  ->  0
$ grep -rEc -- "0.933" <published surfaces>  ->  0
$ grep -rEc -- "0.830" <published surfaces>  ->  0
$ grep -rEc -- "0.533" <published surfaces>  ->  0
$ grep -rEc -- "0.800" <published surfaces>  ->  2
$ grep -rEc -- "0.900" <published surfaces>  ->  3
$ grep -rEc -- "0.600" <published surfaces>  ->  2

The three non-zero counts are not results — Fira Sans font weights and arXiv ids:

$ grep -rn -- "0.800" <published surfaces>
site/results.html:23:  <link href="https://fonts.googleapis.com/css2?family=Fira+Code:wght@400;500;600&family=Fira+Sans:
site/index.html:23:  <link href="https://fonts.googleapis.com/css2?family=Fira+Code:wght@400;500;600&family=Fira+Sans:wg
$ grep -rn -- "0.900" <published surfaces>
docs/COMPARISON.md:58:| **EvoTool** (arXiv:2603.04900) | **original τ-Bench** airline, GPT-4.1 | ReAct 35.9 → **39.1*
docs/COMPARISON.md:67:- EvoTool figures are quoted from its Table 1 (verified against arXiv:2603.04900).
README.md:114:external tool-optimization work ([EvoTool](https://arxiv.org/abs/2603.04900) on the
$ grep -rn -- "0.600" <published surfaces>
site/results.html:23:  <link href="https://fonts.googleapis.com/css2?family=Fira+Code:wght@400;500;600&family=Fira+Sans:
site/index.html:23:  <link href="https://fonts.googleapis.com/css2?family=Fira+Code:wght@400;500;600&family=Fira+Sans:wg

11. Links and anchors resolve on disk (CI docs-links runs --include-fragments)

$ # every relative link + in-page anchor in docs/RESULTS.md
  OK   runs/
  OK   ../examples/tau2_airline/run_full/
  OK   REPRODUCE_tau2.md
  OK   OPTIMIZATION_EXAMPLES.md
  OK   ../examples/tau2_airline/DEMO.md
  OK   COMPARISON.md
  OK   AGENT_ORCHESTRATION.md
  OK   REPRODUCE_tau2.md#8-agent-mode-reproduction-held-out-3020-litellm-proxy
  OK   runs/local-20260729-skillsbench-baseline-opus46.md
  OK   runs/local-20260729-skillsbench-baseline-gptoss120b.md
  OK   #skillsbench-87-baselines
  OK   runs/local-20260730-skillsbench-opus46-optimize.md
  OK   #skillsbench-87-baselines
  OK   sources.bib
  OK   ../examples/skillsbench/run_full/
  OK   REPRODUCE_skillsbench.md
-> 16 resolve, 0 broken

$ # every in-page anchor + relative href in site/results.html
  OK   #at-a-glance
  OK   #qwen-all
  OK   #qwen-tools
  OK   #skillsbench
  OK   #skillsbench-87
  OK   #skillsbench-87-baselines
  OK   #tau2-agent
  OK   #tau2-fit
  OK   #tau2-heldout
  OK   #toy-calc
  OK   ./
  OK   agent-orchestration.html
  OK   architecture.html
  OK   assets/apple-touch-icon.png
  OK   assets/favicon.svg
  OK   benchmarks.html
  OK   getting-started.html
  OK   harbor.html
  OK   results.html
  OK   run-end-to-end.html
  OK   style.css
-> 21 resolve, 0 broken

12. Site pages still parse; benchmarks.js / fixture untouched

site/results.html            unclosed=none         mismatched=none
site/benchmarks.html         unclosed=none         mismatched=none
presentation/index.html      unclosed=none         mismatched=none
$ python -c "import json; json.load(open('site/benchmarks.fixture.json'))"
valid JSON
$ git diff --stat origin/main -- site/benchmarks.js site/benchmarks.fixture.json site/index.html README.md OUTREACH.md CHANGELOG.md llms.txt docs/COMPARISON.md
(no output = all untouched: no site/index.html, no README, no CHANGELOG, no counts, no host labels, no benchmarks.js, no fixture)

13. Files this PR touches

$ git diff --stat origin/main
 core/tests/test_published_results_consistency.py | 106 +++++++++++++++++++++++
 docs/RESULTS.md                                  |  81 ++++++++++++-----
 presentation/index.html                          |  22 ++---
 site/README.md                                   |  11 ++-
 site/benchmarks.html                             |   3 +-
 site/results.html                                |  88 ++++++++++++++++++-
 6 files changed, 272 insertions(+), 39 deletions(-)

14. Commit authorship — Osher Elhadad, no Claude trailer

$ git log --format="%H%n%an <%ae>%n%cn <%ce>" -1
0c0febf9a95b24756a0d8b454ee1f7dc902062fb
Osher Elhadad <Osher.Elhadad@ibm.com>
Osher Elhadad <Osher.Elhadad@ibm.com>
$ git log -1 --format=%B | grep -icE "co-authored|claude|generated with"
0
(0 = none)

15. #196's generator --check on the merged tree

Two trees, built identically except for this PR:

  • CONTROL /tmp/wt-ctl = main + #196
  • MERGED /tmp/wt-m = main + this PR + #182 + #196

The harbor.html / harbor-openshift.html failure is pre-existingharbor pages landed on main after #196 branched. Byte-identical message in both trees:

=== CONTROL (main + #196, my changes ABSENT) ===
$ scripts/sync-site-chrome.py --check
PAGES is out of step with site/: unregistered=['harbor-openshift.html', 'harbor.html'] missing=[]
register the page in scripts/sync-site-chrome.py (PAGES, and NAV_LINKS if it belongs in the nav), or add it to UNMANAGED
  EXIT=1

=== MERGED (main + #100 + #182 + #196, my changes PRESENT) ===
$ scripts/sync-site-chrome.py --check
PAGES is out of step with site/: unregistered=['harbor-openshift.html', 'harbor.html'] missing=[]
register the page in scripts/sync-site-chrome.py (PAGES, and NAV_LINKS if it belongs in the nav), or add it to UNMANAGED
  EXIT=1

Registering harbor* in UNMANAGED in both trees (so the pre-existing failure is out of the way) gives the same next message in both:

=== CONTROL ===                              === MERGED ===
site chrome out of sync: index.html,         site chrome out of sync: index.html,
getting-started.html, run-end-to-end.html,   getting-started.html, run-end-to-end.html,
results.html, benchmarks.html,               results.html, benchmarks.html,
architecture.html, optimize-your-own.html,   architecture.html, optimize-your-own.html,
adapter-templates.html                       adapter-templates.html
  EXIT=1                                       EXIT=1

Running the sync in both, then hashing each page's three chrome regions (chrome:head:*, chrome:nav:*, chrome:footer:*) — byte-identical, all 9 pages, so this PR causes zero chrome drift:

$ scripts/sync-site-chrome.py    # in both trees
$ # sha1 of the chrome regions, CONTROL vs MERGED
  IDENTICAL chrome: index.html                 (8d86c83d3bc1c1c1ae06baa1d8f580c47cb03b5e)
  IDENTICAL chrome: getting-started.html       (2b84e6c905d853e4cbc41adc80d01d02af808c62)
  IDENTICAL chrome: run-end-to-end.html        (7cd3fe30f8a3ceb509d4cfeb60d431324809df8a)
  IDENTICAL chrome: results.html               (dacc4505d47308a6bf3f0b3e81223c45cf74eabf)
  IDENTICAL chrome: benchmarks.html            (a628345ce9320e53de267177b5a39dd68641c8e6)
  IDENTICAL chrome: architecture.html          (6d98e185837b5ad07769df315de96bd41b727395)
  IDENTICAL chrome: optimize-your-own.html     (99415951a8e26b5abac858c779fb603e0ad6effb)
  IDENTICAL chrome: adapter-templates.html     (93b9d7c364b1e0015284bbfd24243efcd489898a)
  IDENTICAL chrome: agent-orchestration.html   (1653e6091a2052cd894a0630b9f2908ca2978849)

--check exits 0 on the merged tree:

$ cd /tmp/wt-m && scripts/sync-site-chrome.py --check
site chrome in sync (9 pages)
  EXIT=0

And both guard suites pass together on main + #100 + #182 + #196:

$ PYTHONPATH=/tmp/wt-m/core pytest core/tests/test_published_results_consistency.py core/tests/test_site_chrome_guard.py -q
...........                                                              [100%]
11 passed in 2.17s

16. #182 co-existence — actual test merge

Merging #182 into this branch conflicts in exactly two spots, both where my new sections sit adjacent to #182's stamped SkillsBench — skill-package optimization heading. Resolution is a union: keep my two sections, take #182's stamped heading.

$ git merge origin/fix/issue-99-sealed-headline
Auto-merging docs/RESULTS.md
CONFLICT (content): Merge conflict in docs/RESULTS.md
Auto-merging presentation/index.html        <- NO conflict (my slide-9 vs its slide-6)
Auto-merging site/index.html                <- NO conflict (I never touch it)
Auto-merging site/results.html
CONFLICT (content): Merge conflict in site/results.html

presentation/index.html auto-merges cleanly, confirming zero overlap with #182's presentation edits. After union resolution, all 9 sections are individually stamped on both surfaces (7 from #182, 2 from this PR) and the guard passes.

Merging #196 afterwards conflicts only in pure chrome (agent-orchestration.html nav aria-current, benchmarks.html ?v= hashes) — nothing of mine; taking --theirs and re-applying my one <main>-body sentence is the whole resolution.

@OsherElhadad

Copy link
Copy Markdown
Collaborator Author

🔍 Review — PR #241

Verdict: APPROVE WITH NITS.

I independently re-ran every claim. All 8 removals are correct — I could not substantiate a single one at any of the 599 commits on any branch. All 9 survivors I spot-checked hit their cited artifact field exactly, with correct split labels. The guard is non-vacuous and reads from disk; it also caught a real regression I accidentally introduced during the #196 merge test (below). Nothing blocking. Two non-blocking gaps: the guard's invariant is one-directional (it would not have caught the #236 divergence that this PR is fixing), and the surviving aggregate claim is falsifiable as written.


Blocking

None.


Non-blocking

1. core/tests/test_published_results_consistency.py:52-72 — the guard's invariant is one-directional, and the direction it misses is the one that actually broke.

The PR says "the divergence did not close, it swapped direction" — #236 added a docs/RESULTS.md section the site never got. The guard checks site → doc. It does not check doc → site. I deleted the site's whole skillsbench-87 section plus its TOC entry, leaving docs/RESULTS.md untouched — i.e. reintroduced exactly the #236 divergence this PR exists to fix — and all 3 tests passed:

=== PROBE G: delete the site's skillsbench-87 section + TOC entry, RESULTS.md intact ===
$ grep -c 'id="skillsbench-87"' site/results.html
0
$ pytest core/tests/test_published_results_consistency.py -q
...                                                                      [100%]
3 passed in 0.01s

Consequence: after merge, the next docs/RESULTS.md-only section (the #236 pattern, which has already happened once) re-opens #100 silently. The guard pins the original direction, not the current one.

Fix: one assertion, no allowlist. Add a SITE_OPTIONAL set for genuinely doc-only sections (agent-mode n=3 re-eval, head-to-head) and assert every other ## heading in docs/RESULTS.md has a site anchor. Since ANCHOR_TO_CANON_HEADING is already the explicit table, this is the reverse lookup over the same dict:

def test_every_canonical_run_section_is_published_on_the_site():
    headings = re.findall(r"^## (.+)$", _read("docs/RESULTS.md"), re.M)
    site_anchors = set(re.findall(r'<h2 id="([^"]+)"', _read("site/results.html")))
    mapped = {v for k, v in ANCHOR_TO_CANON_HEADING.items() if v and k in site_anchors}
    unpublished = [h for h in headings if not any(m in h for m in mapped)]
    assert not unpublished, (
        f"docs/RESULTS.md sections with no site/results.html section: {unpublished}\n"
        "Add the site section, or add the heading to SITE_OPTIONAL with a reason."
    )

2. docs/RESULTS.md:17-22 — the surviving aggregate claim is false as written, in two spots. Detail in its own section below.

3. presentation/index.html:1000 — row 4's 1.000 → 1.000 became , and that one deletion loses information the others don't. Rows 1-3 quoted a gain; row 4 quoted a null result (0%). A ceiling-hit is the one figure nobody fabricates for flattery, and the row's own status is still Done — so a reader now sees four Done rows with no numbers and cannot tell "we ran it, it saturated" from "we ran it, it improved and we're hiding it." Same evidentiary standing as the other three (no artifact), so removing it is defensible and consistent; flagging only because it is the one row where deletion costs more than it buys. Fix (optional): keep 1.000 → 1.000 with 0% and a provenance cell, or add a visible Δ not published — no artifact legend entry so the four rows read as deliberate rather than as missing data. Right now the explanation lives only in the <aside class="notes"> speaker note (presentation/index.html:1118-1123), which nobody in the audience sees.


Were the 8 removals right?

Searched, for each: git log --all -S, git grep at all 599 commits on all branches (incl. origin/benchmark-history, 37 records + 5843 live/ files), every run_full/, docs/runs/, ci/benchmarks/, and sources.bib.

# Number Searched where Substantiated? Remove vs ⚠️-label
1 71.1% (EvoSkills) git log --all -S'71.1%' → only the claim's own commit + #236 that wrote it. docs/sources.bib has no EvoSkills/EvoSkill entry (grep -in evoskill → 0 hits; 10 @misc entries, none of them). String appears nowhere else in the tree. No Remove correct. An uncited external number cannot be ⚠️-labelled — ⚠️ means "we ran it, artifact missing," which is a provenance claim. There is no run and no citation, so there is no provenance to stamp. Replacement text names arXiv:2604.01687v1, and I verified that id is present in the tree (presentation/index.html:1241,1421), so it isn't invented.
2 ~$32 (spend) grep -inE "cost|usd|spend|\$|budget" docs/runs/local-20260730-skillsbench-opus46-optimize.md → hits only Iterations (actual / cap): 3 / 3 and the artifact-path line. No cost field. git log --all -S'~$32' → only the claim's own commit. No total_usd for this run anywhere. No Remove correct. Notably the engine does record spend — runs_run_full.json has total_usd: 148.4339 and per-iteration optimizer_usd for the τ² run — so a real $32 would have left a trace in the same schema. It did not.
3 $400 cap grep -rn 400 examples/*/capevolve.yamlmax_optimizer_usd: 400.0 does exist in both examples/skillsbench/capevolve.yaml:49 and examples/tau2_airline/capevolve.yaml:54-55. ⚠️ Partly — the cap is real config Remove still correct. The cap alone is true, but the published sentence was ~$32 total (well below the $400 cap). The cap only carries meaning as the denominator of the unsubstantiated $32; quoting "spend was well below the $400 cap" with no spend figure asserts the comparison, which is the unbacked part. Removing the pair is right. Nit if you want it back: runs/...optimize.md could state the configured cap alone as a setting, not a result.
4 0.600 → 0.800 (slide-9 row 1) git log --all -S → hits are presentation/index.html itself + origin/benchmark-history live/30518857693__smoke-skillsbench dashboard blobs (a different, unrelated SkillsBench smoke run, matched only because a bare 0.600 appears in its JSON). No baseline_val: 0.6/best_val: 0.8 pairing in any of the 37 records/*.json. No Remove correct — see the joint note below.
5 0.533 → 0.933 (row 2) Same. Exhaustive: for all 599 commits, git grep -lF 0.933 -- '*.json' '*.jsonl' excluding /ui/, live/, records/zero files. 0.5333/0.9333 in origin/benchmark-history records/ + benchmarks.jsonzero. Confirms the orchestrator's finding: published at presentation/index.html:980, in zero artifacts. No Remove correct.
6 0.830 → 0.900 (row 3) Same. No Remove correct.
7 1.000 → 1.000 (row 4) Same. No Remove correct (but see Non-blocking 3 — the one row where deletion costs information).
8 "hand-verified" / "cross-check against the committed run artifacts" site/benchmarks.html:64, site/README.md:66-69. Not numbers — aggregate claims. False as written Remove correct, and this is the only one where the guard now enforces the removal (test 3).

Judgement on the presentation: deletion is right here, and the #192 PATCHES.md precedent does not apply. I read PATCHES.md (origin/fix/issue-120-cdn-font:examples/tau2_airline/run_full/PATCHES.md). It is a ledger attached to a committed artifact, and its whole force is the line "No result data has ever been modified — all JSON under ui/data/, final.json, and demo.cast are byte-identical to what the run produced." It documents a cosmetic edit to a directory a reader can open and diff. Slide-9 has the inverse shape: nothing to diff, no run record, no artifact at any commit — and the numbers sat under a speaker note asserting they were "already committed." A provenance note there would have to read "no artifact, no run record, no commit, origin unknown," which is not provenance, it is an admission that the figure cannot be traced. An untraceable number with a note attached is still an untraceable number on a slide someone screenshots. Deletion is the honest call, and keeping the row (with counts, Done status, for the numbers) preserves the true part — that the experiment happened — while dropping the part that cannot be defended. That is a better outcome than either reverting or annotating.

Cross-check that the removals actually landed. 71.1 → 0 hits in the entire tree at HEAD (present at origin/main:docs/RESULTS.md:235). $32, $400 cap, 0.933, 0.830, 0.533 → 0 on published surfaces. The PR's claim that residual 0.800/0.900/0.600 hits are Fira Sans weights and arXiv ids reproduced exactly (site/*.html:23 font URLs; docs/COMPARISON.md:58,67, README.md:114 = arXiv:2603.04900).


Survivor spot-checks

Nine checked, against the artifact field cited (not the PR's transcription of it).

Claim Cited source Found? Split label correct?
τ² sealed test 0.694 tau2_airline/run_full/final.jsontest.reward 0.6940000000000001 and this is the fix that matters — labeled fit — test == val, NOT held out, n per_task test = 50 confirms test == val == 50. This is the epic's original defect (a fit metric presented as sealed) and it is correctly labeled here.
τ² pass² 0.584 final.jsontest.pass_k["2"] 0.5844444444444444 ✅ fit (50), 10 trials
τ² cand_0007 final.jsonbest_id cand_0007 n/a
τ² baseline/best val 0.536/0.712 run_full/ui/data/runs_run_full.jsonsummary.baseline_val/best_val 0.536 / 0.7120000000000001; delta_abs 0.176, delta_pct 32.8, test_sealed True ✅ fit (50)
τ² stair + "5 of 10 accepted" summary.per_iteration + summary.evaluations ✅ recomputed from statuses: accepted: 5 of 10, stair [0.536, 0.582, 0.634, 0.67, 0.684, 0.712] — matches docs/RESULTS.md:59-60 exactly ✅ fit (50)
toy_calc 0.0 → 1.0 core/tests/test_e2e_slice.py assert base.reward == 0.0 (:56), assert payload["test"]["reward"] == 1.0 (:76) — genuinely re-derived, not transcribed ✅ sealed test
SkillsBench held-out val 0.333 skillsbench/run_full/baseline.jsonval.reward 0.33333333333333337, n=7 val (7, train==val)
SkillsBench optimized val 0.714 not in final.json — found in run_full/JOURNAL.md:85-86, history.jsonl, rejected.jsonl, events.jsonl (cand_0004: ACCEPTED val=0.714) ✅ (in-artifact, different file than a reader would guess) ✅ val fit (7)
SkillsBench sealed test 0.556 → 0.667 final.jsontest_baseline.reward / test.reward 0.5555555555555555 / 0.6666666666666666, n=3, best_id cand_0004, test_delta 0.111… sealed test (3), genuinely held out — separate test_baseline field proves both were scored on the sealed split

Plus the ⚠️ rows, recomputed rather than trusted:

  • 87-task baselines: docs/RESULTS.md:197-198 0.281 ± 0.048 / 23 (26.4%) and 0.040 ± 0.019 / 2 (2.3%) vs run records 0.2809 ± 0.0475 / 23/87 (26.4%) and 0.0396 ± 0.0191 / 2/87 (2.3%). ✅ correct 3-dp rounding, correctly labeled fit (87).
  • 87-task task-level delta: recomputed the passing-task sets from both run records' - ✓ lists: baseline 23, optimized 28, net +5; newly-passing and regressed lists identical to docs/RESULTS.md:252-259, all 11 task names. ✅ genuinely derived.
  • 87-task iterations: 0.357 / 0.325 / 0.170 vs run record 0.3574 / 0.3247 / 0.1703. ✅

Five τ² runs with no committed artifact — confirmed, exhaustively. For each of 0.475, 0.387, 0.240, 0.270, at all 599 commits, git grep -lF restricted to final.json|baseline.json|splits?.json|split_ids|run_full and excluding /ui/:

--- 0.475 ---  (none)
--- 0.387 ---  (none)
--- 0.240 ---  (none)
--- 0.270 ---  (none)

And no Qwen file was ever added: git log --all --diff-filter=A --name-only | grep -i qwen → empty. Keeping these ⚠️ rather than deleting is correct and consistent with #99/#182's reviewer decision.

Split labels — all correct, including the two the epic previously caught. The 30/10/1030(=val)/20 correction is applied on both surfaces (docs/RESULTS.md:69,90; site/results.html:168,208), and the units are now unified: 0.567 (56.7%) / 0.300 (30.0%) in the doc matches site/results.html:172-173 byte-for-byte. That format is applied consistently to the Qwen sections too (:151-152, :173-174), not just the one section the fix targeted.


Guard sufficiency

Reads disk, not a cache. REPO = Path(__file__).resolve().parents[2], _read() is (REPO / rel).read_text(). No manifest, no fixture, no derived copy — unlike #189 (read a committed manifest instead of disk) and #213 (parser couldn't read block scalars, so checks passed vacuously). Confirmed by mutation: every probe below changed only the file on disk and the guard's verdict tracked it.

Non-vacuous — fails before, passes after. Reintroduced the original #100 bug (renamed the canonical qwen-tools heading):

$ pytest core/tests/test_published_results_consistency.py -q
E       assert not {'qwen-tools': 'Qwen 2.5 14B (tools only'}
core/tests/test_published_results_consistency.py:70: AssertionError
1 failed, 2 passed in 0.02s
$ # restored
...                                                                      [100%]
3 passed in 0.01s

Test 2 also bites — removing one TOC entry fails test_site_results_toc_matches_its_sections.

What it catches: a site section whose canonical heading is deleted or renamed (the explicit table beats a set-difference here — a rename fails loudly); a new site anchor with no canonical mapping; TOC/section drift; the two blanket claims on the surfaces it owns. Test 3 is not decoration: during my #196 merge test I resolved site/benchmarks.html with --theirs and silently lost this PR's <main> edit, restoring hand-verified. The guard caught my mistake:

E  assert not {'site/benchmarks.html': ['67: See the hand-verified canonical numbers on the <a href="results.html">Results</a> page; this']}
1 failed, 2 passed in 0.02s

That is a real merge-regression class, not a hypothetical.

What it misses (all six probed, all 3 tests passed):

Probe Mutation Result
B 0.281 ± 0.0480.999 ± 0.048 on the site ✅ passed — number edited in place, undetected
C val (87, fit)val (87, held-out) on the site ✅ passed — split label falsified, undetected
D site/index.html gains a fake FakeBench sealed test 0.111 → 0.999 ✅ passed — new surface uncovered
E new file site/results2.html publishes <h2 id="totally-new-run"> ✅ passed — only site/results.html is read
F delete a ⚠️ reported, artifact not committed marker from a site heading ✅ passed — marker parity unenforced
G delete the site's skillsbench-87 section, RESULTS.md intact ✅ passed — the #236 direction; Non-blocking 1

Verdict on dropping the numeric scrape: correct, and I'd have made the same call. The stated reasons hold up under inspection — site/results.html:23 really is wght@400;500;600;700;800, docs/COMPARISON.md:58 really does carry 2603.04900, 35.9, 39.1, 14.4, 15.7, and all three HTML files are dense with SVG path coordinates (I hit them parsing: 19 self-closing SVG elements in results.html alone). A whole-file scrape would need an allowlist entry per font URL, arXiv id, external figure and SVG glyph, and "a guard whose allowlist keeps growing stops guarding" is the right principle — this repo has already shipped two guards that passed vacuously.

But the reason the scrape fails is scope, not the idea. Take the middle path you already named. Numbers inside a declared results table are unambiguous — no fonts, no SVG, no arXiv ids live in <td class="num">. A scrape confined to <td class="num"> / <td class="gain"> within <section>s that have an <h2 id> in ANCHOR_TO_CANON_HEADING, compared against the matching docs/RESULTS.md table rows, needs zero allowlist and closes probes B and C — the two most likely future failures, because editing a published number is a far more common edit than deleting a section. I'd take this over Non-blocking 1 if only one lands; it and the reverse-direction check are the two highest-value additions. Neither blocks: section identity is a real invariant, mechanically enforced, and strictly better than the nothing that shipped before.


The surviving aggregate claim

docs/RESULTS.md:17-22 (:25 on the merged tree):

This page is the superset. Every result claim on any other surface — README.md, site/results.html, site/index.html, docs/COMPARISON.md, OUTREACH.md, presentation/ — must correspond to a section here, with the same value, the same split label, and the same evidence marker. If a number is not on this page, it is not published.

The PR's defence is right about the reading, wrong about the fact. I agree a reader parses "must correspond" as normative, not as "all numbers are artifact-backed" — the sentence never says verified, and it sits directly above a paragraph saying uncommitted runs are ⚠️-labelled. That distinguishes it cleanly from the claims #182 deletes, and it does not rot the way "every number ships its artifact" does. I also confirmed the surrounding sweep: on the merged #182 + #241 tree this is the only hit for the aggregate-claim regex across all 8 surfaces, and the genuinely false line (docs/RESULTS.md:3 / site/results.html:55, "Every number here is derived from a committed run artifact") is rewritten by #182 into per-marker prose — correctly left to it, verified in #182's diff.

But it is false as written on the merged tree, in two places I can point at:

  1. docs/COMPARISON.md:58-60 publishes EvoTool 35.9 → 39.1, 14.4 → 15.7, and Evolutionary Context Search +23.3%. docs/COMPARISON.md is named in the list. There is no EvoTool or ECS section in docs/RESULTS.md (grep → only a prose mention at :103). Ironically the ECS row is the weakest number in the repo — self-labelled "its primary-source citation is still to be confirmed" — so the one surface the rule most needs to cover is the one that falsifies it.
  2. presentation/index.html:1206-1207 publishes query_babylon_catalog errors on 43.2% of calls, query_aap2 on 17.2%. presentation/ is named in the list. Neither is in docs/RESULTS.md.

(A third, softer one: site/index.html:116 renders 53.6%71.2% in the hero caption; docs/RESULTS.md carries 0.536/0.712. Same result, different units — the doc's own "label percentages explicitly" rule makes this defensible, unlike the two above.)

Consequence: a reader who tests the rule finds a counter-example in one grep, on a page whose entire premise is that its claims survive checking. On an honesty PR that is the expensive kind of wrong. And note the guard does not enforce this scope — it reads only site/results.html, docs/RESULTS.md, site/README.md, site/benchmarks.html. So "mechanically enforced" is true of the site-results ↔ doc mapping only, not of the six-surface sentence.

Fix — smallest edit that makes it true, scope it to what it governs:

**This page is the superset for cap-evolve's own runs.** Every result claim about a
cap-evolve run on any other surface — `README.md`, `site/results.html`,
`site/index.html`, `OUTREACH.md`, `presentation/` — must correspond to a section here,
with the same value, the same split label, and the same evidence marker. If one of our
numbers is not on this page, it is not published. Externally reported results (other
papers' figures) live in [`COMPARISON.md`](COMPARISON.md) with their own citations and
are out of scope for this rule.

That covers 43.2%/17.2% too (they are Parsec tool-API error rates, not cap-evolve run results, so "result claim about a cap-evolve run" excludes them cleanly), and it drops docs/COMPARISON.md from the list with a stated reason rather than by omission. One paragraph, no code, no sibling coordination.


Coordination boundaries — did anything left behind contradict what changed?

Verified untouched (git diff --stat origin/main shows exactly 6 files): README.md, CHANGELOG.md, site/index.html, llms.txt, docs/COMPARISON.md, OUTREACH.md, site/benchmarks.js, site/benchmarks.fixture.json, all counts and host labels. No stray edits.

Contradiction sweep on the merged #182 + #241 tree — the more important half:

presentation/index.html auto-merged with no conflict against #182, confirming the claimed zero overlap (slide 9 vs slide 6).


#196 compatibility

Built main + #241 + #182 + #196 and a control main + #196 (no #241), and drove both to exit 0. Both claims reproduce.

Harbor failure — byte-identical in both trees, so pre-existing:

=== MERGED (main + #241 + #182 + #196) ===        === CONTROL (main + #196, no #241) ===
PAGES is out of step with site/:                  PAGES is out of step with site/:
unregistered=['harbor-openshift.html',            unregistered=['harbor-openshift.html',
'harbor.html'] missing=[]                         'harbor.html'] missing=[]
EXIT=1                                            EXIT=1

With harbor* in UNMANAGED and the chrome conflicts resolved in favour of #196 in both:

=== MERGED ===                                    === CONTROL ===
$ ./scripts/sync-site-chrome.py && --check        $ ./scripts/sync-site-chrome.py && --check
site chrome in sync (9 pages)                     site chrome in sync (9 pages)
MERGED EXIT=0                                     CONTROL EXIT=0

Edits confined to <main>: the #196 conflicts were only site/benchmarks.html ?v= hashes and site/agent-orchestration.html nav aria-current — zero overlap with this PR's content. After running the generator on the merged tree, this PR's <main> edits survive intact (site/benchmarks.html:68 evidence-marker sentence present; site/results.html id="skillsbench-87" present) and the guard still passes. ✅ Confirmed the <head> is untouched, which matters because #196 makes it generator-owned via chrome:head:* sentinels.


docs/RESULTS.md structural fixes

Dangling reference — fixed, and it was genuinely dangling. On origin/main the 87-task optimization section said "the same four shared office-document skills as the baseline section above" with no such section anywhere in the file — a reference to a section never written. Now docs/RESULTS.md:214 reads [the baselines section above](#skillsbench-87-baselines) and the target exists (<a id="skillsbench-87-baselines"> at :186). The fix is structural: it created the missing section rather than just softening the wording.

Units — unified. docs/RESULTS.md:76-77 now 0.567 (56.7%) / 0.300 (30.0%), byte-identical to site/results.html:172-173. Applied consistently to the Qwen sections too (:151-152, :173-174).

No other dangling cross-reference survives. Swept section above|below, (above), (below), table above, described above: 2 hits, both at :214 and :277, both anchored to #skillsbench-87-baselines, both resolve.


Links / anchors

docs/RESULTS.md16 resolve, 0 broken, matching the claim exactly (fragments checked against target-file headings, not just file existence):

OK   runs/ ・ ../examples/tau2_airline/run_full/ ・ REPRODUCE_tau2.md ・ OPTIMIZATION_EXAMPLES.md
OK   ../examples/tau2_airline/DEMO.md ・ COMPARISON.md ・ AGENT_ORCHESTRATION.md
OK   REPRODUCE_tau2.md#8-agent-mode-reproduction-held-out-3020-litellm-proxy
OK   runs/local-20260729-skillsbench-baseline-opus46.md ・ runs/local-20260729-skillsbench-baseline-gptoss120b.md
OK   #skillsbench-87-baselines ・ runs/local-20260730-skillsbench-opus46-optimize.md
OK   #skillsbench-87-baselines ・ sources.bib ・ ../examples/skillsbench/run_full/ ・ REPRODUCE_skillsbench.md
-> 16 resolve, 0 broken

site/results.html10/10 anchors (<h2 id> count = 10), 20 of 21 hrefs resolve. One nit, pre-existing and not this PR's: style.css?v=20260724b fails a naive existence check because of the cache-busting query. site/style.css exists; my checker didn't strip ?. Not a defect. Effective: 21/21.

All three HTML files parse — unclosed=none on site/results.html, site/benchmarks.html, presentation/index.html. The "mismatched" entries my parser reports are all self-closing SVG elements (path, circle, line, rect, polygon), which html.parser doesn't know are void — an artifact of my checker, present identically on origin/main. site/benchmarks.fixture.json untouched and valid.


Merge-order note

Concur with the PR, with one strengthening.

#189 (counts) ─┐
#184 (changelog)─┤ independent, land any time
#182 (#99) ──► #241 (#100) ──► #202 (#126) / #192 (#120) ──► #196 (#123) LAST

Verification I re-ran

Full suite (181 passed = 178 baseline + 3 new; sole failure is the #200 port flake, 7888 vs expected — pre-existing, unrelated):

$ PYTHONPATH=/tmp/rv-241/core /tmp/ce-venv/bin/python -m pytest core/tests -q
core/tests/test_dashboard_launch.py:56: AssertionError
FAILED core/tests/test_dashboard_launch.py::test_maybe_launch_spawns_when_available
1 failed, 181 passed in 63.72s (0:01:03)

$ /tmp/ce-venv/bin/python -m compileall -q core skills
EXIT=0

Guard fail-before / pass-after:

=== reintroduce the #100 bug (rename the canonical qwen-tools heading) ===
E       assert not {'qwen-tools': 'Qwen 2.5 14B (tools only'}
1 failed, 2 passed in 0.02s
=== restored ===
3 passed in 0.01s

Guard blind spots (all 3 tests passed under each mutation):

=== PROBE B: 0.281 -> 0.999 on the site ===                   3 passed
=== PROBE C: val (87, fit) -> val (87, held-out) ===           3 passed
=== PROBE D: site/index.html gains a FakeBench row ===         3 passed
=== PROBE E: new site/results2.html publishes a run ===        3 passed
=== PROBE F: ⚠️ marker deleted from a site heading ===          3 passed
=== PROBE G: site's skillsbench-87 section deleted ===         3 passed   <- the #236 direction

Exhaustive artifact search, all 599 commits, all branches:

=== 0.475 / 0.387 / 0.240 / 0.270 in final.json|baseline.json|splits|split_ids|run_full ===
--- 0.475 ---  (none)
--- 0.387 ---  (none)
--- 0.240 ---  (none)
--- 0.270 ---  (none)
$ git log --all --diff-filter=A --name-only | grep -i qwen
(empty — no Qwen file ever added)
$ for c in $(git rev-list --all); do git grep -lF 0.933 $c -- '*.json' '*.jsonl'; done | grep -v 'ui/\|live/\|records/'
(empty — 0.933 in zero artifacts at any commit)
$ grep -in evoskill docs/sources.bib
(none)
$ grep -inE "cost|usd|spend|\$" docs/runs/local-20260730-skillsbench-opus46-optimize.md
16:- **Iterations (actual / cap):** 3 / 3  ·  1 accepted     <- no cost field

Survivors, read out of the artifacts:

tau2 final.json test.reward = 0.6940000000000001    test.pass_k = {'1': 0.694…, '2': 0.5844444444444444}
tau2 best_id = cand_0007                             n per_task test = 50   (test == val == 50)
tau2 summary: baseline_val 0.536  best_val 0.7120000000000001  delta_abs 0.176  delta_pct 32.8  test_sealed True
tau2 accepted: 5 of 10   stair [0.536, 0.582, 0.634, 0.67, 0.684, 0.712]
sb  baseline val.reward = 0.33333333333333337 n=7
sb  final test_baseline.reward = 0.5555555555555555   test.reward = 0.6666666666666666 n=3   best_id cand_0004
sb  JOURNAL.md:86  <!-- cand_0004: ACCEPTED val=0.714 Δ=+0.095 -->
toy core/tests/test_e2e_slice.py:56  assert base.reward == 0.0
toy core/tests/test_e2e_slice.py:76  assert payload["test"]["reward"] == 1.0
87  baseline passing: 23  optimized passing: 28  net 5
    newly: bike-rebalance, energy-ac-optimal-power-flow, energy-market-pricing, exceltable-in-ppt,
           grid-dispatch-operator, paper-anonymizer, paratransit-routing, weighted-gdp-calc
    regressed: citation-check, crystallographic-wyckoff-position-analysis, pptx-reference-formatting
    -> both lists identical to docs/RESULTS.md:252-259

Merged-tree --check, both trees:

MERGED (main + #241 + #182 + #196):  site chrome in sync (9 pages)   EXIT=0
CONTROL (main + #196, no #241):      site chrome in sync (9 pages)   EXIT=0
harbor failure before UNMANAGED:     byte-identical message, EXIT=1 in both
after generator run on MERGED:  site/benchmarks.html:68 evidence sentence intact,
                                site/results.html id="skillsbench-87" intact, guard 3 passed

Aggregate-claim sweep on the merged tree — 1 hit, and the two counter-examples to it:

$ grep -rniE "every number|every result|all results|cross-check|ships its artifact|committed run artifact|hand-verified|every figure" \
    README.md docs/RESULTS.md docs/COMPARISON.md OUTREACH.md llms.txt site/*.html site/README.md presentation/index.html
docs/RESULTS.md:25:**This page is the superset.** Every result claim on any other surface — `README.md`,

$ grep -n "EvoTool\|Evolutionary Context" docs/RESULTS.md
103:gain sits next to external tool-optimization work (EvoTool, Evolutionary Context Search)
   -> prose mention only; docs/COMPARISON.md:58-60 publishes 35.9→39.1, 14.4→15.7, +23.3% with no section here
$ grep -n "43.2\|17.2" docs/RESULTS.md
NOT in RESULTS.md
   -> presentation/index.html:1206-1207 publishes 43.2% / 17.2% with no section here

Links / dangling refs:

docs/RESULTS.md:    16 resolve, 0 broken
site/results.html:  10/10 <h2 id> anchors, 21/21 hrefs (style.css?v= is a query-string artifact of my checker)
dangling "section above/below" sweep: 2 hits, :214 and :277, both -> #skillsbench-87-baselines, both resolve
HTML parse: site/results.html, site/benchmarks.html, presentation/index.html -> unclosed=none
site/benchmarks.fixture.json: valid JSON, untouched

…rs in mapped tables

Review fixes for PR #241.

1. The guard's invariant was one-directional. It checked site -> doc, but the
   divergence this PR fixes (#236) ran doc -> site: a docs/RESULTS.md section the
   site never got. Deleting the site's skillsbench-87 section with the doc intact
   passed all 3 tests. Adds the reverse lookup over the existing mapping, with an
   explicit (currently empty) SITE_OPTIONAL escape hatch so a doc-only section has
   to be declared rather than appear silently.

2. Adds a numeric scrape confined to <td class="num"> / <td class="gain"> inside
   mapped sections, matched row-by-row against the same-labelled docs/RESULTS.md
   row. Editing a published number is a far more likely edit than deleting a
   section, and both were undetected. The confinement is what keeps the no-growing-
   allowlist principle: fonts, arXiv ids and SVG coordinates never live inside a
   results <td>, so this needs zero allowlist entries. Known limits (a new surface,
   a new results page, evidence-marker parity) are documented as limits rather than
   implied as coverage.

   That scrape immediately found a real divergence: site/results.html labelled the
   agent-mode row "sealed test (20, held-out)" where the canonical doc says
   "(20, held-out, scored once)". Fixed on the site, per the superset rule.

3. The superset claim was false as written: it named docs/COMPARISON.md and
   presentation/, both of which publish figures with no RESULTS.md section.
   Rewritten to scope it to cap-evolve's own run results, to say which part is
   mechanically enforced (the site-results <-> doc mapping, not the six-surface
   sentence), and to place external figures in COMPARISON.md explicitly rather than
   by omission. Reconciles COMPARISON.md's one in-scope row too: its cap-evolve
   held-out row now carries the same evidence marker its RESULTS.md section does.

4. Slide 9's row-4 null result was a real finding whose explanation lived only in a
   speaker note. The reasoning is now visible on the slide: what "-" means, and that
   row 4 was already at the scorer's ceiling on 5 easy tasks, so optimization had no
   headroom. A saturated benchmark stated in words, not a bare dash a reader can
   mistake for withheld data.
Copilot AI review requested due to automatic review settings July 30, 2026 15:35

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@OsherElhadad

Copy link
Copy Markdown
Collaborator Author

🔧 Review fixes

All 3 non-blocking findings fixed. New commit e5c33332 (not amended). Full suite 183 passed (181 + 2 new tests), 1 pre-existing #200 port flake.

Guard probes

Probe Mutation Caught before? Caught now?
A reintroduce the original #100 bug (rename the canonical qwen-tools heading) ✅ yes (1 failed) ✅ yes (2 failed — both directions now bite)
B 0.281 ± 0.0480.999 ± 0.048 on the site (number edited in place) no — 3 passed yestest_site_results_numbers_match_the_canonical_doc
C val (87, fit)val (87, held-out) on the site (split label falsified) no — 3 passed yes — same test, "no row labelled …"
D site/index.html gains a fake FakeBench 0.111 → 0.999 (new surface) ❌ no no — documented limit
E new site/results2.html publishes <h2 id="totally-new-run"> (new results page) ❌ no no — documented limit
F delete a ⚠️ reported, artifact not committed marker from a site heading ❌ no no — documented limit
G delete the site's skillsbench-87 section + TOC entry, RESULTS.md intact (the #236 direction) no — 3 passed yestest_every_canonical_run_section_is_published_on_the_site

D/E/F are now stated as limits in the module docstring rather than left as implied coverage — they are section-registration problems, not consistency problems, and covering them means either a new-file watcher or an emoji-parity rule that would need the growing allowlist we're avoiding.

Literal output — before (on 0c0febf9):

=== PROBE G (before): delete site skillsbench-87 section AND its TOC anchor ===
site anchors: 0
doc still has heading: 1
...                                                                      [100%]
3 passed in 0.01s

=== PROBE B (before): 0.281 ± 0.048 -> 0.999 ± 0.048 on the site ===
3 passed in 0.01s

=== PROBE C (before): val (87, fit) -> val (87, held-out) on the site ===
3 passed in 0.01s

After (e5c33332):

=== PROBE A: reintroduce the original #100 bug (rename canonical qwen-tools heading) ===
0
FAILED core/tests/test_published_results_consistency.py::test_every_canonical_run_section_is_published_on_the_site
2 failed, 3 passed in 0.02s

=== PROBE B: 0.281 ± 0.048 -> 0.999 ± 0.048 on the site (number edited in place) ===
FAILED core/tests/test_published_results_consistency.py::test_site_results_numbers_match_the_canonical_doc
1 failed, 4 passed in 0.03s

=== PROBE C: val (87, fit) -> val (87, held-out) on the site (split label falsified) ===
FAILED core/tests/test_published_results_consistency.py::test_site_results_numbers_match_the_canonical_doc
1 failed, 4 passed in 0.02s

=== PROBE D: site/index.html gains a fake FakeBench row (new surface) ===
5 passed in 0.01s
=== PROBE E: new site/results2.html publishes a run (new results page) ===
5 passed in 0.01s
=== PROBE F: delete a ⚠️ evidence marker from a site heading ===
5 passed in 0.01s

=== PROBE G: delete site skillsbench-87 section + TOC entry, RESULTS.md intact ===
site anchors for skillsbench-87: 0
doc heading intact: 1
FAILED core/tests/test_published_results_consistency.py::test_every_canonical_run_section_is_published_on_the_site
1 failed, 4 passed in 0.02s

=== restored ===
5 passed in 0.01s

Finding 1b — the confined scrape needs zero allowlist entries, and found a real bug

Implemented as you described: <td class="num"> / <td class="gain"> inside a <section> whose <h2 id> is in ANCHOR_TO_CANON_HEADING, matched to the docs/RESULTS.md table row with the same label and the same column count.

Two normalisation rules do the work an allowlist would otherwise do, and neither is an exception list:

  1. Percent restatements are dropped. The site writes 0.536 (53.6%) where the doc writes 0.536. _numbers() drops any value that is exactly 100× another value in the same cell — which is the doc's own "label percentages explicitly" rule expressed mechanically, not a per-number carve-out.
  2. Only num/gain columns are compared, positionally. That is what excludes the doc's run record link column (local-20260729-… parses as 20260729) without naming it.

Proof — zero false positives across all 9 mapped sections, zero allowlist entries:

$ python /tmp/align2.py
CLEAN: zero mismatches, zero allowlist entries

It immediately caught a real divergence in this PR's own branch, which is the best argument for it. site/results.html:213 labelled the agent-mode row:

sealed test (20, held-out)                    <- site
sealed test (20, held-out, scored once)       <- docs/RESULTS.md:85

Fixed on the site, per the superset rule (the doc is the superset, so the doc's label wins). That is exactly probe C's failure mode occurring for real, undetected until the scrape existed.

Finding 2 — the rewritten aggregate sentence

docs/RESULTS.md:17-27:

This page is the superset for cap-evolve's own runs. Every result claim about a cap-evolve run on any other surface — README.md, site/results.html, site/index.html, OUTREACH.md, presentation/ — must correspond to a section here, with the same value, the same split label, and the same evidence marker. If one of our numbers is not on this page, it is not published. Externally reported results (other papers' figures) live in COMPARISON.md with their own citations and are out of scope for this rule; measurements that are not cap-evolve run results (e.g. a partner system's tool-API error rates) are out of scope too. The site/results.html ↔ this-page mapping is the part that is mechanically enforced, by core/tests/test_published_results_consistency.py (section identity in both directions, plus the numbers in every mapped results table); the other surfaces are convention.

Your paragraph, plus two additions: the tool-API-error-rates clause (so 43.2%/17.2% are excluded by a stated reason rather than by the reader inferring that "cap-evolve run result" excludes them), and the last sentence, because "mechanically enforced" was the second half of the finding — it is now attached to the mapping it is true of, and the other surfaces are called convention.

Grep proof. Every X → Y result claim on each of the five named surfaces resolves to a number docs/RESULTS.md publishes:

$ python /tmp/prove_agg.py
CLEAN: every X -> Y result claim on every named surface resolves to a number published in
docs/RESULTS.md (0 counter-examples)

Same on the merged #182 + #241 tree, and again on + #196. Your two counter-examples:

presentation/index.html:1206-1207   43.2% / 17.2%  -> Parsec tool-API error rates, not a
                                       cap-evolve run result; excluded by the stated clause
docs/COMPARISON.md                  -> no longer in the list; excluded with a stated reason

docs/COMPARISON.md — fixed here, not handed off

You noted it is the only unreconciled surface on the merged tree. Rather than only exempting it, I reconciled the part that is in scope. Its one cap-evolve row was already numerically correct (30.0 → 47.5, +17.5 pp, +58.3% — all in docs/RESULTS.md), but it was missing the evidence marker the rule requires, so it was in violation of "the same evidence marker" even under the narrowed sentence:

- | **cap-evolve** … | sealed test 30.0 → **47.5** (+17.5 pp) | **+58.3%** |
+ | **cap-evolve** … | sealed test 30.0 → **47.5** (+17.5 pp) — ⚠️ reported, artifact not committed | **+58.3%** |

Plus two caveat bullets: one linking that ⚠️ to its RESULTS.md section, one stating that the EvoTool/ECS rows are externally reported and therefore out of scope. So the exemption is symmetric — the doc says the rows are out of scope, and the page itself says the same thing. Nothing left to hand off.

Finding 3 — the null result is restored, on the slide

Agreed: a null result is a real finding, and it was only in the speaker note. Row 4's 1.000 → 1.000 stays deleted (it has no artifact, same as rows 1–3, and re-publishing an untraceable figure is what this PR is removing) — but the finding is now visible, which is what was actually lost. Added under the slide-9 legend:

in the Δ columns means the run happened but has no artifact and no run record in this repo, so no figure is quoted; it does not mean the result was withheld. One of the four is a null result: row 4 (5 easy τ²-bench tasks) was already at the scorer's ceiling and optimization changed nothing — a saturated subset, not a hidden gain. Every published figure is on slide 6 and in docs/RESULTS.md.

That fixes the reading you identified: four Done rows with four no longer read as missing data, and "we ran it, it saturated" is now distinguishable from "we ran it, it improved and we're hiding it" — without putting a number back that no artifact supports. The speaker note now points at the slide instead of being the only place the reasoning lives.

⚠️ Warning for whoever resolves #196

Resolving site/benchmarks.html with --theirs silently reverts this PR's <main> edit and restores hand-verified. sync-site-chrome.py --check will not tell you. Reproduced your exact mistake on the merged tree:

=== WRONG resolution (--theirs wholesale on benchmarks.html) ===
site/benchmarks.html:67:  See the hand-verified canonical numbers on the <a href="results.html">Results</a> page; this

-- sync-site-chrome.py --check says: --
site chrome in sync                                        CHECK EXIT=0     <- passes!
-- the new guard says: --
FAILED …::test_no_blanket_artifact_claim_on_the_surfaces_this_guard_owns
1 failed, 4 passed in 0.02s

After resolving #196, run pytest core/tests/test_published_results_consistency.py, not just --check. --check compares generator-owned chrome; the reverted sentence is in <main>, which the generator does not own, so --check is structurally incapable of seeing it.

Correct resolution: --theirs for the chrome (?v= hashes, nav aria-current), then re-apply this PR's one <main> sentence. On the correctly-resolved tree:

=== MERGED (main + #241 + #182 + #196) ===
site chrome in sync (9 pages)                              CHECK EXIT=0
#241 <main> sentence survives: 1
site/results.html id="skillsbench-87" survives: 1
5 passed in 0.01s
CLEAN: zero mismatches, zero allowlist entries
CLEAN: every X -> Y result claim on every named surface resolves … (0 counter-examples)

(#196 also still needs harbor.html / harbor-openshift.html in PAGES or UNMANAGED — pre-existing, fails identically on main + #196 without #241.)

Merge order — upgraded as you requested

#189 (counts) ─┐
#184 (changelog)─┤ independent, land any time
#182 (#99) ──► #241 (#100) ──► #202 (#126) / #192 (#120) ──► #196 (#123) LAST

#182 is REQUIRED before #241, not preferred. Adopting your reasoning: landing #241 first ships a window in which its own superset rule is contradicted by two lines only #182 fixes. Verified both are gone on the merged tree:

=== OUTREACH.md 30/10/10 gone on merged tree? ===
(none — #182 fixed it)
=== docs/RESULTS.md:3 'committed run artifact' gone? ===
(none — #182 rewrote it)

The reverse order has no such window. Merge is mechanical: 2 conflicts (docs/RESULTS.md, site/results.html), both a union where this PR's new sections abut #182's stamped SkillsBench — skill-package optimization heading. presentation/index.html still auto-merges with no conflict even with the new slide-9 legend note — confirming zero overlap with #182's slide-6 edits.

Verification

$ PYTHONPATH=/tmp/fx-241/core /tmp/ce-venv/bin/python -m pytest core/tests -q
FAILED core/tests/test_dashboard_launch.py::test_maybe_launch_spawns_when_available
1 failed, 183 passed in 72.91s (0:01:12)

  -> 183 = 181 baseline + 2 new. Sole failure is the #200 port flake, unchanged:
     AssertionError: assert 'http://127.0.0.1:7888' == 'http://127.0.0.1:7878'

$ /tmp/ce-venv/bin/python -m compileall -q core skills
EXIT=0

$ links / anchors
docs/RESULTS.md + docs/COMPARISON.md:  21 resolve, 0 broken
  (incl. the new COMPARISON.md -> RESULTS.md#τ²-bench-airline--held-out-30val20-run fragment)
site/results.html:  h2 anchors=10  toc=10  match=True

$ HTML parse (void/SVG-aware)
site/results.html:        unclosed=none  stray_end=none
site/benchmarks.html:     unclosed=none  stray_end=none
presentation/index.html:  unclosed=none  stray_end=none
site/index.html:          unclosed=none  stray_end=none

Files touched

File Change
core/tests/test_published_results_consistency.py +2 tests (reverse direction, confined numeric scrape), SITE_OPTIONAL, 5 helpers, docstring rewrite incl. stated limits
site/results.html sealed test (20, held-out)(20, held-out, scored once) — divergence found by the new scrape
docs/RESULTS.md superset paragraph rewritten (scoped + enforcement scope stated)
docs/COMPARISON.md cap-evolve row gains its ⚠️ marker; 2 caveat bullets (⚠️ provenance, external-rows scope)
presentation/index.html slide-9 .exp-note explaining and row 4's null result; .exp-note CSS; speaker note updated

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Reconcile results across README, docs/RESULTS.md, and the site (Qwen runs live only on the site)

3 participants