Hosted swe-bench_verified test run has 0 completed rows after accepted upload
Hello SWE-bench team,
We submitted a swe-bench_verified test run via sb-cli. The hosted API
accepted the uploaded predictions, but the final report has zero completed
evaluation rows.
Run details
- Subset:
swe-bench_verified
- Split:
test
- Run id:
tris_codex_helper_verified_20260629T002207Z
- Prediction rows uploaded:
500
- CLI upload result:
500 new predictions uploaded - these cannot be changed
Hosted report summary
- Total instances:
500
- Submitted instances:
478
- Completed/successful evaluation runs:
0
- Failed runs:
478
- Resolved:
0
- Unresolved:
0
- Errors:
0
- Incomplete:
22
This does not look like ordinary model/test failure because the report contains
no completed evaluation rows. A normal patch miss would usually appear as a
completed-but-unresolved row.
Local official-harness sanity check
To check whether the submitted JSONL itself was malformed, we ran the exact
same prediction file through the official local SWE-bench harness on a small
subset.
One-row diagnostic:
astropy__astropy-12907: 1 submitted / 1 completed / 1 resolved / 0 errors
Mixed three-row diagnostic:
astropy__astropy-12907
psf__requests-1921
pydata__xarray-6992
Result:
3 completed / 3 resolved / 0 unresolved / 0 errors
So the same submitted prediction file can produce completed/resolved rows under
the official local harness.
Related issue shape
This appears similar to:
Could you inspect or re-trigger the hosted evaluation containers for run id
tris_codex_helper_verified_20260629T002207Z?
Thank you.
Hosted
swe-bench_verified testrun has 0 completed rows after accepted uploadHello SWE-bench team,
We submitted a
swe-bench_verified testrun viasb-cli. The hosted APIaccepted the uploaded predictions, but the final report has zero completed
evaluation rows.
Run details
swe-bench_verifiedtesttris_codex_helper_verified_20260629T002207Z500500 new predictions uploaded - these cannot be changedHosted report summary
500478047800022This does not look like ordinary model/test failure because the report contains
no completed evaluation rows. A normal patch miss would usually appear as a
completed-but-unresolved row.
Local official-harness sanity check
To check whether the submitted JSONL itself was malformed, we ran the exact
same prediction file through the official local SWE-bench harness on a small
subset.
One-row diagnostic:
astropy__astropy-12907:1 submitted / 1 completed / 1 resolved / 0 errorsMixed three-row diagnostic:
astropy__astropy-12907psf__requests-1921pydata__xarray-6992Result:
3 completed / 3 resolved / 0 unresolved / 0 errorsSo the same submitted prediction file can produce completed/resolved rows under
the official local harness.
Related issue shape
This appears similar to:
Could you inspect or re-trigger the hosted evaluation containers for run id
tris_codex_helper_verified_20260629T002207Z?Thank you.