Showing that you can take one of the top models available and push it across another top model is neat, but that doesn't really tell us much. Showing how good a wide range of models do with your harness would be far more valuable in evaluating its value.
Some models that would be useful to see scores for inside Zenith:
Opus 4.8 - Unclear why you would benchmark GPT 5.5 but not Opus 4.8
GLM-5.2 - Best open source model.
Qwen 3.6 27b - Arguably the top consumer-hardware sized coding agent. Also the most popular target to fine-tune coding agents on.
Qwen 3.6 35b - Very fast local model, popular for when you need things quickly and "free" (electricity).
Qwen 3.5 397B - One of the most fine tuned models out there, the foundation many build on.
Minimax M3 - Second place open source coding model.
Kimi K2.7 - Popular open source large coding model.
DeepSeek v4 Pro - Novel tech behind it, useful to see how it competes.
At a minimum, I feel like GLM 5.2 and Qwen 3.6:27b should be benchmarked. Making those surpass or at least approach (in the case of Qwen 3.6:27b) naive opus/gpt would be impressive and much more meaningful than dragging an already top model across the finish line.
Showing that you can take one of the top models available and push it across another top model is neat, but that doesn't really tell us much. Showing how good a wide range of models do with your harness would be far more valuable in evaluating its value.
Some models that would be useful to see scores for inside Zenith:
Opus 4.8 - Unclear why you would benchmark GPT 5.5 but not Opus 4.8
GLM-5.2 - Best open source model.
Qwen 3.6 27b - Arguably the top consumer-hardware sized coding agent. Also the most popular target to fine-tune coding agents on.
Qwen 3.6 35b - Very fast local model, popular for when you need things quickly and "free" (electricity).
Qwen 3.5 397B - One of the most fine tuned models out there, the foundation many build on.
Minimax M3 - Second place open source coding model.
Kimi K2.7 - Popular open source large coding model.
DeepSeek v4 Pro - Novel tech behind it, useful to see how it competes.
At a minimum, I feel like GLM 5.2 and Qwen 3.6:27b should be benchmarked. Making those surpass or at least approach (in the case of Qwen 3.6:27b) naive opus/gpt would be impressive and much more meaningful than dragging an already top model across the finish line.