Skip to content

Please include benchmark results for other models. #7

Description

@MicahZoltu

Showing that you can take one of the top models available and push it across another top model is neat, but that doesn't really tell us much. Showing how good a wide range of models do with your harness would be far more valuable in evaluating its value.

Some models that would be useful to see scores for inside Zenith:
Opus 4.8 - Unclear why you would benchmark GPT 5.5 but not Opus 4.8
GLM-5.2 - Best open source model.
Qwen 3.6 27b - Arguably the top consumer-hardware sized coding agent. Also the most popular target to fine-tune coding agents on.
Qwen 3.6 35b - Very fast local model, popular for when you need things quickly and "free" (electricity).
Qwen 3.5 397B - One of the most fine tuned models out there, the foundation many build on.
Minimax M3 - Second place open source coding model.
Kimi K2.7 - Popular open source large coding model.
DeepSeek v4 Pro - Novel tech behind it, useful to see how it competes.

At a minimum, I feel like GLM 5.2 and Qwen 3.6:27b should be benchmarked. Making those surpass or at least approach (in the case of Qwen 3.6:27b) naive opus/gpt would be impressive and much more meaningful than dragging an already top model across the finish line.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions