Eval suite exposing a failure mode in the Gemini 3.5 Flash model (exception-blindness in tabular reconciliation)
-
Updated
May 28, 2026 - Python
Eval suite exposing a failure mode in the Gemini 3.5 Flash model (exception-blindness in tabular reconciliation)
Terminal-Bench 3 task: reverse-engineer a custom binary log format from a partial spec, samples, and a black-box validator
Terminal-Bench 3 task for evaluating agentic log analysis, mixed-format parsing, and deterministic JSON reporting.
Smallest readable coding-agent harness that scores on benchmarks. ~970 lines, 59.6% on Terminal-Bench 2.0.
Add a description, image, and links to the terminal-bench-3 topic page so that developers can more easily learn about it.
To associate your repository with the terminal-bench-3 topic, visit your repo's landing page and select "manage topics."