A lightweight, general-purpose harness for tool-using LLM agents, fair benchmark evaluation, harness baselines, and personal assistant workflows.
-
Updated
Jul 16, 2026 - Python
A lightweight, general-purpose harness for tool-using LLM agents, fair benchmark evaluation, harness baselines, and personal assistant workflows.
Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents
Add a description, image, and links to the researchclawbench topic page so that developers can more easily learn about it.
To associate your repository with the researchclawbench topic, visit your repo's landing page and select "manage topics."