Python framework for evaluating LLM tool-calling behavior with comprehensive metrics on accuracy, efficiency, and correctness
-
Updated
Jun 19, 2026 - Python
Python framework for evaluating LLM tool-calling behavior with comprehensive metrics on accuracy, efficiency, and correctness
Track AI-agent task metrics: token cost, retry pressure, and outcome quality for Claude Code and similar tools.
Open-source LLM observability, FinOps, and governance platform - normalize token usage, cost, latency, errors, retries, quotas, and billing across OpenAI, Anthropic, Gemini, Azure OpenAI, and AWS Bedrock. OpenTelemetry-native, Prometheus-ready, self-hosted. Apache-2.0.
Add a description, image, and links to the llm-metrics topic page so that developers can more easily learn about it.
To associate your repository with the llm-metrics topic, visit your repo's landing page and select "manage topics."