Windows Agent Arena (WAA) 🪟 is a scalable OS platform for testing and benchmarking of multi-modal AI agents.
-
Updated
Apr 13, 2026 - Python
Windows Agent Arena (WAA) 🪟 is a scalable OS platform for testing and benchmarking of multi-modal AI agents.
An agent benchmark with tasks in a simulated software company.
Frontier Models playing the board game Diplomacy.
Ranking LLMs on agentic tasks
The definitive benchmark for AI agents on OpenClaw. 45 tasks across 4 tiers. Powered by MyClaw.ai
AI clothes swap prompt cookbook with 50 virtual try-on prompts, KIE GPT Image before/after examples, failure fixes, and a benchmark rubric for AIClothSwap.
开源劳动者AI研究与评测项目:建设劳动知识、公开评测基准与共同治理机制|Open worker-centered AI research and evaluation project.
llmBench is a high-depth benchmarking tool designed to measure the raw performance of local LLM runtimes (Ollama, llama.cpp) while providing deep hardware intelligence.
🤖 A curated list of resources for testing AI agents - frameworks, methodologies, benchmarks, tools, and best practices for ensuring reliable, safe, and effective autonomous AI systems
A visual benchmark testing whether leading coding agents repeat the same design patterns across 100 neutral website briefs.
Benchmarking and evaluation resources for trust, robustness, and reliability analysis in graph-based AI systems.
📊 Daily auto-updated snapshots of all Arena AI (LMSYS Chatbot Arena) leaderboards — LLM, Vision, Code, Video, Image & more. Structured JSON with historical tracking.
Daily evidence-first radar for AI benchmarks, evaluations, datasets, and data quality.
MindTrial: Evaluate and compare AI language models (LLMs) on text-based tasks with optional file/image attachments and tool use. Supports multiple providers (OpenAI, Google, Anthropic, DeepSeek, Mistral AI, xAI, Alibaba, Moonshot AI, OpenRouter), custom tasks in YAML, and HTML/CSV/JSON reports.
Precision-aware retrieval benchmark for LLM memory systems.
GDB: GraphicDesignBench - A real-world benchmark for evaluating AI on graphic design tasks
Benchmark for evaluating AI epistemic reliability - testing how well LLMs handle uncertainty, avoid hallucinations, and acknowledge what they don't know.
A curated list of evaluation tools, benchmark datasets, leaderboards, frameworks, and resources for assessing model performance.
Benchmark abierto en español de 170 modelos de IA (118 con 20+ runs, 69 rankeados, juez Phi-4 independiente). Calidad, costo, velocidad, long-context y fuga de credenciales como dimensiones separadas. Alternativas a Claude, GPT y Gemini para agentes n8n/Hermes. Calculadora interactiva con tus propios pesos.
GTA (Guess The Algorithm) Benchmark - A tool for testing AI reasoning capabilities
Add a description, image, and links to the ai-benchmark topic page so that developers can more easily learn about it.
To associate your repository with the ai-benchmark topic, visit your repo's landing page and select "manage topics."