From a790ed94868a30ce58405bb93929e370c579c4b7 Mon Sep 17 00:00:00 2001 From: reacher-z Date: Mon, 27 Jul 2026 19:46:11 +0800 Subject: [PATCH] docs: add ClawBench to related benchmarks --- README.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/README.md b/README.md index acf2302..2c7fb12 100644 --- a/README.md +++ b/README.md @@ -33,6 +33,10 @@ **Headline** โ€” Best model-runtime pairing (Claude Opus 4.7 + Claude Code) = **41.2% PassRate**, vs >78% the same backbones reach on OSWorld-Verified. +## Related Benchmarks + +- **[ClawBench](https://claw-bench.com/)** โ€” A complementary live-web benchmark with 153 real-world tasks across 144 websites and safe request interception for reproducible evaluation. [Paper](https://arxiv.org/abs/2604.08523) ยท [Code](https://github.com/TIGER-AI-Lab/ClawBench) + ## ๐ŸŽฌ Demo https://github.com/weavebench/WeaveBench/raw/main/docs/media/rabbitmq_dlq_demo.mp4