Skip to content

Added Knows Benchmark#336

Open
farhanishmam wants to merge 1 commit into
ServiceNow:mainfrom
farhanishmam:add-knows-benchmark
Open

Added Knows Benchmark#336
farhanishmam wants to merge 1 commit into
ServiceNow:mainfrom
farhanishmam:add-knows-benchmark

Conversation

@farhanishmam

Copy link
Copy Markdown
Contributor

Added the KNOWS benchmark corresponding to the BrowserGym Knows PR #397

…s, Slides)

Knows evaluates browser agents on long-horizon document authoring in Google
Workspace. An agent drives a real browser to build a Doc, Sheet, or Slide deck
from a natural-language goal, and the result is graded checkpoint-by-checkpoint
via the Workspace APIs, yielding a fractional reward rather than binary success.

110 tasks across 22 families and 3 Workspace apps (25 docs / 45 sheets /
40 slides), exposed as a single `knows` split.

This is the AgentLab-side counterpart to the BrowserGym change that adds the
`knows` action subset, benchmark config, and task metadata. As with TimeWarp,
the benchmark itself lives in an external package (`browsergym-knows`), so
AgentLab needs only the lazy import and a docs entry:

- lazy `import browsergym.knows` in `_get_env_name`, so worker processes
  register the gym envs before `gym.make` (joblib/ray workers get a fresh
  `sys.modules`, so the parent's `prepare_backends()` import does not carry
  over)
- a `Knows` row in the supported-benchmarks table

No dependency change, matching the TimeWarp precedent: AgentLab declares no
per-benchmark extras, and benchmark packages are installed out of band per the
setup link.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant