What it is

Web agents that operate a browser, and the benchmarks used to measure them.

Key sources

  • WebArena — the canonical realistic web-agent benchmark; outcome-based evaluation and the 14% vs 78% human gap
  • WorkArena — enterprise (ServiceNow) tasks and BrowserGym, the shared environment for web-agent evaluation
  • Agent-E — agent-side paper; hierarchy, DOM denoising, change observation, and metrics beyond success rate
  • WorkArena++ — compositional planning/reasoning tasks (L2/L3) where agents score near zero; skill taxonomy and trace generation
  • TheAgentCompany — Self-hosted software startup with Sotopia-backed simulated colleagues; weighted checkpoint scoring; best agent completes 30.3% of 175 tasks
  • CRMArena-Pro — Salesforce B2B/B2C Orgs, 19 tasks, four business skills; 58% single-turn falls to 35% multi-turn against an incremental-release user simulator

To ingest

(none)