What it is
The hub for benchmark work in this wiki. It used to be one bucket; on 2026-09-10 the sources were split by what they contribute:
- how the human side is simulated: User Simulators
- whether a benchmark measures what it claims: Benchmark Validity
- agents authoring tasks, environments, and verifiers: Environment Generation
- simulated companies and IT operations: Enterprise Environments
- browser and UI agents: Web Agents
- retrieval and QA evaluation: Retrieval Benchmarks
- agents doing the ML engineer’s job: Agents Automating ML Work
Cross-cutting design papers stay here: HELM for taxonomy-first multi-metric design, tau-bench for pass^k, AppWorld for state-diff verification with allowed change sets, the knowledge-work reporting schema, and the reliability metrics. Multi-agent benchmarks sit under Enterprise Environments. Domains still unmapped as concept pages: ecommerce, cloning, and computer-use benchmarks.
Key sources
- Tau-Bench — origin of pass^k and the small-tasks-many-trials argument; end-state grading and its blind spot
- AppWorld — 9-app simulated world, 457 APIs; state-diff verification with expected/allowed change sets catches collateral damage; best 48.8% TGC
- Designing Benchmarks for Knowledge Work — Four-field reporting schema plus 18 O*NET-derived work activities; shows benchmarks evaluate less than the work product they claim
- Science of AI Agent Reliability — Twelve accuracy-independent reliability metrics over four dimensions; 15 models show reliability plateaued while capability rose for 24 months
- Holistic Evaluation of Language Models — the structural model for a benchmark: taxonomy, multi-metric by design, standardised adaptation, explicit gaps
- Tau-Tau-Bench — scores a builder end-to-end on one artifact under a serving budget with the eval suite withheld; best 23.9% vs 82.2% expert ceiling, and developers weaken their own failing tests
- The Meta-Agent Challenge — meta-agents build agents against a hidden test set under time and quota budgets; 5/39 beat human scaffolds, σ>0.1 on a third of configs, and zero-resource pressure induces hacking that direct prompting cannot
- AIxCC SoK — agent teams building under 50K LLM budgets; non-linear accuracy multiplier (90% free, 40% → −13%) reorders the ranking, and automatic patch validation admits 38-46% semantically wrong work
- HiL-Bench — Necessity/sufficiency task admission checked from rollouts; ask_human() is a frozen-LLM expert oracle without the price; Ask-F1’s harmonic mean builds the anti-hack into the metric
- EnterpriseBench Corecraft — Expert-built enterprise world whose rubrics are LLM-judged and never validated; supplies the transfer evidence and the three OOD benchmarks, and is the position WorldSmith argues against
Related
- WorldSmith
- ITSMBench
- User Simulators
- Benchmark Validity
- Environment Generation
- Enterprise Environments
- Web Agents
- Retrieval Benchmarks
- Agents Automating ML Work
To ingest
(none)