What it is

The hub for benchmark work in this wiki. It used to be one bucket; on 2026-09-10 the sources were split by what they contribute:

Cross-cutting design papers stay here: HELM for taxonomy-first multi-metric design, tau-bench for pass^k, AppWorld for state-diff verification with allowed change sets, the knowledge-work reporting schema, and the reliability metrics. Multi-agent benchmarks sit under Enterprise Environments. Domains still unmapped as concept pages: ecommerce, cloning, and computer-use benchmarks.

Key sources

  • Tau-Bench — origin of pass^k and the small-tasks-many-trials argument; end-state grading and its blind spot
  • AppWorld — 9-app simulated world, 457 APIs; state-diff verification with expected/allowed change sets catches collateral damage; best 48.8% TGC
  • Designing Benchmarks for Knowledge Work — Four-field reporting schema plus 18 O*NET-derived work activities; shows benchmarks evaluate less than the work product they claim
  • Science of AI Agent Reliability — Twelve accuracy-independent reliability metrics over four dimensions; 15 models show reliability plateaued while capability rose for 24 months
  • Holistic Evaluation of Language Models — the structural model for a benchmark: taxonomy, multi-metric by design, standardised adaptation, explicit gaps
  • Tau-Tau-Bench — scores a builder end-to-end on one artifact under a serving budget with the eval suite withheld; best 23.9% vs 82.2% expert ceiling, and developers weaken their own failing tests
  • The Meta-Agent Challenge — meta-agents build agents against a hidden test set under time and quota budgets; 5/39 beat human scaffolds, σ>0.1 on a third of configs, and zero-resource pressure induces hacking that direct prompting cannot
  • AIxCC SoK — agent teams building under 50K LLM budgets; non-linear accuracy multiplier (90% free, 40% → −13%) reorders the ranking, and automatic patch validation admits 38-46% semantically wrong work
  • HiL-Bench — Necessity/sufficiency task admission checked from rollouts; ask_human() is a frozen-LLM expert oracle without the price; Ask-F1’s harmonic mean builds the anti-hack into the metric
  • EnterpriseBench Corecraft — Expert-built enterprise world whose rubrics are LLM-judged and never validated; supplies the transfer evidence and the three OOD benchmarks, and is the position WorldSmith argues against

To ingest

(none)