CLSEJun 22

EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions

arXiv:2606.2365425.6Has Code
Predicted impact top 18% in CL · last 90 daysOriginality Incremental advance
AI Analysis

Provides a reproducible benchmark for evaluating enterprise agents in realistic workplace scenarios, addressing the lack of standardized evaluation in this domain.

EnterpriseClawBench is a benchmark for enterprise agents built from 852 real-world workplace sessions, where the best configuration (Codex with GPT-5.5) achieves only 0.663, highlighting the need for multi-faceted evaluation.

Enterprise agents increasingly operate inside workspaces: they read heterogeneous files, invoke tools, and deliver business artifacts. We introduce EnterpriseClawBench, an enterprise agent benchmark constructed from proprietary, real-world agent sessions. Starting from a large archive of workplace sessions, the EnterpriseClawBench produces 852 reproducible tasks, each paired with recovered fixtures, rewritten prompts, role classes, skill subclasses, hard rules, and semantic rubrics. Because the sessions contain internal enterprise content, we do not release the benchmark data; instead, our reusable contribution is the construction and evaluation protocol. On EnterpriseClawBench, the best configuration reaches only 0.663 (Codex with GPT-5.5). These results show that enterprise agent evaluation must report harness--model combinations, artifact delivery, visual quality, cost, runtime, and skill-transfer behavior, rather than collapsing performance into a single score. Code: https://github.com/FrontisAI/EnterpriseClawBench

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes