AIAug 3

Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce

arXiv:2608.0244125.0
Predicted impact top 4% in AI · last 90 daysOriginality Incremental advance
AI Analysis

This work provides a new auditable and verifiable environment for evaluating AI agents in complex, multi-agent commerce scenarios, which is important for researchers and developers building autonomous commerce systems.

This paper introduces Agentic Commerce World (ACWorld), an environment designed for evaluating AI agents in a 'vibe commerce' setting where agents handle buying and selling tasks based on natural language goals. The ACWorld Benchmark, comprising a 200-task capability-coverage track and a 60-task large-catalog track (searching 785,022 listings), was used to evaluate ten models, achieving mean scores ranging from 65.9% to 85.6% and 56.1% to 91.4% respectively.

In vibe coding, people describe software in natural language and delegate implementation to AI agents. By analogy, vibe commerce allows people to express buying or selling goals in natural language and delegate the corresponding tasks to agents. Commerce, however, requires independently controlled Buyer and Merchant agents to interact in a shared market while preserving their private objectives and distinct authority. We introduce Agentic Commerce World (ACWorld), an environment for evaluating such agents across ongoing transactions. Through its Vibe Commerce Protocol (VCP), ACWorld validates agent actions before updating shared transaction state and records the resulting interactions, making agent behavior auditable and evaluation reproducible. The ACWorld Benchmark contains a 200-task capability-coverage track and a 60-task large-catalog track that searches 785,022 transactable listings. Across ten models, mean scores range from 65.9% to 85.6% and from 56.1% to 91.4%, respectively. Our analysis shows that process-level evidence is necessary: final state alone can miss evaluated errors, incomplete trajectories still retain useful process signals, and large-catalog tasks expose bottlenecks across stages.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes