Position: Behavioral Systems Require Behavioral Tests
For AI researchers and evaluators, this paper advocates a shift in evaluation focus from performance outcomes to behavioral processes, but it is a position paper without empirical results, so its immediate impact is limited.
This position paper argues that AI agents should be evaluated as behavioral systems using methods from behavioral sciences, proposing a research agenda for developing rigorous behavioral tests. It outlines directions such as recovering decision strategies, constructing isolating environments, and probing multi-agent dynamics.
Artificial agentic systems increasingly operate as behavioral systems by interacting with dynamic environments, pursuing goals, and adapting over time. Yet, current evaluation methods largely focus on performance outcomes, not the underlying behavioral processes that produce them. This paper argues that AI agents must be evaluated like other behavioral systems: through systematic observation, perturbation, and interpretation of their actions. We draw on lessons from the behavioral sciences to motivate this position, and propose a research agenda focused on developing rigorous behavioral tests. These include methods for recovering decision strategies from action sequences, constructing environments that isolate behavioral differences, and probing emergent dynamics in multi-agent systems. Taken together, these directions offer a roadmap for developing a science of AI behavior.