OCLGJun 16

Sequential Hiring of Contingent Workers Through Learning-Based Optimization

arXiv:2606.184384.2
Predicted impact top 58% in OC · last 90 daysOriginality Incremental advance
AI Analysis

For firms managing contingent labor under uncertainty, this work provides a learning-based policy with theoretical guarantees for a problem with operational frictions.

The paper tackles sequential workforce management with costly worker replacement and random hiring delays, proposing DR-UCB policy that achieves regret matching the lower bound and outperforms benchmarks in numerical experiments.

In this paper, we study a sequential workforce management problem in a contingent labor setting with uncertainty in both worker production and labor supply. A firm seeks to maximize cumulative profit by maintaining an active team of fixed size while learning worker productivity over time. We emphasize two critical operational frictions in this problem: replacing workers is costly, and workers may not be available immediately for hiring because of, for example, prior job commitments, scheduling constraints, or onboarding procedures. Thus, hiring decisions take effect only after a random delay. We formulate this problem as a stochastic multi-play bandit with costly switching and delayed actions, and develop a learning-based hiring policy, DR-UCB (DelayedReplacement-UCB), that makes replacement and hiring decisions sequentially through learning cycles. In each cycle, the policy uses real-time production data to determine when to initiate workforce changes and which workers to replace and hire. We show that the leading-order regret of the proposed policy matches its lower bound in its dependence on the time horizon. Our numerical experiments show that DR-UCB outperforms benchmark policies.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes