StackingNet: Collective Inference Across Independent AI Foundation Models
Provides a practical method for coordinating independently developed black-box foundation models, addressing the problem of isolated AI systems for practitioners needing trustworthy ensemble predictions.
StackingNet, a meta-ensemble framework, aggregates outputs from independent black-box foundation models at inference, improving accuracy, reducing error and disparities, and ranking reliability without accessing internal parameters or training data. Across language, vision, and paper rating tasks, it consistently outperforms individual models and classic ensembles, with gains widening as model diversity increases.
Artificial intelligence built on large foundation models has transformed language understanding, computer vision, and reasoning, yet these systems remain isolated and cannot readily share their capabilities. Coordinating the complementary strengths of independently developed, black-box foundation models is essential for trustworthy intelligent systems, yet no established method exists. Here we show that such coordination can be achieved through a meta-ensemble framework termed StackingNet, which aggregates the output predictions of independent models at inference. StackingNet improves accuracy, reduces individual-model error and group-wise disparities, ranks model reliability, and identifies or prunes models that degrade performance, all without access to internal parameters or training data. Across language comprehension, visual attribute estimation, and academic paper rating, it consistently outperforms individual models and classic ensembles, with gains that persist when the base models are uniformly strong. These gains stem from variance reduction and consensus alignment among independent models rather than from any emergent group cognition, and they widen as the model pool grows more diverse. By turning model diversity from a source of inconsistency into a resource for cooperation, StackingNet offers a practical path toward coordinated artificial intelligence, where progress emerges not only from larger single models but from principled cooperation among many specialized ones.