Contravariance Theory: Strong Alignment for Minimal Solutions to Hard Tasks
For the NeuroAI community, this provides theoretical justification for why different neural networks converge to similar representations when solving hard tasks, addressing core questions about comparing DNNs to brains.
The paper proves that for minimal deep neural network solutions to hard tasks, weak alignment of representations guarantees strong alignment of privileged axes, which propagates up the network hierarchy. This suggests that with sufficiently strong tasks, inter-network comparison metrics are not sensitive and convergent evolution is inevitable.
A series of results from the NeuroAI over the past fifteen years have raised core questions both about how to compare Deep Neural Network (DNN) models to the brain, and about how much convergent evolution to expect between artificial networks and real brain networks. Here, we show that for any two minimal DNN solutions to a sufficiently hard task: (i) "weak" alignment of network representations based on affine mappings guarantees "strong" alignment of privileged axes, and (ii) alignment "zippers" up the network hierarchy, causing the emergence of privileged axes from end-to-end task optimization. These results formalize the notion of contravariance from Cao and Yamins [2024], and illustrate important consequences for the theory of NeuroAI: with sufficiently strong tasks, choice of metric for inter-network comparison is not all that sensitive, and that convergent evolution is probably inevitable.