CVLGJun 24

REViT: Roto-reflection Equivariant Convolutional Vision Transformer

arXiv:2606.253184.4
Predicted impact top 83% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For tasks requiring orientation awareness, this work provides a simpler and more effective equivariant vision transformer, though it is an incremental improvement over prior equivariant methods.

The paper proposes a discrete roto-reflection group equivariant vision transformer with convolutional attention, achieving improved image classification performance over existing equivariant networks.

In this paper, we propose a discrete roto-reflection group equivariant vision transformer with convolutional attention. Roto-reflection equivariant networks preserve the rotational, flip and positional symmetry in feature maps, making them useful for tasks where orientation of the inputs is relevant to the model outputs. In image classification and object detection, most of the studies on roto-reflection equivariant models have focused on using convolutional neural networks rather than vision transformers. In this paper, we examine the challenges involved in achieving equivariance in vision transformers, and we propose a simpler way to implement a discretized roto-reflection group equivariant vision transformer. The experimental results demonstrate that our approach outperforms the existing approaches for developing discrete roto-reflection group equivariant neural networks for image classification.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes