Self-supervised learning of Split Invariant Equivariant representations

Quentin Garrido, Laurent Najman, Yann Lecun

Introduction

Self-supervised learning of image representations has made significant progress in recent years (Chen et al., 2020a; He et al., 2020; Chen et al., 2020b; Grill et al., 2020; Lee et al., 2021b; Caron et al., 2020; Zbontar et al., 2021; Bardes et al., 2021; Tomasev et al., 2022; Caron et al., 2021; Chen et al., 2021; Li et al., 2022a; Zhou et al., 2022a, b; HaoChen et al., 2021; He et al., 2022; Bardes et al., 2022), catching up to supervised baselines in tasks requiring high-level information such as classification. Most of these works are placed in a joint-embedding framework, where two augmented views are generated from a source image. These two views are then fed to an encoder, giving representations, and then through a projection head, giving embeddings. Finally, a loss minimises the distance between the embeddings, i.e. makes them invariant to the augmentations, and is combined with a regularisation loss to spread embeddings in space.

While these invariance based approaches have been very successful for classification when using augmentations that preserve the semantic information of the image, this removal of information may be problematic for downstream tasks. For example, the use of color-jitter removes color information which can be useful for tasks such as flower classification (Lee et al., 2021a). This motivates the goal of introducing equivariance to representations, in order to learn more general representations for more varied downstream tasks. We say that representations are equivariant if the application of data augmentation commutes with the application of the encoder, i.e. can the representations of two related views be mapped similarly as the views themselves. Previous works have introduced ways to enrich usually invariant representations by keeping information about the augmentations. One approach is to use subsets of augmentations to construct partially invariant representations (Xiao et al., 2021). This can also be done by predicting rotations (Dangovski et al., 2021), preserving augmentation strengths in the representations (Xie et al., 2022), or by predicting all of the augmentation parameters (Lee et al., 2021a) or a discretized version of them (Scherr et al., 2022). While a mapping between representations may exist with these approaches, they offer no guarantees on its existence nor on its complexity. There are also no reliable approaches to empirically prove its existence. As such we do not consider these methods to truly be equivariant. Learning equivariant representations requires being able to predict a representation from another in latent space, which has also been a successful paradigm. This can be done by a simple prediction head, either using reconstruction (Winter et al., 2022) or without (Park et al., 2022; Devillers & Lefort, 2022; Shakerinava et al., 2022).

Parameter prediction based methods have been developed for image datasets such as ImageNet (Deng et al., 2009) where there is no clear equivariant task and where augmentations happen in pixel space with no loss of information. On the other hand, equivariance based methods have been used on simpler synthetic datasets(Shakerinava et al., 2022) where we can evaluate equivariance, but where it is hard to evaluate other classical computer vision tasks such as classification. To bridge the gap between those two worlds, we start by introducing a dataset called 3DIEBench, consisting of renderings of over fifty-thousand 3D objects where we can study an equivariance related task (3D rotation prediction) and an invariant one (image classification). This allows us to measure more precisely how invariant classical self-supervised methods are, while also showing limitations of existing equivariant approaches where predictors often collapse to the identity, leading to invariant representations.

We then introduce a hypernetwork (Ha et al., 2016) based predictor which avoids a collapse to the identity by design and show how it can outperform existing predictor architectures. We further show that by separating the representations in equivariant and invariant parts, we can significantly improve performance on equivariance related tasks, allowing us to match supervised baselines. To complement our quantitative results we also analyze qualitatively the learned split invariant-equivariant representations and see that all invariant information is not discarded from the equivariant part, and that the predictor offers a meaningful way to steer the latent space. To summarize:

We introduce 3DIEBench, a new dataset to evaluate representations on tasks that require invariant and equivariant information

We show the limitations of existing predictor architectures and introduce a hypernetwork based one that improves performance on all methods

We show that splitting the representations in invariant and equivariant parts further improves performance on equivariance related tasks

Related works

Two main families of methods can be distinguished: contrastive and non-contrastive. Contrastive methods (Chen et al., 2020a; He et al., 2020; Chen et al., 2020b, 2021; Yeh et al., 2021) mostly rely on the InfoNCE criterion (Oord et al., 2018) except for (HaoChen et al., 2021) which uses squared similarities between the embedding. A clustering variant of contrastive learning has also emerged (Caron et al., 2018, 2020, 2021) and can be thought of as contrastive methods, but between cluster centroids instead of samples. Non-contrastive methods (Grill et al., 2020; Chen & He, 2020; Bardes et al., 2021; Zbontar et al., 2021; Ermolov et al., 2021; Li et al., 2022b; Bardes et al., 2022) aim at bringing together embeddings of positive samples, similar to contrastive learning. However, a key difference with contrastive methods lies in how those methods prevent a representational collapse. In the former, the criterion explicitly pushes away negative samples, i.e., all samples that are not positive, from each other. In the latter, the criterion considers the embeddings as a whole and encourages information content maximization to avoid collapse, e.g., by regularizing the empirical covariance matrix of the embeddings. While we study methods from both families in our experiments, they have been shown to lead to very similar representations (Garrido et al., 2022).

Introducing equivariance in invariant self-supervised learning

While most of the aforementioned works focus on learning representations that are invariant to augmentations, some works have instead tried to learn representations where information about certain transformations is preserved. This can be done by predicting the augmentation parameters (Scherr et al., 2022; Lee et al., 2021a; Gidaris et al., 2018), or by introducing other transformations such as image rotations (Dangovski et al., 2021). Preserving the augmentations’ strength in the representations can also be used to learn less invariant representations (Xie et al., 2022). These methods offer no guarantees on the existence of a mapping between transformed representations in latent space, nor ways to prove its existence or lack thereof. As such these methods cannot be considered to truly be equivariant.

Equivariant representation learning

Previous works have explored equivariant representation learning using autoencoders, such as transforming autoencoders (Hinton et al., 2011), Homeomorphic VAEs (Falorsi et al., 2018) or (Winter et al., 2022). Recent works such as EquiMod (Devillers & Lefort, 2022) or SEN (Park et al., 2022) have also included a predictor that enables the steering of representations in latent space, without requiring reconstruction. These methods form the basis for our comparisons. In (Marchetti et al., 2022), representations are split in class and pose, i.e. invariant and equivariant, and assumes a simple equivariant latent space where the group action is the same as in the underlying data, e.g. 3 dimensions to represent pose. This assumes prior knowledge on the group of transformations, and can prove limited when the transformations cause a loss of information. Transformations are also assumed to be small, similarly as for SEN. We aim at deriving a more general predictor architecture with no such priors. In (Shakerinava et al., 2022), equivariant representations are learned with no knowledge of the group element associated with the transformation, but by having pairs of samples where the same transformation was applied.

3DIEBench: A new benchmark for invariant-equivariant SSL

Existing datasets used to evaluate equivariant or invariant representations have flaws when trying to design a method to learn more general representations. Datasets used to evaluate equivariance in representations often consist of simple images due to the need to control how transformations are applied (Park et al., 2022; Kipf et al., 2019). Conversely, datasets used to evaluate invariant representations (Deng et al., 2009; Krizhevsky et al., 2009) are limited in the sense that position and shape of the objects in these dataset can not be parameterized by controllable transformations, and only pixel-level transformations can be applied on the images. This motivates us to introduce a new dataset called 3D Invariant Equivariant Benchmark (3DIEBench) that aims at bridging the gap between the two. We want a dataset that is not trivial for an invariant task (image classification) but where we still have control on the parameters of the scene and the objects within it to learn meaningful equivariant representations. Taking inspiration from 3DIdent (Zimmermann et al., 2021; von Kügelgen et al., 2021) we use renderings of 3D objects from the subset of ShapeNetCore (Chang et al., 2015) originating from 3d Warehouse (Trimble Inc, ). This gives us a total 52472 objects spread across 55 classes. We are then able to adjust various factors of variations such as the object rotation, the lighting color, or the floor color. We focus on learning representations that are equivariant with respect to object rotations of arbitrary strength due to their inherent difficulty when using a diverse dataset, as well as the loss of information they can cause when looking at 2d renderings of the scene. In our experiments, we constrain the range of rotations to Euler angles between −π2-\frac{\pi}{2} and π2\frac{\pi}{2}. The goal is to make the task tractable, while still remaining challenging, as we show in our experiments. Using arbitrary rotations on arbitrary objects can make the task close to impossible, even for primates (Logothetis et al., 1994). For each object, we generate 50 random values for the factors of variation and then render the scene using Blender (Blender Online Community, ) and BlenderProc (Denninger et al., 2019), for a total of around 2.5 million images. Sample renderings can be found in figure 1. In supplementary section G we provide more details on the dataset generation as well as additional visualizations. The dataset as well as the code to generate the renderings will be released.

Creating a general predictor

Invariant self-supervised learning

Equivariant representations

Given a group GG with representations ρX\rho_{X} and ρY\rho_{Y}, we say that a function f:X→Yf:X\rightarrow Y is equivariant with respect to GG if ∀x∈X,∀g∈G\forall x\in X,\forall g\in G we have

This means that a function ff is equivariant if it commutes with group transformations. It is worth noting that we can see invariance as a special case where ρY(g)=Id\rho_{Y}(g)=Id. As we show in our experiments, this is a common failure mode of existing equivariant approaches.

While prior works have focused on forcing equivariance by the architecture of ff (Cohen & Welling, 2016; Cohen et al., 2018), we focus on a setting where this is not possible, and where we do not even know ρX\rho_{X}. Indeed, in our constructed dataset, it is impossible to apply the transformation on the renderings, even though this was possible in the original 3D space. However, we still know the group elements that parametrized our transformation and are able to make use of them. The goal is then to learn both ff and ρY\rho_{Y} in order to obtain representations that are as equivariant as possible to the original transformation.

2 Our Method: SIE

While we are placed in the joint-embedding framework described previously, we introduce a split in two of the representations before the projection head hϕh_{\phi}. We separate yy (resp. y′y^{\prime}) as yinvy_{\text{inv}} which contains information that is preserved by the transformation, i.e. invariant information, and yequiy_{\text{equi}} which contain information that was changed by the transformation, i.e., equivariant information. To illustrate, if yy is 512-dimensional, we use the first 256 dimensions for yinvy_{\text{inv}} and the 256 last for yequiy_{\text{equi}}. We thus call our approach Split Invariant Equivariant (SIE). Both parts are then fed through separate projection heads hϕ,invh_{\phi,\text{inv}} and hϕ,equih_{\phi,\text{equi}} to ensure that no information is exchanged between the two after the split. This gives us embeddings zinvz_{\text{inv}} and zequiz_{\text{equi}}. While we want the invariant embeddings zinvz_{\text{inv}} and zinv′z^{\prime}_{\text{inv}} to be identical, the equivariant embeddings zequiz_{\text{equi}} and zequi′z^{\prime}_{\text{equi}} should only be identical after the predictor pψ,gp_{\psi,g}, which is parametrized by the transformation between the two views gg. As such, pψ,gp_{\psi,g} is our learnable ρY(g)\rho_{Y}(g) described previously. The representation split can also be interpreted as a single predictor on the whole representations where we force it to be the identity for certain dimensions of the representations (zinvz_{\text{inv}}) and allow more flexibility on the rest of the dimensions (zequiz_{\text{equi}}). This whole process is illustrated in figure 2.

We now discuss the loss function used to train SIE. We use VICReg (Bardes et al., 2021) as our basis since it lends itself well to split representation. In order to adapt its invariance criterion, we define our similarity criterion Lsim\mathcal{L}_{\text{sim}} as

which we use to match both our invariant and equivariant embedding pairs. In order to avoid a collapse of the representations, we use the original variance and covariance criterion to define our regularisation criterion Lreg\mathcal{L}_{\text{reg}} as

The goal of the variance criterion VV is to ensure that all dimensions are used in the embeddings and the goal of the covariance criterion CC is to decorrelate the dimensions to spread out the information in the embeddings. We are now ready to introduce our final criterion as

Notice that the regularisation criterion Lreg\mathcal{L}_{\text{reg}} is applied to the whole embeddings and not separately for the invariant and equivariant part. While this does not matter for the variance criterion, it allows the covariance criterion to decorrelate information between the invariant and equivariant parts which is consistent with our goal. We also add another variance criterion (in gray) on the output of the predictor to help stabilize training. Its goal is to avoid a fully collapsed predictor since it normalizes the predictions. While this helps to stabilize the beginning of the training it does not impact final performance. However, without it, some runs never learn useful representations and fall back to VICReg’s behaviour. It is thus an optional yet recommended component. We use λinv=λV=10\lambda_{\text{inv}}=\lambda_{V}=10,λequi=4.5\lambda_{\text{equi}}=4.5, and λC=1\lambda_{C}=1 in our experiments.

Predictor architecture

Previous works have employed predictors in self-supervised learning, for different purposes. In (Grill et al., 2020; Chen & He, 2020) the predictor used is purely deterministic and does not depend on the target. As such its only solution is to converge to the identity, giving it a limited role in introducing any level of equivariance. Recent works have introduced predictors that depend on the transformation between the two views of the input and have converged to a linear transformation in general (Devillers & Lefort, 2022), or only for 3D rotations (Park et al., 2022). As we study in supplementary section D, this kind of architectures can ignore the transformation parameters and collapse back to behaviours associated with invariance based methods, i.e. pψ,g=Idp_{\psi,g}=Id.

Experiments

We compare our approach to VICReg (Bardes et al., 2021) and SimCLR (Chen et al., 2020a) to have baselines for invariant self-supervised methods. We consider both the scenarios where they have to be invariant about gg as well as a scenario where g=0g=0 and where we apply standard image augmentations, following the protocol of (Grill et al., 2020). The goal is to see if we learn some information about the object pose by considering different poses as different samples instead of augmented views. We also compare our approach to SimCLR+AugSelf (Lee et al., 2021a) as a parameter prediction method. It is trained to predict gg, but since it does not provide a transformation in embedding space (ρY(g)\rho_{Y}(g)) it cannot be considered equivariant and is mostly included for completeness. Finally we compare ourselves to SEN (Park et al., 2022) and EquiMod (Devillers & Lefort, 2022) for equivariant methods. We consider them both with their original predictor as well as with our hypernetwork based predictor to demonstrate both its benefits as well as the benefits of the invariant-equivariant split. For SEN, we use the same contrastive loss as SimCLR instead of the original triplet loss to limit hyperparameter tuning. For clarity we label this change as Only Equivariance.

Training protocols

All methods use a ResNet-18 (He et al., 2016) as their encoder and a three layer MLP as projection head. To obtain asymptotic behaviours they are all trained for 2000 epochs using the Adam optimizer (Kingma & Ba, 2014), with learning rate 10−310^{-3} and default β\beta parameters. We give more details on the pretraining strategies in supplementary section A.

2 Representations evaluation

As is common practice in self-supervised learning we start by evaluating the quality of the representations on dowsntream tasks. We use linear classification on top of frozen representations as our representative invariant task. It is worth noting that this is not purely invariant since some information about the transformation is helpful in practice (Bordes et al., 2022). For our representative equivariant task, we use rotation prediction. The representations from two transformed views of the same object are fed through a 3 layer MLP that is trained to regress the rotation between the two. We also include linear regression of the floor and spot hue to study any side effects on a task that is comparatively simple. We give more details on the evaluation protocols in supplementary section A.

Quantitative results

Our results are summarized in table 1. We first notice that invariant self-supervised learning methods offer classification accuracies that are very close to the supervised baselines, but they offer worse performance in rotation prediction. When considering rotated objects as different instances and not as a transformation, we notice an increase of performance on rotation prediction at the cost of classification performance. Overall, VICReg offers a higher level of invariance than SimCLR across all metrics. AugSelf yields a significant boost over the SimCLR baseline for rotation prediction performance, albeit with a small drop in classification. Its pretraining task is identical to our evaluation which explains why its performance is close to the supervised baseline. Looking at equivariant methods, we see that both EquiMod and Only Equivariance perform very well on classification but offer no increase in performance for rotation prediction compared to SimCLR which serves as their base. This would suggest that their original predictor does not induce more equivariance that the implicit equivariance offered by the projection head. We discuss this behaviour below.

When using our hypernetwork based predictor with EquiMod or Only Equivariance, we notice a clear boost in performance in rotation prediction, showing that it is able to improve the performance on equivariance related tasks. The performance is however still far from the supervised baseline. When combining both this predictor architecture and our split representations, SIE is able to further improve performance on rotation prediction, but by incurring a small performance drop compared to its VICReg base. Looking into more details about SIE’s results, we see that the equivariant part of the representations still contains a significant amount of information that is helpful for classification. The rotations that we apply can be so extreme that knowing the nature of the object is important to learn how to apply a rotation. When looking at the invariant part of the representations, we see that it achieves the lowest performance, showing that SIE learned to most invariant representations. As for color, we also obtain almost perfect invariance, highlighting again the invariance of our representations to transformations that have no incentive to be preserved. Overall, both our predictor architecture and invariant-equivariant split help to greatly improve performance on equivariance related tasks, while also obtaining the highest level of invariance when desired. We study the task of learning representations that are both equivariant to rotation and color in supplementary section E, where we are able to show increased performance in color prediction with only slight drops in performance for rotation prediction.

Implicit equivariance of the projection head

While a significant part of the performance on equivariance related tasks can be attributed to the predictor, even invariant methods have various level of performance. This can be explained by the use of the projection head which absorbs some of the bias from the invariance criterion (Bordes et al., 2022; Chen et al., 2020a). However, this is not enough to achieve satisfactory performance, nor does it give a way to steer the latent space. Nonetheless, even predictor based methods can benefit from it and it remains an easy way to improve quantitative performance. We provide a more in depth analysis of the role of the projection head in supplementary section C, where we see that SIE suffers from the smallest drops in performance after the projection head, matching the performance on representations of EquiMod, whereas EquiMod falls to a level of performance similar to VICReg before its projection head.

Predictor collapse to the identity

The predictor architecture that was originally used in (Park et al., 2022; Devillers & Lefort, 2022) was a linear layer which takes as input the concatenation of the representations and the transformation parameters. This means that it can simply choose to ignore the transformation parameters by setting the appropriate weights to 0. This behaviour is also accompanied by a collapse to the identity of the predictor, since solving the invariant task is easier than learning equivariant representations. This happens in practice for every method, and is the reason why the performance with this predictor architecture is similar to what an invariant method gives. Confer supplementary section D for a study of this phenomenon.

Qualitative results

In order to visualize the information present in the representations, we perform a retrieval of nearest representations on our validation set. We expect the representations of invariant methods to contain similar objects in various poses but representations of equivariant methods to contain objects in similar poses to the queried representations. As we can see in figure 3 all methods lead to nearest-neighbours in similar poses as the queried object except for the invariant part of the representations from SIE. Nonetheless, for SimCLR, Only Equivariance and EquiMod, the nearest neighbour is not in a similar pose as the queried object, highlighting an imperfectly learned equivariant mapping. Invariant methods preserve transformation related information in a way that is stronger than expected due to the implicit equivariance introduced by the projection head. We reproduce the same figure on embeddings in supplementary section C and notice that invariant methods do not lead to nearest neighbours in similar poses, whereas equivariant methods tend to perform better, especially SIE. We do notice that all objects are cars, which would suggest that the class information is still present, confirming our quantitative results.

3 Predictor evaluation

Quantitative results

As we can see in table 2, no matter the considered dataset, SIE consistently outperforms EquiMod and Only Equivariance, achieving 0.290.29 PRE on the validation set compared to 0.480.48 for EquiMod and Only Equivariance. The results are similar with MRR and H@1/5 where on both the training and validation set SIE outperforms EquiMod and Only Equivariance. To interpret better what this means, the H@1 of 0.30.3 for SIE means that 30%30\% of the time the nearest neighbour is the target embedding, whereas this is only true 5%5\% of the time for EquiMod and Only Equivariance. This is only slightly better than random which would be 2%2\%. Since the same predictor is used for all methods, this highlights the importance of using split representations.

Qualitative results

To give a clearer picture of the predictor’s influence on embedding, we show the closest neighbours of predicted embedding in figure 4. Starting from a pair of embeddings, we apply the predictor on the starting embedding with the goal of rotating it to the pose of the target embedding. When retrieving the nearest neighbours of the predicted embeddings, we expect them to be objects in a similar pose as the target. For both Only Equivariance and EquiMod, we see that the predicted embeddings are dissimilar to the target embedding, whereas for SIE, we do find other cars in the same pose.

While both quantitative and qualitative results would suggest that the predictor is of lower quality in EquiMod and Only Equivariance compared to SIE, these results must be interpreted with care. Since our visualization and metrics rely on nearest neighbour retrieval on the whole dataset, if equivariant information contributes less to the norm of the embeddings compared to invariant information then it may not be captured well. Nonetheless, this still shows that the invariant-equivariant split in SIE plays a significant role in making equivariant information easily accessible. As mentioned previously, the drop in quality when going from the representations to the embeddings can be partially attributed to the implicit equivariance induced by the projection head.

Limitations

While we have shown improved performance over existing methods both thanks to our hypernetwork-based predictor and split representations, SIE currently requires knowledge about the group elements. This can limit its applicability in settings where they are unknown or where only partial information is available. While existing works also suffer from this limitations, removing the need for this knowledge is an important future line of work. We have also shown that using split representations significantly helps equivariance-related performance, however this comes at a small cost in invariance performance. When optimizing for different tasks a trade-off has to be made for performance and different methods will be optimal for a given use-case. SIE gives a satisfactory trade-off, maximizing equivariance performance while preserving most of the invariance performance, but other choices may be better suited for different targeted downstream tasks.

Conclusion

We have introduced a method to learn both invariant and equivariant representations based on self-supervised learning. By introducing 3DIEBench, we create an experimental setting which is more challenging than existing datasets to learn equivariant representations and that also enables us to evaluate on image classification, a task that can be assimilated to invariant representations. By using a predictor based on a linear hypernetwork and by splitting representations in an invariant and equivariant part, SIE is able to beat existing equivariant methods on both qualitative and quantatitative metrics. Reproducing previous works with our hypernetwork-based predictor further enabled us to show the positive impact on performance of both the predictor design and the use of split representations. We hope that SIE can serve as a basis towards the design of equivariant methods in more complex settings, enabling the use of self-supervised learning to learn richer representations.

Acknowledgments

We wish to thank Adrien Bardes for insightful discussions as well as inspiring the first experiments that led this work. We also want to thank Randall Balestrierio, Grégoire Mialon, Pascal Vincent, Florian Bordes, Nicolas Ballas, Mido Assran and Yubei Chen for insightful discussions throughout the development of this work.

References

Appendix A Exact training and evaluation protocols

As previously mentioned, all methods are trained for 2000 epochs using a resnet-18 encoder and MLP projection head. We use a batch size of 1024 with the Adam optimizer with learning rate 10−310^{-3}, β1=0.9\beta_{1}=0.9 , β2=0.999\beta_{2}=0.999. All experiments are done using 4 NVIDIA V100 GPUs and take around 24 hours. As we discuss in supplementary section B, shorter training regimens can be used and lead to similar performance. We discuss method-specific hyperparameters.

In the supervised baselines we used the same protocol and used the same evaluation heads that are used when evaluating self-supervised approaches on top of a resnet-18, see next subsection for details. We train for 2000 epochs to also obtain asymptotic results and to provide better upper bounds on performance.

VICReg

We use a projection head with intermediate dimensions 2048-2048-2048, as well as loss weights λinv=λvar=10\lambda_{inv}=\lambda_{var}=10 and λcov=1\lambda_{cov}=1.

SimCLR, Only Equivariance

We use a projection head with intermediate dimensions 2048-2048-2048, and temperature τ=0.1\tau=0.1 for the loss.

SimCLR + AugSelf

We use a projection head with intermediate dimensions 2048-2048-2048, and temperature τ=0.1\tau=0.1. For the parameter prediction head, we use a MLP with intermediate dimensions 1024-1024-4. We weigh the two losses (SimCLR and parameter prediction) equally.

EquiMod

We follow the original protocol and use projection heads with intermediate dimensions 1024-1024-128. We use these dimensions to coincide with the ones used for SIE, giving us a fair comparison. We use τ=0.1\tau=0.1 as our loss temperature, and weigh the two losses (invariance and equivariance) equally. We tried with different weights and found the original equivariance weight of 1 to work best.

SIE

For both our invariance and equivariant projection heads we use intermediate dimensions 1024-1024-1024. We use λinv=λV=10\lambda_{\text{inv}}=\lambda_{V}=10, λequi=4.5\lambda_{\text{equi}}=4.5, and λC=1\lambda_{C}=1 in our experiments. See supplementary section B for ablations on these parameters.

A.2 Evaluation

Following common protocols we train a linear classification head on top of frozen representations for 300 epochs using a batch size of 256. We rely on the Adam optimizer with learning rate 10−310^{-3}, β1=0.9\beta_{1}=0.9 , β2=0.999\beta_{2}=0.999. We use a cross entropy loss to train our classifier. Performance is then reported on the validation set.

Rotation prediction

We train a MLP with intermediate dimensions 1024-1024-4 on top of our frozen encoder. Its inputs are pairs of representations that are concatenated. We train for 300 epochs using a batch size of 256. We rely on the Adam optimizer with learning rate 10−310^{-3}, β1=0.9\beta_{1}=0.9 , β2=0.999\beta_{2}=0.999. We use a MSE loss as our regression loss.Performance is then reported on the validation set using R2R^{2}, which contains unseen objects with the same pose distribution as seen during training.

Color prediction

We train a linear regression head on top of frozen representations.Its inputs are pairs of representations that are concatenated. We train it for 50 epochs using a batch size of 256. We rely on the Adam optimizer with learning rate 10−310^{-3}, β1=0.9\beta_{1}=0.9 , β2=0.999\beta_{2}=0.999. We use a MSE loss as our regression loss. Performance is then reported on the validation set using R2R^{2}.

Appendix B Ablations

We study more carefully the impact of each component of SIE in tables S1, S2 and S3.

We can see that having either no invariance criterion on a part of the representation, or using a linear or MLP (4 layers) predictor leads to performance in the ballpark of VICReg. However, the use of the hypernetwork-based predictor significantly boosts performance, without necessarily increasing the parameter count of the model. We also see that the split representations lead to optimal representations, compared to using two projection heads on the full representations like EquiMod, or using only equivariance. For fairness, we used the same grid of equivariance weights for the VICReg-EquiMod scenario as we did for the split representations.

When looking at the performance for different training schedules, we see that performance plateaus after 1000 epochs, suggesting that shorter training regimens will lead to similar results for a much more manageable compute cost.

Classification

We see that the performance is in general not dependent on the predictor architecture. Even though we notice a small drop for the MLP predictor, the hypernetwork achieves similar performance as the linear predictor, which performed significantly worse on the rotation prediction task. Interestingly, using VICReg-EquiMod does not lead to a significant drop in performance compared to classical VICReg, whereas using only equivariance or split representations led to bigger drops.

Contrary to what we found for rotation prediction, training for longer always improve classification performance, even when going from 1500 to 2000 epochs. We also notice that SIE is the method that benefits the most from longer training, gaining over 5 points in top-1 accuracy going from 500 to 2000 epochs, compared to less than 2.52.5 for EquiMod. This indicates that longer training are beneficial in general, even if they are not for rotation prediction.

Choice of base SSL criterion

As we have seen in table 1, SimCLR’s criterion leads to less invariant representations compared to VICReg’s. However when looking at tables S1 and S2 we see that when using VICReg’s criterion, EquiMod seems to perform better on equivariant tasks compared to the base method using SimCLR’s criterion. This shows an opposite behaviour for equivariant performance compared to the default versions of SimCLR and VICReg. We also evaluate the predictor quality for VICReg-EquiMod in table S3. We can see that while VICReg-EquiMod achieves better performance than the classical EquiMod, SIE still outperforms it by a significant margin. This further demonstrates the usefulness of splitting representations.

Appendix C Performance on embeddings

In order to better understand the role of the projection head in absorbing invariance, we evaluate methods on the embeddings instead of the predictor.

As we can see in table S4, all methods suffer from a drop in performance after the projection head, highlighting the importance of the projection head’s implicit equivariance. However the drop is significantly smaller for SIE, which achieves similar performance after the projection head to what other method achieve before it. This further help explain the quality of the predictor learned by SIE. We can notice that VICReg and SimCLR achieve a level of invariance after the projection head that is similar to SIE’s before the projection head on the invariant part. As such the split of information was helpful to achieve invariance even before the projection head.

From a qualitative point of view, we reproduce figure 4 which was done on the representations in figure S1 which is done on the embeddings. We notice that SIE still achieves a similar level of equivariant performance, where the retrieved objects are in the same pose as the query. But for other methods the invariance is much stronger, and interestingly Only Equivariance does not appear to have much information about the object’s pose. This further corroborates what we saw quantitatively.

Appendix D Predictor collapse to the identity

A phenomenon that motivated the introduction of the hypernetwork-based predictor is the collapse of certain predictors to the identity. This is something that can be observed in invariant self-supervised learning methods such as SimSiam (Chen & He, 2020) where the predictor does not depend on the chosen images, but this can also happen in our framework. To illustrate this, we look at the weight matrix of the predictor when using a linear predictor where transformation parameters are concatenated to the embeddings, such as used by SEN or EquiMod.

As we can see in figure S2, the first four columns which are associated with the rotation are zero, and so the rotation parameters are ignored. The rest of the predictor is very close to the identity, which would suggest that the predictor effectively does nothing. As such the method collapsed to a classical invariant method in this case.

This highlights the need for a predictor where the information about the transformation cannot be ignored, such as our hypernetwork based design. This is a problem that must be taken into consideration when designing predictor based methods in complex scenarios.

Appendix E Results on rotation and color equivariance

While we previously trained models to be equivariant to rotation only, we study here the performance when trying to learn representations that are equivariant to both rotation and changes in color of the floor and light.

As we can see in table S5, all methods achieve a similar level of performance for classification and rotation prediction, although we can notice a slight drop for SIE. Looking at color prediction, we see that the performance is significantly increased from models trained to be equivariant only to rotation, achieving almost perfect performance on color prediction. The hypernetwork-based predictor also brings increased performance here, and the split representations bring a slight advantage in performance, though less noticeable than on rotation alone.

Appendix F Generalization to unseen rotations

While we have trained methods on a certain set of rotations, we can wonder what happens when confronted with objects in unseen poses. To evaluate this, we generated rotations along the zz-axis spanning the full circle, in increments of 5 degrees.

As we can see in figure S3, while the model manages to extrapolate reasonably well for the table, it fails to do so for the car, giving retrieved images that are at the maximum rotation that could have been seen during training. There may be multiple reasons for this. This could come from the encoder that isn’t able to encode these rotated images properly, but this could also come from the predictor that is not able to learn a transformation for these unseen angles. Most likely, this phenomenon is a combination of the two.

Appendix G Dataset generation

The basis of 3DIEBench is the subset of ShapeNetCore coming from 3D Warehouse. This gives 52462 models spanning 55 classes. We split the dataset into a training and validation part, containing respectively 80% and 20% of the objects. Starting from a given 3D model, we generate 50 different scenes by changing factors of variation using the ranges described in table S6. We change the light position to ensure that the objects shadow is not informative and cannot be used as a crutch by the network. The range of rotations is limited in order to make the problem more tractable. As we see in the results, it is still very challenging for existing approaches.

The generation of all images takes around 500 hours on a single NVIDIA V100 GPU, but can be easily parallelized.

G.2 Additional data samples