Non-isotropy Regularization for Proxy-based Deep Metric Learning

Karsten Roth, Oriol Vinyals, Zeynep Akata

Introduction

Visual similarity plays a crucial role for applications in image & video retrieval and clustering , face re-identification or general supervised and unsupervised contrastive representation learning. A majority of approaches used in these fields employ or can be derived from Deep Metric Learning (DML). DML aims to learn highly nonlinear distance metrics parametrized by deep networks. These networks span a representation space in which semantic relations between images are expressed as distances between respective representations.

In the field of DML, methods utilizing proxies have shown to provide among the most consistent and highest performances in addition to fast convergence . While other methods introduce ranking tasks over samples for the network to solve, proxy-based methods require the network to contrast samples against a proxy representation, commonly approximating generic class prototypes. Their utilization addresses sampling complexity issues inherent to purely sample-based approaches, resulting in improved convergence and benchmark performance.

However, there is no free lunch. Relying on sample-proxy relations, relations between samples within a class can not be explicitly captured. This is exacerbated by proxy-based objectives optimizing for distances between samples and proxies using non-bijective distance functions. This means, for a particular proxy, that alignment to a sample is non-unique - as long as the angle between sample and proxy is retained, i.e. samples being aligned isotropically around a proxy (see Fig. 1), their distances and respective loss remain the same. This means that samples lie on a hypersphere centered around a proxy with same distance and thus incurring the same training loss. This incorporates an undesired prior over sample-proxy distributions which doesn’t allow local structures to be resolved well. By incorporating multiple classes and proxies (which is automatically done when applying proxy-based losses such as to training data with multiple classes), this is extended to a mixture of sample distributions around proxies. While this offers an implicit workaround to address isotropy around modes by incorporating relations of samples to proxies from different classes, relying only on other unrelated proxies potentially far away makes fine-grained resolution of local structures difficult. Furthermore, as training progresses and proxies move further apart. As a consequence, the distribution of samples around proxies, which proxy-based objectives optimize for, comprises modes with high affinity towards local isotropy. This introduces semantic ambiguity, as semantic relations between samples within a class are not resolved well. However, a lot of recent work has shown that understanding and incorporating these non-discriminative relations drives generalization performance .

Related Works

Deep Metric Learning (DML) has driven research in image and video retrieval & zero-shot clustering applications , with particular applications for example in person re-identification and as an auxiliary tool for improved supervised and unsupervised representation learning . Commonly proposed DML methods introduce ranking tasks for networks to solve as training surrogates. These can involve ranking constituents in tuples (such as pairs , triplets or higher-order tuples ) to contrasting between sample and prototypical representations . These prototypical- or proxy-based approaches are commonly introduced in order to address sampling complexity issues when sampling tuples for a network to solve, which are otherwise addressed through various sampling heuristics . We propose a natural extension to proxy-based approaches by addressing a major shortcoming introduced when only contrasting between samples and proxies while retaining beneficial properties of these methods. Finally, recent work has focused on generic extensions to DML to improve the quality of learned representation spaces through divide-and-conquer , synthetic data , adversarial and graph-based training , bypassing representation bottlenecks , attention and auxiliary or few-shot feature learning . These works offer distinct, orthogonal benefits.

Non-isotropic Deep Metric Learning

Proxy-based objectives use contrastive operations not between samples (e.g. via the cosine similarity sψ(xi,xj)s_{\psi}(x_{i},x_{j}) ), but between class-prototypical (class proxies) representations ρj∈P\rho_{j}\in\mathcal{P} s(ψi,ρj)s(\psi_{i},\rho_{j}) for classes yiy_{i} and yjy_{j}. This removes the need for complex sampling operations in methods reliant on sample-based tuples, which allows proxy-based objectives to benefit from fast convergence and good generalization performance.

However, this property also incurs the strongest shortcoming, because relying on sample-proxy pairs and the non-bijective similarity measure s(ψ,ρ):=s(ψi,ρyψi)s(\psi,\rho):=s(\psi_{i},\rho_{y_{\psi_{i}}}) can induce features to locally follow an isotropic distribution around the proxy. This can be seen more explicitly when looking at the sample-proxy distributions various proxy-objectives optimize for. Take for example the foundational ProxyNCA objective . ProxyNCA is heavily connected to various recent, state-of-the-art objectives (such as ProxyAnchor or SoftTriple ) and has the form

with the complete set of class proxies P\mathcal{P} and with class yy removed P−y\mathcal{P}^{-y}We use cosine similarity instead of the euclidean distance as done in , as both are equivalent on the hypersphere. that are trained jointly during training. Minimizing the distance of samples to their respective class proxies while maximizing it for non-related proxies, this objective can be regarded as implicitly maximizing the log-likelihood of samples ψ\psi belonging to proxy ρ\rho (such that yψ=yρy_{\psi}=y_{\rho}) under a von-Mises-Fisher (vMFAn assumption found also e.g. in self-supervised learning .) mixture model around directions ρ\rho

assuming a class-independent concentration parameter κρ=κ\kappa_{\rho}=\kappa and mixture πρ=π\pi_{\rho}=\pi such that Cd(κρ)=Cd(κ)=constC_{d}(\kappa_{\rho})=C_{d}(\kappa)=constCdC_{d} incorporates the modified Bessel function IpI_{p} of the first kind and order pp, which can be neglected here as CdC_{d} cancels out. Even more, show that performance of LPNCA\mathcal{L}_{\text{PNCA}} improves when actually optimize the proxy-assignment probability (by replacing P−y\mathcal{P}^{-y} with P\mathcal{P} in the denominator, giving LPNCA++\mathcal{L}_{\text{PNCA++}}) directly, which equals an explicit negative log-likelihood minimization of pvMFmmp_{\text{vMFmm}}:

Recent and state-of-the-art proxy objectives that extend upon ProxyNCA, such as the ProxyAnchor objective

operate under similar assumptions, suggesting slight, more hyperparameter-heavy variations to the loss terms. While ProxyAnchor specifically suggests pulling samples towards proxies instead of proxies towards samples as done in LPNCA\mathcal{L}_{\text{PNCA}}, it similarly relies on the same sample-proxy contrastive operations to learn a respective metric space. This means that these methods can only learn sample distributions around proxies that are closely related to pvMFmmp_{\text{vMFmm}}-like distributions learned via LPNCA\mathcal{L}_{\text{PNCA}} . Indeed, our experiments (see §4.2 and Tab. 1) show that when adapting the ProxyNCA objective to the ProxyAnchor formulation

performance becomes much more similar than indicated in . And while ProxyAnchor may not optimize for the exact p(ψ∣ρ)p(\psi|\rho) formulation, the results support a very strong distributional relation between these proxy-objectives.

However, mixture distributions such as pvMFmmp_{\text{vMFmm}} suffer from several issues. Firstly, each mode on its own is isotropic as s(ψ,ρ)s(\psi,\rho) returns the same value as long as the angle θ(ψ,ρ)\theta(\psi,\rho) is retained. This means that class-specific structures can only be resolved implicitly through relations with proxies of different classes. Secondly, this intraclass resolution becomes worse as training progress, since proxies from different classes continue to contrast further and sample-proxy relations for same-class pairs are overshadowed (see e.g. Eq. 1, 5, 6). Similarly, for samples closer towards each respective class proxy, resolving local structures becomes harder. Effectively, this results in learned sample distributions to have a strong affinity towards local isotropy.

As such, proxy-based objectives are inherently handicapped in resolving local intraclass clusters and structures. Consequently, semantic relations between samples within a class are not well encoded in the representational structure of the learned deep metric spaces. However, the ability to explain and represent intraclass sample relations has been consistently shown to be a key driver for downstream generalization performance in DML .

While sample-based objectives similarly suffer from the non-bijectivity of s(∙,∙)s(\bullet,\bullet), the usage of sample-to-sample operations explicitly introduces intersample relational context that allows for better structuring of samples within a class. However, just incorporating a sample-based contrastive operation into the training process is not a sufficient remedy as it re-introduces the sampling complexity issue; whose removal was what made proxy-based methods and their fast convergence attractive in the first place.

2 Non-isotropy Regularization

Motivation. To address the non-learnability of intraclass context in proxy-based DML, we therefore have to address the inherent issue of local isotropy in the learned sample-proxy distribution p(ψ∣ρ)p(\psi|\rho). However, to retain the convergence (and generalization) benefits of proxy-based methods, this has to be achieved without resorting to additional augmentations that move the overall objective away from a purely proxy-based one. As such, we aim to find p(ψ∣ρ)p(\psi|\rho) whose optimization better resolves the distribution of sample representations ψ\psi around our proxy ρ\rho. This can be achieved by breaking down the fundamental issue of non-bijectivity in the used similarity measure s(ψ,ρ)s(\psi,\rho), which (on its own) introduces non-unique sample-proxy relations. To do so, we look for some regularization that specifically encourages unique sample-proxy relations to exists. For such unique sample-proxy relations to exist, we must have access to some bijective and thus invertible (deterministic) translation ψ=τ(ζ∣ρ)\psi=\tau(\zeta|\rho) which, given some residual ζ\zeta from some prior distribution q(ζ)q(\zeta), allows to uniquely translate from the respective proxy ρ\rho to ψ\psi. Given such a unique translation of samples and proxies within a class, the local alignment of samples would then no longer rely on relations to proxies and samples from different classes which, as noted, do not scale well locally and as training progresses. Assuming proxies and sample representation to have the same dimensionality, this can be achieved through some affine transformation. However, to capture non-linear relations and proxy-to-sample translations, it is much more beneficial for τ\tau to be non-linear.

Normalizing Flows. Such invertible, non-linear functions are naturally expressed through Normalizing Flows (NF) or more generally invertible neural networks . A Normalizing Flow can be generally seen as a transformation between two probability distributions, most commonly between simple, well-defined ones and complex multimodal ones . More specifically, we leverage flows similar to the one proposed in and (and as used e.g. in ), which introduces a sequence of non-linear, but still invertible coupling operations as showcased in Fig. 2 (⇋\leftrightharpoons). Given some input representation ψ\psi, a coupling block splits ψ\psi into ψ1\psi_{1} and ψ2\psi_{2}, which are scaled and translated in succession with non-linear scaling and translation networks η1\eta_{1} and η2\eta_{2}, respectively. Note that following , each network ηi\eta_{i} provides both scaling ηis\eta_{i}^{s} and translation values ηit\eta_{i}^{t}, such that

where ψ∗\psi^{*} denotes ψ\psi after passing through a respective coupling block. Successive application of different ηi\eta_{i} then gives our non-linear invertible transformation τ\tau from some prior distribution over residuals q(ζ)q(\zeta) with explicit density and CDF (for sampling) to our target distribution.

with Jacobian JJ for translation τ−1\tau^{-1} and proxies ρyx\rho_{y_{x}}, where yxy_{x} denotes the class of sample xx. To arrive at above equation, we simply leveraged the change of variables formula

In practice, by setting our prior q(ζ)q(\zeta) to be a standard zero-mean unit-variance normal distribution N(0,1)\mathcal{N}(0,1), we get

i.e. given sample representations ψ(x)\psi(x), we project them onto our residual space ζ\zeta via τ−1\tau^{-1} and compute Eq. 10. By selecting suitable normalizing flows such as GLOW , we make sure that the Jacobian is cheap to compute.

Experiments

Implementations use PyTorch . ImageNet-pretrainings are taken from torchvision and timm. Our experiments were run on compute servers with NVIDIA 2080Ti. Our Normalizing Flow utilizes 8 coupling blocks and subnets η\eta comprising linear layers with 128 nodes. Optimization is done using Adam (learning rate 10−510^{-5}, weight decay 4⋅10−34\cdot 10^{-3}, ). We set ω∈[0.001,0.01]\omega\in[0.001,0.01] depending on the choice for LPDML\mathcal{L}_{\text{PDML}}. In general, we found consistent improvements in this interval. Following we utilize a high learning rate multiplication for the proxies (4000). We also saw this helping the Normalizing Flow and use 50 for all experiments. Finally, we found a warmup epoch to help; adapting the translation τ\tau over pretrained features first before joint training.

Datasets. We use the standard benchmarks CUB200-2011 (11,788 bird images, 200 classes), CARS196 (16,185 car images, 196 classes) and Stanford Online Products (SOP, 120,053 images, 22,634 product instances).

2 Effectiveness of Non-isotropy Regularisation

3 Convergence Properties

4 Qualitative differences in alignment

5 Ablations

6 Self-regularization

Conclusion

Acknowledgements

This work has been partially funded by the ERC (853489 - DEXIM) and by the DFG (2064/1 – Project number 390727645). We thank the International Max Planck Research School for Intelligent Systems (IMPRS-IS) for supporting Karsten Roth. Karsten Roth further acknowledges his membership in the European Laboratory for Learning and Intelligent Systems (ELLIS) PhD program.

References