Unsupervised Learning of Invariant Representations in Hierarchical Architectures

Fabio Anselmi, Joel Z. Leibo, Lorenzo Rosasco, Jim Mutch, Andrea Tacchetti, Tomaso Poggio

Invariant representations and sample complexity

One could argue that the most important aspect of intelligence is the ability to learn. How do present supervised learning algorithms compare with brains? One of the most obvious differences is the ability of people and animals to learn from very few labeled examples. A child, or a monkey, can learn a recognition task from just a few examples. The main motivation of this paper is the conjecture that the key to reducing the sample complexity of object recognition is invariance to transformations. Images of the same object usually differ from each other because of simple transformations such as translation, scale (distance) or more complex deformations such as viewpoint (rotation in depth) or change in pose (of a body) or expression (of a face).

The conjecture is supported by previous theoretical work showing that almost all the complexity in recognition tasks is often due to the viewpoint and illumination nuisances that swamp the intrinsic characteristics of the object . It implies that in many cases, recognition—i.e., both identification, e.g., of a specific car relative to other cars—as well as categorization, e.g., distinguishing between cars and airplanes—would be much easier (only a small number of training examples would be needed to achieve a given level of performance, i.e. n→1n\rightarrow 1), if the images of objects were rectified with respect to all transformations, or equivalently, if the image representation itself were invariant. In SI Appendix, section 0 we provide a proof of the conjecture for the special case of translation (and for obvious generalizations of it).

The case of identification is obvious since the difficulty in recognizing exactly the same object, e.g., an individual face, is only due to transformations. In the case of categorization, consider the suggestive evidence from the classification task in Fig. 2. The figure shows that if an oracle factors out all transformations in images of many different cars and airplanes, providing “rectified” images with respect to viewpoint, illumination, position and scale, the problem of categorizing cars vs airplanes becomes easy: it can be done accurately with very few labeled examples. In this case, good performance was obtained from a single training image of each class, using a simple classifier. In other words, the sample complexity of the problem seems to be very low. We propose that the ventral stream in visual cortex tries to approximate such an oracle, providing a quasi-invariant signature for images and image patches.

Invariance and uniqueness

How can two orbits be characterized and compared? There are several possible approaches. A distance between orbits can be defined in terms of a metric on images, but its computation is not obvious (especially by neurons). We follow here a different strategy: intuitively two empirical orbits are the same irrespective of the ordering of their points. This suggests that we consider the probability distribution PIP_{I} induced by the group’s action on images II (gIgI can be seen as a realization of a random variable). It is possible to prove (see Theorem 2 in SI Appendix section 2) that if two orbits coincide then their associated distributions under the group GG are identical, that is

The distribution PIP_{I} is thus invariant and discriminative, but it also inhabits a high-dimensional space and is therefore difficult to estimate. In particular, it is unclear how neurons or neuron-like elements could estimate it.

As argued later, neurons can effectively implement (high-dimensional) inner products, ⟨⋅,⋅⟩\left\langle{\cdot},{\cdot}\right\rangle, between inputs and stored “templates” which are neural images. It turns out that classical results (such as the Cramer-Wold theorem , see Theorem 3 and 4 in section 2 of SI Appendix) ensure that a probability distribution PIP_{I} can be almost uniquely characterized by KK one-dimensional probability distributions P⟨I,tk⟩P_{\left\langle{I},{t^{k}}\right\rangle} induced by the (one-dimensional) results of projections ⟨I,tk⟩{\left\langle{I},{t^{k}}\right\rangle}, where tk,  k=1,...,Kt^{k},\;k=1,...,K are a set of randomly chosen images called templates. A probability function in dd variables (the image dimensionality) induces a unique set of 1-D projections which is discriminative; empirically a small number of projections is usually sufficient to discriminate among a finite number of different probability distributions. Theorem 4 in SI Appendix section 2 says (informally) that an approximately invariant and unique signature of an image II can be obtained from the estimates of KK 1-D probability distributions P⟨I,tk⟩P_{\left\langle{I},{t^{k}}\right\rangle} for k=1,⋯ ,Kk=1,\cdots,K. The number KK of projections needed to discriminate nn orbits, induced by nn images, up to precision ϵ\epsilon (and with confidence 1−δ21-\delta^{2}) is K≥2cϵ2log⁡nδK\geq\frac{2}{c\epsilon^{2}}\log{\frac{n}{\delta}}, where cc is a universal constant.

Thus the discriminability question can be answered positively (up to ϵ\epsilon) in terms of empirical estimates of the one-dimensional distributions P⟨I,tk⟩P_{\left\langle{I},{t^{k}}\right\rangle} of projections of the image onto a finite number of templates tk, k=1,...,Kt^{k},~{}k=1,...,K under the action of the group.

Memory-based learning of invariance

Notice that the estimation of P⟨I,tk⟩P_{\left\langle{I},{t^{k}}\right\rangle} requires the observation of the image and “all” its transforms gIgI. Ideally, however, we would like to compute an invariant signature for a new object seen only once (e.g., we can recognize a new face at different distances after just one observation, i.e. n→1n\rightarrow 1). It is remarkable and almost magical that this is also made possible by the projection step. The key is the observation that ⟨gI,tk⟩=⟨I,g−1tk⟩\left\langle{gI},{t^{k}}\right\rangle=\left\langle{I},{g^{-1}t^{k}}\right\rangle. The same one-dimensional distribution is obtained from the projections of the image and all its transformations onto a fixed template, as from the projections of the image onto all the transformations of the same template. Indeed, the distributions of the variables ⟨I,g−1tk⟩\left\langle{I},{g^{-1}t^{k}}\right\rangle and ⟨gI,tk⟩\left\langle{gI},{t^{k}}\right\rangle are the same. Thus it is possible for the system to store for each template tkt^{k} all its transformations gtkgt^{k} for all g∈Gg\in G and later obtain an invariant signature for new images without any explicit knowledge of the transformations gg or of the group to which they belong. Implicit knowledge of the transformations, in the form of the stored templates, allows the system to be automatically invariant to those transformations for new inputs (see eq. $$ in SI Appendix).

Estimates of the one-dimensional probability density functions (PDFs) P⟨I,tk⟩P_{\left\langle{I},{t^{k}}\right\rangle} can be written in terms of histograms as μnk(I)=1/∣G∣∑i=1∣G∣ηn(⟨I,gitk⟩)\mu_{n}^{k}(I)=1/|G|\sum_{i=1}^{|G|}\eta_{n}(\left\langle{I},{g_{i}t^{k}}\right\rangle), where ηn,  n=1,⋯ ,N\eta_{n},\;n=1,\cdots,N is a set of nonlinear functions (see remark 1 in SI Appendix section 1 or Theorem 6 in section 2 but also ). A visual system need not recover the actual probabilities from the empirical estimate in order to compute a unique signature. The set of μnk(I)\mu_{n}^{k}(I) values is sufficient, since it identifies the associated orbit (see box 1 in SI Appendix). Crucially, mechanisms capable of computing invariant representations under affine transformations for future objects can be learned and maintained in an unsupervised, automatic way by storing and updating sets of transformed templates which are unrelated to those future objects.

A theory of pooling

The arguments above make a few predictions. They require an effective normalization of the elements of the inner product (e.g. ⟨I,gitk⟩↦⟨I,gitk⟩∥I∥∥gitk∥\left\langle{I},{g_{i}t^{k}}\right\rangle\mapsto\frac{\left\langle{I},{g_{i}t^{k}}\right\rangle}{\|I\|\|g_{i}t^{k}\|}) for the property ⟨gI,tk⟩=⟨I,g−1tk⟩\left\langle{gI},{t^{k}}\right\rangle=\left\langle{I},{g^{-1}t^{k}}\right\rangle to be valid (see remark 8 of SI Appendix section 1 for the affine transformations case). Notice that invariant signatures can be computed in several ways from one-dimensional probability distributions. Instead of the μnk(I)\mu_{n}^{k}(I) components directly representing the empirical distribution, the moments mnk(I)=1/∣G∣∑i=1∣G∣(⟨I,gitk⟩)nm^{k}_{n}(I)=1/|G|\sum_{i=1}^{|G|}(\left\langle{I},{g_{i}t^{k}}\right\rangle)^{n} of the same distribution can be used (this corresponds to the choice ηn(⋅)≡(⋅)n\eta_{n}(\cdot)\equiv(\cdot)^{n} ). Under weak conditions, the set of all moments uniquely characterizes the one-dimensional distribution P⟨I,tk⟩P_{\left\langle{I},{t^{k}}\right\rangle} (and thus PIP_{I}). n=1n=1 corresponds to pooling via sum/average (and is the only pooling function that does not require a nonlinearity); n=2n=2 corresponds to ”energy models” of complex cells and n=∞n=\infty is related to max-pooling. In our simulations, just one of these moments usually seems to provide sufficient selectivity to a hierarchical architecture (see SI Appendix section 6). Other nonlinearities are also possible. The arguments of this section begin to provide a theoretical understanding of “pooling”, giving insight into the search for the “best” choice in any particular setting—something which is normally done empirically . According to this theory, these different pooling functions are all invariant, each one capturing part of the full information contained in the PDFs.

Implementations

The theory has strong empirical support from several specific implementations which have been shown to perform well on a number of databases of natural images. The main support is provided by HMAX, an architecture in which pooling is done with a max operation and invariance, to translation and scale, is mostly hardwired (instead of learned). Its performance on a variety of tasks is discussed in SI Appendix section 6. Good performance is also achieved by other very similar architectures . This class of existing models inspired the present theory, and may now be seen as special cases of it. Using the principles of invariant recognition the theory makes explicit, we have now begun to develop models that incorporate invariance to more complex transformations which cannot be solved by the architecture of the network, but must be learned from examples of objects undergoing transformations. These include non-affine and even non-group transformations, allowed by the hierarchical extension of the theory (see below). Performance for one such model is shown in Figure 3 (see caption for details).

Extensions of the Theory

The core of the theory applies without qualification to compact groups such as rotations of the image in the image plane. Translation and scaling are however only locally compact, and in any case, each of the modules of Fig. 1 observes only a part of the transformation’s full range. Each ⋀\bigwedge-module has a finite pooling range, corresponding to a finite “window” over the orbit associated with an image. Exact invariance for each module, in the case of translations or scaling transformations, is equivalent to a condition of localization/sparsity of the dot product between image and template (see Theorem 6 and Fig. 5 in section 2 of SI Appendix). In the simple case of a group parameterized by one parameter rr the condition is (for simplicity II and tt have support center in zero):

Since this condition is a form of sparsity of the generic image II w.r.t. a dictionary of templates tkt^{k} (under a group), this result provides a computational justification for sparse encoding in sensory cortex.

It turns out that localization yields the following surprising result (Theorem 7 and 8 in SI Appendix): optimal invariance for translation and scale implies Gabor functions as templates. Since a frame of Gabor wavelets follows from natural requirements of completeness, this may also provide a general motivation for the Scattering Transform approach of Mallat based on wavelets .

The same Equation 2, if relaxed to hold approximately, that is ⟨IC,grtk⟩≈0∣r∣>a\left\langle{I_{C}},{g_{r}t^{k}}\right\rangle\approx 0\quad|r|>a, becomes a sparsity condition for the class of ICI_{C} w.r.t. the dictionary tkt^{k} under the group GG when restricted to a subclass ICI_{C} of similar images. This property (see SI Appendix, end of section 2), which is an extension of the compressive sensing notion of “incoherence”, requires that II and tkt^{k} have a representation with sharply peaked correlation and autocorrelation. When the condition is satisfied, the basic HW-module equipped with such templates can provide approximate invariance to non-group transformations such as rotations in depth of a face or its changes of expression (see Proposition 9, section 2, SI Appendix). In summary, Equation 2 can be satisfied in two different regimes. The first one, exact and valid for generic II, yields optimal Gabor templates. The second regime, approximate and valid for specific subclasses of II, yields highly tuned templates, specific for the subclass. Note that this argument suggests generic, Gabor-like templates in the first layers of the hierarchy and highly specific templates at higher levels. (Note also that incoherence improves with increasing dimensionality.)

2 Hierarchical architectures

The property of compositionality discussed above is related to the efficacy of hierarchical architectures vs. one-layer architectures in dealing with the problem of partial occlusion and the more difficult problem of clutter in object recognition. Hierarchical architectures are better at recognition in clutter than one-layer networks because they provide signatures for image patches of several sizes and locations. However, hierarchical feedforward architectures cannot fully solve the problem of clutter. More complex (e.g. recurrent) architectures are likely needed for human-level recognition in clutter (see for instance ) and for other aspects of human vision. It is likely that much of the circuitry of visual cortex is required by these recurrent computations, not considered in this paper.

Visual Cortex

The theory described above effectively maps the computation of an invariant signature onto well-known capabilities of cortical neurons. A key difference between the basic elements of our digital computers and neurons is the number of connections: 33 vs. 103−10410^{3}-10^{4} synapses per cortical neuron. Taking into account basic properties of synapses, it follows that a single neuron can compute high-dimensional (103−10410^{3}-10^{4}) inner products between input vectors and the stored vector of synaptic weights . Consider an HW-module of “simple” and “complex” cells looking at the image through a window defined by their receptive fields (see SI Appendix, section 2, POG). Suppose that images of objects in the visual environment undergo affine transformations. During development—and more generally, during visual experience—a set of ∣G∣|G| simple cells store in their synapses an image patch tkt^{k} and its transformations g1tk,...,g∣G∣tkg_{1}t^{k},...,g_{|G|}t^{k}—one per simple cell. This is done, possibly at separate times, for KK different image patches tkt^{k} (templates), k=1,⋯ ,Kk=1,\cdots,K. Each gtkgt^{k} for g∈Gg\in G is a sequence of frames, literally a movie of image patch tkt^{k} transforming. There is a very simple, general, and powerful way to learn such unconstrained transformations. Unsupervised (Hebbian) learning is the main mechanism: for a “complex” cell to pool over several simple cells, the key is an unsupervised Foldiak-type rule: cells that fire together are wired together. At the level of complex cells this rule determines classes of equivalence among simple cells – reflecting observed time correlations in the real world, that is, transformations of the image. Time continuity, induced by the Markovian physics of the world, allows associative labeling of stimuli based on their temporal contiguity.

Later, when an image is presented, the simple cells compute ⟨I,gitk⟩\left\langle{I},{g_{i}t^{k}}\right\rangle for i=1,...,∣G∣i=1,...,|G|. The next step, as described above, is to estimate the one-dimensional probability distribution of such a projection, that is, the distribution of the outputs of the simple cells. It is generally assumed that complex cells pool the outputs of simple cells. Thus a complex cell could compute μnk(I)=1/∣G∣∑i=1∣G∣σ(⟨I,gitk⟩+nΔ)\mu_{n}^{k}(I)=1/|G|\sum_{i=1}^{|G|}\sigma(\left\langle{I},{g_{i}t^{k}}\right\rangle+n\Delta) where σ\sigma is a smooth version of the step function (σ(x)=0\sigma(x)=0 for x≤0x\leq 0, σ(x)=1\sigma(x)=1 for x>0x>0) and n=1,...,Nn=1,...,N (this corresponds to the choice ηn(⋅)≡σ(⋅+nΔ)\eta_{n}(\cdot)\equiv\sigma(\cdot+n\Delta)) . Each of these NN complex cells would estimate one bin of an approximated CDF (cumulative distribution function) for P⟨I,tk⟩P_{\left\langle{I},{t^{k}}\right\rangle}. Following the theoretical arguments above, the complex cells could compute, instead of an empirical CDF, one or more of its moments. n=1n=1 is the mean of the dot products, n=2n=2 corresponds to an energy model of complex cells ; very large nn corresponds to a maxmax operation. Conventional wisdom interprets available physiological data to suggest that simple/complex cells in V1 may be described in terms of energy models, but our alternative suggestion of empirical histogramming by sigmoidal nonlinearities with different offsets may fit the diversity of data even better.

As described above, a template and its transformed versions may be learned from unsupervised visual experience through Hebbian plasticity. Remarkably, our analysis and empirical studies show that Hebbian plasticity, as formalized by Oja, can yield Gabor-like tuning—i.e., the templates that provide optimal invariance to translation and scale (see SI Appendix section 2).

The localization condition (Equation 2) can also be satisfied by images and templates that are similar to each other. The result is invariance to class-specific transformations. This part of the theory is consistent with the existence of class-specific modules in primate cortex such as a face module and a body module . It is intriguing that the same localization condition suggests general Gabor-like templates for generic images in the first layers of a hierarchical architecture and specific, sharply tuned templates for the last stages of the hierarchy. This theory also fits physiology data concerning Gabor-like tuning in V1 and possibly in V4 (see ). It can also be shown that the theory, together with the hypothesis that storage of the templates takes place via Hebbian synapses, also predicts properties of the tuning of neurons in the face patch AL of macaque visual cortex .

From the point of view of neuroscience, the theory makes a number of predictions, some obvious, some less so. One of the main predictions is that simple and complex cells should be found in all visual and auditory areas, not only in V1. Our definition of simple cells and complex cells is different from the traditional ones used by physiologists; for example, we propose a broader interpretation of complex cells, which in the theory represent invariant measurements associated with histograms of the outputs of simple cells or of moments of it. The theory implies that invariance to all image transformations could be learned, either during development or in adult life. It is, however, also consistent with the possibility that basic invariances may be genetically encoded by evolution but also refined and maintained by unsupervised visual experience. Studies on the development of visual invariance in organisms such as mice raised in virtual environments could test these predictions.

Discussion

The goal of this paper is to introduce a new theory of learning invariant representations for object recognition which cuts across levels of analysis . At the computational level, it gives a unified account of why a range of seemingly different models have recently achieved impressive results on recognition tasks. HMAX , Convolutional Neural Networks and Deep Feedforward Neural Networks are examples of this class of architectures—as is, possibly, the feedforward organization of the ventral stream. At the algorithmic level, it motivates the development, now underway, of a new class of models for vision and speech which includes the previous models as special cases. At the level of biological implementation, its characterization of the optimal tuning of neurons in the ventral stream is consistent with the available data on Gabor-like tuning in V1 and the more specific types of tuning in higher areas such as in face patches.

Despite significant advances in sensory neuroscience over the last five decades, a true understanding of the basic functions of the ventral stream in visual cortex has proven to be elusive. Thus it is interesting that the theory of this paper follows from a novel hypothesis about the main computational function of the ventral stream: the representation of new objects/images in terms of a signature which is invariant to transformations learned during visual experience, thereby allowing recognition from very few labeled examples—in the limit, just one. A main contribution of our work to machine learning is a novel theoretical framework for the next major challenge in learning theory beyond the supervised learning setting which is now relatively mature: the problem of representation learning, formulated here as the unsupervised learning of invariant representations that significantly reduce the sample complexity of the supervised learning stage.

Supplementary Information

Invariance significantly reduces sample complexity

For example, in the case of linear learning rules, the sample complexity is proportional to the logarithm of the covering number.

Consider the simplest and most intuitive example: an image made of a single pixel and its translations in a square of dimension p×pp\times p, where p2=dp^{2}=d. In the pixel basis the space of the image and all its translates has dimension p2p^{2} meanwhile the image dimension is one. The associated covering numbers are therefore

where NI\mathcal{N}^{I} stands for the covering number of the image space and NTI\mathcal{N}^{TI} the covering number of the translated image space. The sample complexity associated to the image space (see e.g. ) is O(1)O(1) and that associated to the translated images O(p2)O(p^{2}). The sample complexity reduction of an invariant representation is therefore given by

The above reasoning is independent on the choice of the basis since it depends only on the dimensionality of the ball containing all the images. For example we could have determined the dimensionality looking the cardinality of eigenvectors (with non null eigenvalue) associated to a circulant matrix of dimension p×pp\times p i.e. using the Fourier basis. In the simple case above, the cardinality is clearly p2p^{2}. In general any transformation of an abelian group can be analyzed using the Fourier transform on the group. We conjecture that a similar reasoning holds for locally compact groups using a wavelet representation instead of the Fourier representation. The example and ideas above leads to the following theorem:

Consider a space of images of dimensions p×pp\times p pixels which may appear in any position within a window of size rp×rprp\times rp pixels. The usual image representation yields a sample complexity (of a linear classifier) of order mimage=O(r2p2)m_{image}=O(r^{2}p^{2}); the invariant representation yields (because of much smaller covering numbers) a sample complexity of order

Setup and Definitions

Normalized dot products of signals (e.g. images or “neural activities”) are usually assumed throughout the theory, for convenience but also because they provide the most elementary invariances – to measurement units (origin and scale). We assume that the dot products are between functions or vectors that are zero-mean and of unit norm. Thus ⟨I,t⟩\left\langle{I},{t}\right\rangle sets I=I′−I′ˉ∥I′−I′ˉ∥I=\frac{I^{\prime}-\bar{I^{\prime}}}{\left\lVert{I^{\prime}-\bar{I^{\prime}}}\right\rVert}, t=t′−t′ˉ∥t′−t′ˉ∥t=\frac{t^{\prime}-\bar{t^{\prime}}}{\left\lVert{t^{\prime}-\bar{t^{\prime}}}\right\rVert} with (⋅)ˉ\bar{(\cdot)} the mean. This normalization stage before each dot product is consistent with the convention that the empty surround of an isolated image patch has zero value (which can be taken to be the average “grey” value over the ensemble of images). In particular the dot product of a template – in general different from zero – and the “empty” region outside an isolated image patch will be zero. The dot product of two uncorrelated images – for instance of random 2D noise – is also approximately zero. Remarks:

The kk-th component of the signature associated with a simple-complex module is (see Equation (10) or (13)) \mu_{n}^{k}(I)=\frac{1}{|G_{0}|}\sum_{g\in{G}_{0}}\eta_{n}\Big{(}\left\langle{gI},{t^{k}}\right\rangle\Big{)} where the functions ηn\eta_{n} are such that Ker(ηn)={0}Ker(\eta_{n})=\{0\}: in words, the empirical histogram estimated for ⟨gI,tk⟩\left\langle{gI},{t^{k}}\right\rangle does not take into account the value, since it does not carry any information about the image patch. The functions ηn\eta_{n} are also assumed to be positive and bijective.

The domain of the dot products ⟨gI,tk⟩\left\langle{gI},{t^{k}}\right\rangle corresponding to templates and to simple cells is in general different from the domain of the pooling ∑g∈G0\sum_{g\in{G}_{0}}. We will continue to use the commonly used term receptive field – even if it mixes these two domains.

The main part of the theory characterizes properties of the basic HW module – which computes the components of an invariant signature vector from an image patch within its receptive field.

It is important to emphasize that the basic module is always the same throughout the paper. We use different mathematical tools, including approximations, to study under which conditions (e.g. localization or linearization, see end of section 2) the signature computed by the module is invariant or approximatively invariant.

The pooling ∑g∈G0,  G0⊆G\sum_{g\in{G}_{0}},\;G_{0}\subseteq G is effectively over a pooling window in the group parameters. In the case of 1D scaling and 1D translations, the pooling window corresponds to an interval, e.g. [aj,aj+k][a^{j},a^{j+k}], of scales and an interval, e.g. [−xˉ,xˉ][-\bar{x},\bar{x}], of xx translations, respectively.

All the results in this paper are valid in the case of a discrete or a continuous compact (locally compact) group: in the first case we have a sum over the transformations, in the second an integral over the Haar measure of the group.

Normalized dot products also eliminate the need of the explicit computation of the determinant of the Jacobian for affine transformations (which is a constant and is simplified dividing by the norms) assuring that ⟨AI,At⟩=⟨I,t⟩\left\langle{AI},{At}\right\rangle=\left\langle{I},{t}\right\rangle, where AA is an affine transformation.

Invariance and uniqueness: Basic Module

Given an image I∈XI\in\mathcal{X} and a group representation gg, the orbit OI={I′∈X  s.t.  I′=gI,g∈G}O_{I}=\{I^{\prime}\in\mathcal{X}~{}~{}s.t.~{}~{}I^{\prime}=gI,g\in G\} is uniquely associated to an image and all its transformations. The orbit provides an invariant representation of II, i.e. OI=OgIO_{I}=O_{gI} for all g∈Gg\in G. Indeed, we can view an orbit as all the possible realizations of a random variable with distribution PIP_{I} induced by the group action. From this observation, a signature Σ(I)\Sigma(I) can be derived for compact groups, by using results characterizing probability distributions via their one dimensional projections. In this section we study the signature given by

2 Orbits and probability distributions

The distribution PIP_{I} is invariant and unique i.e. I∼I′  ⇔  PI=PI′I\sim I^{\prime}\;\Leftrightarrow\;P_{I}=P_{I^{\prime}}.

Proof: We first prove that I∼I′  ⇒  PI=PI′I\sim I^{\prime}\;\Rightarrow\;P_{I}=P_{I^{\prime}}. By definition PI=PI′P_{I}=P_{I^{\prime}} iff ∫AdPI(s)=∫AdPI′(s)\int_{A}dP_{I}(s)=\int_{A}dP_{I^{\prime}}(s), ∀  A⊆X\forall\;A\subseteq{\cal X}, that is ∫ZI−1(A)dg=∫ZI′−1(A)dg\int_{Z^{-1}_{I}(A)}dg=\int_{Z^{-1}_{I^{\prime}}(A)}dg, where,

∀  A⊆X\forall\;A\subseteq{\cal X}. Note that ∀  A⊆X\forall\;A\subseteq{\cal X} if gI∈A  ⇒  ggˉ−1gˉI=ggˉ−1I′∈AgI\in A\;\Rightarrow\;g\bar{g}^{-1}\bar{g}I=g\bar{g}^{-1}I^{\prime}\in A, so that g∈ZI−1(A)  ⇒  ggˉ−1∈ZI′−1(A)g\in Z^{-1}_{I}(A)\;\Rightarrow\;g\bar{g}^{-1}\in Z^{-1}_{I^{\prime}}(A), i.e. ZI−1(A)⊆ZI′−1(A)Z^{-1}_{I}(A)\subseteq Z^{-1}_{I^{\prime}}(A). Conversely g∈ZI′−1(A)  ⇒  ggˉ∈ZI−1(A)g\in Z^{-1}_{I^{\prime}}(A)\;\Rightarrow\;g\bar{g}\in Z^{-1}_{I}(A), so that ZI−1(A)=ZI′−1(A)gˉ,  ∀AZ^{-1}_{I}(A)=Z^{-1}_{I^{\prime}}(A)\bar{g},\;\forall A. Using this observation we have,

where in the last integral we used the change of variable g^=ggˉ−1\hat{g}=g\bar{g}^{-1} and the invariance property of the Haar measure: this proves the implication. To prove that PI=PI′  ⇒  I∼I′P_{I}=P_{I^{\prime}}\;\Rightarrow\;I\sim I^{\prime}, note that PI(A)−P_{I}(A)- PI′(A)=0P_{I^{\prime}}(A)=0 for some A⊆XA\subseteq{\cal X} implies that the support of the probability distribution of II has non null intersection with that of I′I^{\prime} i.e. the orbits of II and I′I^{\prime} intersect. In other words there exist g′,g′′∈Gg^{\prime},g^{\prime\prime}\in{G} such that g′I=g′′I′g^{\prime}I=g^{\prime\prime}I^{\prime}. This implies I=g′−1g′′I′=gˉI′,  gˉ=g′−1g′′I={g^{\prime}}^{-1}g^{\prime\prime}I^{\prime}=\bar{g}I^{\prime},\;\bar{g}={g^{\prime}}^{-1}g^{\prime\prime}, i.e. I∼I′I\sim I^{\prime}. Q.E.D.

3 Random Projections for Probability Distributions.

In words, two probability distributions are equal if and only if their projections on any of the unit sphere directions is equal. The above result can be equivalently stated as saying that the probability of choosing tt such that P⟨I,t⟩=Q⟨I,t⟩P_{\left\langle{I},{t}\right\rangle}=Q_{\left\langle{I},{t}\right\rangle} is equal to 1 if and only if P=QP=Q and the probability of choosing tt such that P⟨I,t⟩=Q⟨I,t⟩P_{\left\langle{I},{t}\right\rangle}=Q_{\left\langle{I},{t}\right\rangle} is equal to 0 if and only if P≠QP\neq Q (see Theorem 3.4 in ). The theorem suggests a way to define a metric on distributions (orbits) in terms of

where d0d_{0} is any metric on one dimensional probability distributions and dλ(t)d\lambda(t) is a distribution measure on the projections. Indeed, it is easy to check that dd is a metric. In particular note that, in view of the Cramer Wold Theorem, d(P,Q)=0d(P,Q)=0 if and only if P=QP=Q. As mentioned in the main text, each one dimensional distribution P⟨I,t⟩P_{\left\langle{I},{t}\right\rangle} can be approximated by a suitable histogram μt(I)=(μnt(I))n=1,…,N∈RN\mu^{t}(I)=(\mu_{n}^{t}(I))_{n=1,\dots,N}\in R^{N}, so that, in the limit in which the histogram approximation is accurate

where dμd_{\mu} is a metric on histograms induced by d0d_{0}.

A natural question is whether there are situations in which a finite number of projections suffice to discriminate any two probability distributions, that is PI≠PI′⇔d(PI,PI′)≠0P_{I}\neq P_{I}^{\prime}\Leftrightarrow d(P_{I},P_{I^{\prime}})\neq 0. Empirical results show that this is often the case with a small number of templates (see and HMAX experiments, section 6). The problem of mathematically characterizing the situations in which a finite number of (one-dimensional) projections are sufficient is challenging. Here we provide a partial answer to this question. We start by observing that the metric (3) can be approximated by uniformly sampling KK templates and considering

where μk=μtk\mu^{k}=\mu^{t^{k}}. The following result shows that a finite number KK of templates is sufficient to obtain an approximation within a given precision ϵ\epsilon. Towards this end let

Consider nn images Xn{\cal X}_{n} in X\cal X. Let K≥2cϵ2log⁡nδK\geq\frac{2}{c\epsilon^{2}}\log{\frac{n}{\delta}}, where cc is a universal constant. Then

with probability 1−δ21-\delta^{2}, for all I,I′∈XnI,I^{\prime}\in{\cal X}_{n}.

4 Memory based learning of invariance

The signature Σ(I)=(μ11(I),…,μNK(I))\Sigma(I)=(\mu_{1}^{1}(I),\dots,\mu^{K}_{N}(I)) is obviously invariant (and unique) since it is associated to an image and all its transformations (an orbit). Each component of the signature is also invariant – it corresponds to a group average. Indeed, each measurement can be defined as

when GG is a (locally) compact group. Here, the non linearity ηn\eta_{n} can be chosen to define an histogram approximation; in general is a bijective positive function. Then, it is clear that from the properties of the Haar measure we have

Note that in the r.h.s. of eq. (8) the transformations are on templates: this mathematically trivial (for unitary transformations) step has a deeper computational aspect. Invariance is now in fact achieved through transformations of templates instead of those of the image, not always available.

5 Stability

and assume that μnk(I)=∫dg  ηn(⟨gI,tk⟩)\mu_{n}^{k}(I)=\int dg\;\eta_{n}(\left\langle{gI},{t^{k}}\right\rangle) for n=1,…,Nn=1,\dots,N and k=1,…,Kk=1,\dots,K. If L<1L<1 we call the signature map contractive. In the following we prove a stronger form of eq. 10 where the L2L^{2} norm is substituted with the Hausdorff norm on the orbits (which is independent of the choice of II and I′I^{\prime} in the orbits) defined as ∥I−I′∥H=ming,g′∈G∥gI−g′I′∥2\left\lVert{I-I^{\prime}}\right\rVert_{H}=min_{g,g^{\prime}\in G}\left\lVert{gI-g^{\prime}I^{\prime}}\right\rVert_{2}, I,I′∈XI,I^{\prime}\in\mathcal{X}, i.e. we have:

Assume normalized templates and let Lη=max⁡n(Lηn)L_{\eta}=\max_{n}(L_{\eta_{n}}) s.t. NLη≤1NL_{\eta}\leq 1, where LηnL_{\eta_{n}} is the Lipschitz constant of the function ηn\eta_{n}. Then

Proof: By definition, if the non linearities ηn\eta_{n} are Lipschitz continuous, for all n=1,…,Nn=1,\dots,N, with Lipschitz constant LηnL_{\eta_{n}}, it follows that for each kk component of the signature we have

where we used the linearity of the inner product and Jensen’s inequality. Applying Schwartz’s inequality we obtain

where Lη=max⁡n(Lηn)L_{\eta}=\max_{n}(L_{\eta_{n}}). If we assume the templates and their transformations to be normalized to unity then we finally have,

In particular being NLη<1NL_{\eta}<1 the map is non expansive summing each component and dividing by 1/K1/K we have eq. (11). Q.E.D. The above result shows that the stability of the empirical signature

provided with the metric (4) (together with (5)) holds for nonlinearities with Lipschitz constants LηnL_{\eta_{n}} such that Nmaxn(Lηn)<1Nmax_{n}(L_{\eta_{n}})<1.

6 Partially Observable Groups case: invariance implies localization and sparsity

This section outlines invariance, uniqueness and stability properties of the signature obtained in the case in which transformations of a group are observable only within a window “over” the orbit. The term POG (Partially Observable Groups) emphasizes the properties of the group – in particular associated invariants – as seen by an observer (e.g. a neuron) looking through a window at a part of the orbit. Let GG be a finite group and G0⊆GG_{0}\subseteq G a subset (note: G0G_{0} is not usually a subgroup). The subset of transformations G0G_{0} can be seen as the set of transformations that can be observed by a window on the orbit that is the transformations that correspond to a part of the orbit. A local signature associated to the partial observation of GG can be defined considering

and ΣG0(I)=(μnk(I))n,k\Sigma_{G_{0}}(I)=(\mu_{n}^{k}(I))_{n,k}. This definition can be generalized to any locally compact group considering,

Note that the constant V0V_{0} normalizes the Haar measure, restricted to G0G_{0}, so that it defines a probability distribution. The latter is the distribution of the images subject to the group transformations which are observable, that is in G0G_{0}. The above definitions can be compared to definitions (7) and (8) in the fully observable groups case. In the next sections we discuss the properties of the above signature. While stability and uniqueness follow essentially from the analysis of the previous section, invariance requires developing a new analysis.

7 POG: Stability and Uniqueness

A direct consequence of Theorem 9.2 is that any two orbits with a common point are identical. This follows from the fact that if gI,g′I′gI,g^{\prime}I^{\prime} is a common point of the orbits, then

Thus the two images are transformed versions of one another and OI=OI′O_{I}=O_{I^{\prime}}. Suppose now that only a fragment of the orbits – the part within the window – is observable; the reasoning above is still valid since if the orbits are different or equal so must be any of their “corresponding” parts. Regarding the stability of POG signatures, note that the reasoning in the previous section, Theorem 9.5, can be repeated without any significant change. In fact, only the normalization over the transformations is modified accordingly.

8 POG: Partial Invariance and Localization

Since the group is only partially observable we introduce the notion of partial invariance for images and transformations G0G_{0} that are within the observation window. Partial invariance is defined in terms of invariance of

We recall that when gIgI and tkt^{k} do not share any common support on the plane or II and tt are uncorrelated, then ⟨gI,tk⟩=0\left\langle{gI},{t^{k}}\right\rangle=0. The following theorem, where G0G_{0} corresponds to the pooling range states, a sufficient condition for partial invariance in the case of a locally compact group:

Proof: To prove the implication note that if ⟨gI,tk⟩=0,  ∀g∈G/(G0∩gˉG0)\left\langle{gI},{t^{k}}\right\rangle=0,\;\forall g\in G/(G_{0}\cap\bar{g}G_{0}), being G0ΔgˉG0⊆G/(G0∩gˉG0)G_{0}\Delta\bar{g}G_{0}\subseteq G/(G_{0}\cap\bar{g}G_{0}) (Δ\Delta is the symbol for symmetric difference (AΔB=(A∪B)/(A∩B)    A,B    setsA\Delta B=(A\cup B)/(A\cap B)\;\;A,B\;\;sets) we have:

The second equality is true since, being ηn\eta_{n} positive, the fact that the integral is zero implies ⟨gI,tk⟩=0  ∀g∈G/(G0∩gˉG0)\left\langle{gI},{t^{k}}\right\rangle=0\;\forall g\in G/(G_{0}\cap\bar{g}G_{0}) (and therefore in particular ∀g∈G0ΔgˉG0\forall g\in G_{0}\Delta\bar{g}G_{0}). Being the r.h.s. of the inequality positive, we have

i.e. μnk(I)=μnk(gˉI)\mu^{k}_{n}(I)=\mu^{k}_{n}(\bar{g}I) (see also Fig. 5 for a visual explanation). Q.E.D. Equation (• ‣ 4. Synopsis of Mathematical Results) describes a localization condition on the inner product of the transformed image and the template. The above result naturally raises question of weather the localization condition is also necessary for invariance. Clearly, this would be the case if eq. (9.8) could be turned into an equality, that is

Indeed, in this case, if μnk(I)−μnk(gˉI)=0\mu^{k}_{n}(I)-\mu^{k}_{n}(\bar{g}I)=0, and we further assume the natural condition ⟨gI,tk⟩≠0\left\langle{gI},{t^{k}}\right\rangle\neq 0 if and only if g∈G0g\in G_{0}, then the localization condition (• ‣ 4. Synopsis of Mathematical Results) would be necessary since ηn\eta_{n} is a positive bijective function. The equality in eq. (9.8) in general is not true. However, this is clearly the case if we consider the group of transformations to be translations as illustrated in Fig. 7 a). We discuss in some details this latter case.

for a given cc where TxT_{x} is a unitary representation of the translation operator. We can view S,ScS,S_{c} as sets of simple responses to a given template through two receptive fields. Let S0={⟨TxI,t⟩ :   x∈[0,a+c]}S_{0}=\{\left\langle{T_{x}I},{t}\right\rangle~{}:~{}\;x\in[0,a+c]\}, so that S,Sc⊂S0S,S_{c}\subset S_{0} for all cc. We assume that S0,S,ScS_{0},S,S_{c} to be closed intervals for all cc. Then, recall that a bijective function (in this case ηn\eta_{n}) is strictly monotonic on any closed interval so that the difference of integrals in eq. (9.8) is zero if and only if S=ScS=S_{c}. Since we are interested in considering all the values of cc up to some maximum CC, then we can consider the condition

The above condition can be satisfied in two cases: 1) both dot products are zero, which is the localization condition, or 2) TaI=IT_{a}I=I (or equivalently Tat=tT_{a}t=t) i.e. the image or the template are periodic. A similar reasoning applies to the case of scale transformations.

In the next paragraph we will see how localization conditions for scale and translation transformations imply a specific form of the templates.

In this section we identify G0G_{0} with subsets of the affine group. In particular, we study separately the case of scale and translations (in 1D for simplicity).

In the following it is helpful to assume that all images II and templates tt are strictly contained in the range of translation or scale pooling, PP, since image components outside it are not measured. We will consider images II restricted to PP: for translation this means that the support of II is contained in PP, for scaling, since gsI=I(sx)g_{s}I=I(sx) and I(sx)^=(1/s)I^(ω/s)\widehat{I(sx)}=(1/s)\hat{I}(\omega/s) (where ⋅^\hat{\cdot} indicates the Fourier transform), assuming a scale pooling range of [sm,sM][s_{m},s_{M}], implies a range [ωmI,ωMI],  [ωmt,ωMt][\omega^{I}_{m},\omega^{I}_{M}],\;[\omega_{m}^{t},\omega_{M}^{t}] (mm and MM indicates maximum and minimum) of spatial frequencies for the maximum support of II and tt. As we will see because of Theorem 9.6 invariance to translation requires spatial localization of images and templates and less obviously invariance to scale requires bandpass properties of images and templates. Thus images and templates are assumed to be localized from the outset in either space or frequency. The corollaries below show that a stricter localization condition is needed for invariance and that this condition determines the form of the template. Notice that in our framework images and templates are bandpass because of being zero-mean. Notice that, in addition, neural “images” which are input to the hierarchical architecture are spatially bandpass because of retinal processing.

with S>1S>1. Localization conditions of the support of the dot product for translation and scale are depicted in Figure 7,a) ,b). As shown by the following Lemma 9.7 Eq. (22) and (23) gives interesting conditions on the supports of tt and its Fourier transform t^\hat{t}. For translation, the corollary is equivalent to zero overlap of the compact supports of II and tt. In particular using Theorem 9.6, for I=tI=t, the maximal invariance in translation implies the following localization conditions on tt

Invariance to translation in the range [0,xˉ],  xˉ>0[0,\bar{x}],\;\bar{x}>0 is equivalent to the following localization condition of tt in space

Separately, invariance to dilations in the range [1,sˉ],  sˉ>1[1,\bar{s}],\;\bar{s}>1 is equivalent to the following localization condition of t^\hat{t} in frequency ω\omega

Proof: To prove that supp(t)⊆[−b+xˉ,b]−supp(I)supp(t)\subseteq[-b+\bar{x},b]-supp(I) note that eq. (22) implies that supp(⟨TxI,t⟩)⊆[−b+xˉ,b]supp(\left\langle{T_{x}I},{t}\right\rangle)\subseteq[-b+\bar{x},b] (see Figure 7, a)). In general supp(⟨TxI,t⟩)=supp(I∗t)⊆supp(I)+supp(t)supp(\left\langle{T_{x}I},{t}\right\rangle)=supp(I*t)\subseteq supp(I)+supp(t). The inclusion account for the fact that the integral ⟨TxI,t⟩\left\langle{T_{x}I},{t}\right\rangle can be zero even if the supports of TxIT_{x}I and tt are not disjoint. However if we suppose invariance for a continuous set of translations xˉ∈[0,Xˉ]\bar{x}\in[0,\bar{X}], (where, for any given I,tI,t, Xˉ\bar{X} is the maximum translation for which we have an invariant signature) and for a generic image in X\mathcal{X} the inclusion become an equality, since for the invariance condition in Theorem 9.6 we have

which is possible, given the arbitrariness of xˉ\bar{x} and II only if

where we used the property supp(Txf)=Txf,∀f∈Xsupp(T_{x}f)=T_{x}f,\forall f\in\mathcal{X}. Being, under these conditions, supp(⟨TxI,t⟩)supp(\left\langle{T_{x}I},{t}\right\rangle) =supp(I)+supp(t)=supp(I)+supp(t) we have supp(t)⊆[−b−xˉ,b]−supp(I)supp(t)\subseteq[-b-\bar{x},b]-supp(I), i.e. eq (25). To prove the condition in eq. (9.7) note that eq. (23) is equivalent in the Fourier domain to

The situation is depicted in Fig. 7 b′)b^{\prime}) for SS big enough: in this case in fact we can suppose the support of Dsˉ/SI^\widehat{D_{\bar{s}/S}I} to be on an interval on the left of that of supp(t^)supp(\hat{t}) and DSI^\widehat{D_{S}I} on the right; the condition supp(⟨DsI^,t^⟩)⊆[sˉ/S,S]supp(\left\langle{\widehat{D_{s}I}},{\hat{t}}\right\rangle)\subseteq[\bar{s}/S,S] is in this case equivalent to

and therefore eq. (9.7). Note that for some s∈[sˉ/S,S]s\in[\bar{s}/S,S] the condition that the Fourier supports are disjoint is only sufficient and not necessary for the dot product to be zero since cancelations can occur. However we can repeat the reasoning done for the translation case and ask for ⟨DsI^,t^⟩=0\left\langle{\widehat{D_{s}I}},{\hat{t}}\right\rangle=0 on a continuous interval of scales.Q.E.D.

The results above lead to a statement connecting invariance with localization of the templates:

Maximum translation invariance implies a template with minimum support in the space domain (xx); maximum scale invariance implies a template with minimum support in the Fourier domain (ω\omega).

Proof: We illustrate the statement of the theorem with a simple example. In the case of translations suppose, e.g., supp(I)=[−b′,b′],  supp(t)=[−a,a]supp(I)=[-b^{\prime},b^{\prime}],\;supp(t)=[-a,a], a≤b′≤ba\leq b^{\prime}\leq b. Eq. (25) reads

which gives the condition −a≥−b+b′+xˉ-a\geq-b+b^{\prime}+\bar{x}, i.e. xˉmax=b−b′−a\bar{x}^{max}=b-b^{\prime}-a; thus, for any fixed b,b′b,b^{\prime} the smaller the template support, 2a2a, in space, the greater is translation invariance. Similarly, in the case of dilations, increasing the range of invariance [1,sˉ],  sˉ>1[1,\bar{s}],\;\bar{s}>1 implies a decrease in the support of t^\hat{t} as shown by eq. (29); in fact noting that ∣supp(t^)∣=2Δt|supp(\hat{t})|=2\Delta_{t} we have

i.e. the measure, ∣⋅∣|\cdot|, of the support of t^\hat{t} is a decreasing function w.r.t. the measure of the invariance range [1,sˉ][1,\bar{s}]. Q.E.D. Because of the assumption of maximum possible support of all II being finite there is always localization for any choice of II and tt under spatial shift. Of course if the localization support is larger than the pooling range there is no invariance. For a complex cell with pooling range [−b,b][-b,b] in space only templates with self-localization smaller than the pooling range make sense. An extreme case of self-localization is t(x)=δ(x)t(x)=\delta(x), corresponding to maximum localization of tuning of the simple cells.

9 Invariance, Localization and Wavelets

The conditions equivalent to optimal translation and scale invariance – maximum localization in space and frequency – cannot be simultaneously satisfied because of the classical uncertainty principle: if a function t(x)t(x) is essentially zero outside an interval of length Δx\Delta x and its Fourier transform I^(ω)\hat{I}(\omega) is essentially zero outside an interval of length Δω\Delta\omega then

In other words a function and its Fourier transform cannot both be highly concentrated. Interestingly for our setup the uncertainty principle also applies to sequences.

It is well known that the equality sign in the uncertainty principle above is achieved by Gabor functions of the form

The uncertainty principle leads to the concept of “optimal localization” instead of exact localization. In a similar way, it is natural to relax our definition of strict invariance (e.g. μnk(I)=μnk(g′I)\mu_{n}^{k}(I)=\mu_{n}^{k}(g^{\prime}I)) and to introduce ϵ\epsilon-invariance as ∣μnk(I)−μnk(g′I)∣≤ϵ|\mu_{n}^{k}(I)-\mu_{n}^{k}(g^{\prime}I)|\leq\epsilon. In particular if we suppose, e.g., the following localization condition

where erf is the error function. The differences above, with an opportune choice of the localization ranges σs,σx\sigma_{s},\sigma_{x} can be made as small as wanted. We end this paragraph by a conjecture: the optimal ϵ−\epsilon-invariance is satisfied by templates with non compact support which decays exponentially such as a Gaussian or a Gabor wavelet. We can then speak of optimal invariance meaning “optimal ϵ\epsilon-invariance”. The reasonings above lead to the theorem:

Assume invariants are computed from pooling within a pooling window with a set of linear filters. Then the optimal templates (e.g. filters) for maximum simultaneous invariance to translation and scale are Gabor functions

The Gabor function ψx0,ω0(x)\psi_{x_{0},\omega_{0}}(x) corresponds to a Heisenberg box which has a xx-spread σx2=∫x2∣ψ(x)∣dx\sigma_{x}^{2}=\int x^{2}|\psi(x)|dx and a ω\omega spread σω2=∫ω2∣ψ^(ω)∣dω\sigma_{\omega}^{2}=\int\omega^{2}|\hat{\psi}(\omega)|d\omega with area σxσω\sigma_{x}\sigma_{\omega}. Gabor wavelets arise under the action on ψ(x)\psi(x) of the translation and scaling groups as follows. The function ψ(x)\psi(x), as defined, is zero-mean and normalized that is

A family of Gabor wavelets is obtained by translating and scaling ψ\psi:

Under certain conditions (in particular, the Heisenberg boxes associated with each wavelet must together cover the space-frequency plane) the Gabor wavelet family becomes a Gabor wavelet frame.

Optimal self-localization of the templates (which follows from localization), when valid simultaneously for space and scale, is also equivalent to Gabor wavelets. If they are a frame, full information can be preserved in an optimal quasi invariant way.

Note that the result proven in Theorem 9.9 is not related to that in . While Theorem 9.9 shows that Gabor functions emerge from the requirement of maximal invariance for the complex cells – a property of obvious computational importance – the main result in shows that (a) the wavelet transform (not necessarily Gabor) is covariant with the similitude group and (b) the wavelet transform follows from the requirement of covariance (rather than invariance) for the simple cells (see our definition of covariance in the next section). While (a) is well-known (b) is less so (but see p.41 in ). Our result shows that Gabor functions emerge from the requirement of maximal invariance for the complex cells.

10 Approximate Invariance and Localization

In the previous section we analyzed the relation between localization and invariance in the case of group transformations. By relaxing the requirement of exact invariance and exact localization we show how the same strategy for computing invariants can still be applied even in the case of non-group transformations if certain localization properties of ⟨TI,t⟩\left\langle{TI},{t}\right\rangle holds, where TT is a smooth transformation (to make it simple think to a transformation parametrized by one parameter).

We first notice that the localization condition of theorems 9.6 and 9.9 – when relaxed to approximate localization – takes the (e.g. for the 1D1D translations group supposing for simplicity that the supports of II and tt are centered in zero) form ⟨I,Txtk⟩<δ∀x  s.t.  ∣x∣>a\left\langle{I},{T_{x}t^{k}}\right\rangle<\delta\quad\forall x\;s.t.\;|x|>a, where δ\delta is small in the order of 1/n1/\sqrt{n} (where nn is the dimension of the space) and ⟨I,Txtk⟩≈1∀x  s.t.  ∣x∣<a\left\langle{I},{T_{x}t^{k}}\right\rangle\approx 1\quad\forall x\;s.t.\;|x|<a. We call this property sparsity of II in the dictionary tkt^{k} under GG. This condition can be satisfied by templates that are similar to images in the set and are sufficiently “rich” to be incoherent for “small” transformations. Note that from the reasoning above the sparsity of II in tkt^{k} under GG is expected to improve with increasing nn and with noise-like encoding of II and tkt^{k} by the architecture. Another important property of sparsity of II in tkt^{k} (in addition to allowing local approximate invariance to arbitrary transformations, see later) is clutter-tolerance in the sense that if n1,n2n_{1},n_{2} are additive uncorrelated spatial noisy clutter ⟨I+n1,gtk+n2⟩≈⟨I,gt⟩\left\langle{I+n_{1}},{gt^{k}+n_{2}}\right\rangle\approx\left\langle{I},{gt}\right\rangle. Interestingly the sparsity condition under the group is related to associative memories for instance of the holographic type,. If the sparsity condition holds only for I=tkI=t^{k} and for very small set of g∈Gg\in G, that is, it has the form ⟨I,gtk⟩=δ(g)δI,tk\left\langle{I},{gt^{k}}\right\rangle=\delta(g)\delta_{I,t^{k}} it implies strict memory-based recognition ( see non-interpolating look-up table in the description of ) with inability to generalize beyond stored templates or views.

While the first regime – exact (or ϵ−\epsilon-) invariance for generic images, yielding universal Gabor templates – applies to the first layer of the hierarchy, this second regime (sparsity) – approximate invariance for a class of images, yielding class-specific templates – is important for dealing with non-group transformations at the top levels of the hierarchy where receptive fields may be as large as the visual field.

Several interesting transformations do not have the group structure, for instance the change of expression of a face or the change of pose of a body. We show here that approximate invariance to transformations that are not groups can be obtained if the approximate localization condition above holds, and if the transformation can be locally approximated by a linear transformation, e.g. a combination of translations, rotations and non-homogeneous scalings, which corresponds to a locally compact group admitting a Haar measure.

where R(I)R(I) is the reminder, ee is the identity operator, JIJ^{I} the Jacobian and LrI=e+JIrL^{I}_{r}=e+J^{I}r is a linear operator. Let RR be the range of the parameter rr where we can approximately neglect the remainder term R(I)R(I). Let LL be the range of the parameter rr where the scalar product ⟨TrI,t⟩\left\langle{T_{r}I},{t}\right\rangle is localized i.e. ⟨TrI,t⟩=0,  ∀r∉L\left\langle{T_{r}I},{t}\right\rangle=0,\;\forall r\not\in L. If L⊆RL\subseteq R we have

then μnk(TrˉI)=μnk(I)\mu^{k}_{n}(T_{\bar{r}}I)=\mu^{k}_{n}(I).

Proof: We have, following the reasoning done in Theorem 9.6

Summary of the argument: Different transformations can be classified in terms of invariance and localization.

Compact Groups: consider the case of a compact group transformation such as rotation in the image plane. A complex cell is invariant when pooling over all the templates which span the full group θ∈[−π,+π]\theta\in[-\pi,+\pi]. In this case there is no restriction on which images can be used as templates: any template yields perfect invariance over the whole range of transformations (apart from mild regularity assumptions) and a single complex cell pooling over all templates can provide a globally invariant signature.

Locally Compact Groups and Partially Observable Compact Groups: consider now the POG situation in which the pooling is over a subset of the group: (the POG case always applies to Locally Compact groups (LCG) such as translations). As shown before, a complex cell is partially invariant if the value of the dot-product between a template and its shifted template under the group falls to zero fast enough with the size of the shift relative to the extent of pooling.

In the POG and LCG case, such partial invariance holds over a restricted range of transformations if the templates and the inputs have a localization property that implies wavelets for transformations that include translation and scaling.

General (non-group) transformations: consider the case of a smooth transformation which may not be a group. Smoothness implies that the transformation can be approximated by piecewise linear transformations, each centered around a template (the local linear operator corresponds to the first term of the Taylor series expansion around the chosen template). Assume – as in the POG case – a special form of sparsity – the dot-product between the template and its transformation fall to zero with increasing size of the transformation. Assume also that the templates transform as the input image. For instance, the transformation induced on the image plane by rotation in depth of a face may have piecewise linear approximations around a small number of key templates corresponding to a small number of rotations of a given template face (say at ±30o,±90o,±120o\pm 30^{o},\pm 90^{o},\pm 120^{o}). Each key template and its transformed templates within a range of rotations corresponds to complex cells (centered in ±30o,±90o,±120o\pm 30^{o},\pm 90^{o},\pm 120^{o}). Each key template, e.g. complex cell, corresponds to a different signature which is invariant only for that part of rotation. The strongest hypothesis is that there exist input images that are sparse w.r.t. templates of the same class – these are the images for which local invariance holds. Remarks:

We are interested in two main cases of POG invariance:

partial invariance simultaneously to translations in x,yx,y, scaling and possibly rotation in the image plane. This should apply to “generic” images. The signatures should ideally preserve full, locally invariant information. This first regime is ideal for the first layers of the multilayer network and may be related to Mallat’s scattering transform, . We call the sufficient condition for LCG invariance here, localization, and in particular, in the case of translation (scale) group self-localization given by Equation (24).

partial invariance to linear transformations for a subset of all images. This second regime applies to high-level modules in the multilayer network specialized for specific classes of objects and non-group transformations. The condition that is sufficient here for LCG invariance is given by Theorem 9.6 which applies only to a specific class of II. We prefer to call it sparsity of the images with respect to a set of templates.

For classes of images that are sparse with respect to a set of templates, the localization condition does not imply wavelets. Instead it implies templates that are

similar to a class of images so that ⟨I,g0tk⟩≈1\left\langle{I},{g_{0}t^{k}}\right\rangle\approx 1 for some g0∈Gg_{0}\in G and

complex enough to be “noise-like” in the sense that ⟨I,gtk⟩≈0\left\langle{I},{gt^{k}}\right\rangle\approx 0 for g≠g0g\neq g_{0}.

Templates must transform similarly to the input for approximate invariance to hold. This corresponds to the assumption of a class-specific module and of a nice object class .

For the localization property to hold, the image must be similar to the key template or contain it as a diagnostic feature (a sparsity property). It must be also quasi-orthogonal (highly localized) under the action of the local group.

For a general, non-group, transformation it may be impossible to obtain invariance over the full range with a single signature; in general several are needed.

It would be desirable to derive a formal characterization of the error in local invariance by using the standard module of dot-product-and-pooling, equivalent to a complex cell. The above arguments provide the outline of a proof based on local linear approximation of the transformation and on the fact that a local linear transformation is a LCG.

Hierarchical Architectures

So far we have studied the invariance, uniqueness and stability properties of signatures, both in the case when a whole group of transformations is observable (see (7) and (8)), and in the case in which it is only partially observable (see (13) and (14)). We now discuss how the above ideas can be iterated to define a multilayer architecture. Consider first the case when GG is finite. Given a subset G0⊂GG_{0}\subset G, we can associate a window gG0gG_{0} to each g∈Gg\in G. Then, we can use definition (13) to define for each window a signature Σ(I)(g)\Sigma(I)(g) given by the measurements,

We will keep this form as the definition of signature. For fixed n,kn,k, a set of measurements corresponding to different windows can be seen as a ∣G∣|G| dimensional vector. A signature Σ(I)\Sigma(I) for the whole image is obtained as a signature of signatures, that is, a collection of signatures (Σ(I)(g1),…,Σ(I)(g∣G∣)(\Sigma(I)(g_{1}),\dots,\Sigma(I)(g_{|G|}) associated to each window. Since we assume that the output of each module is made zero-mean and normalized before further processing at the next layer, conservation of information from one layer to the next requires saving the mean and the norm at the output of each module at each level of the hierarchy. We conjecture that the neural image at the first layer is uniquely represented by the final signature at the top of the hierarchy and the means and norms at each layer. The above discussion can be easily extended to continuous (locally compact) groups considering,

Similarly we have assumed that the number of non linearities ηn\eta_{n}, considered at every layer, is the same.

Following the above discussion, the extension to continuous (locally compact) groups is straightforward. We collect it in the following definition.

Remark Note that eq. (41) can be written as:

In the following we study the properties of the complex response, in particular

We call the map Σ\Sigma covariant under GG iff

The covariance property described in proposition 9.11 can be stated equivalently as μ1n,k(I)(g)=μ1n,k(gˉI)(gˉg)\mu_{1}^{n,k}(I)(g)=\mu_{1}^{n,k}(\bar{g}I)(\bar{g}g). This last expression has a more intuitive meaning as shown in Fig. 8.

The covariance property described in proposition 9.11 holds both for abelian and non-abelian groups. However the group average on templates transformations in definition of eq. 41 is crucial. In fact, if we define the signature averaging on the images we do not have a covariant response:

With respect to the range of invariance, the following property holds for multilayer architectures in which the output of a layer is defined as covariant if it transforms in the same way as the input: for a given transformation of an image or part of it, the signature from complex cells at a certain layer is either invariant or covariant with respect to the group of transformations; if it is covariant there will be a higher layer in the network at which it is invariant (more formal details are given in Theorem 9.13), assuming that the image is contained in the visual field. This property predicts a stratification of ranges of invariance in the ventral stream: invariances should appear in a sequential order meaning that smaller transformations will be invariant before larger ones, in earlier layers of the hierarchy.

12 Property 2: partial and global invariance (whole and parts)

The proof follows from the observation that the pooling range over the group is a bigger and bigger subset of GG with growing layer number, in other words, there exists a layer such that the image and its transformations are within the pooling range at that layer (see Fig. 9). This is clear since for any gˉ∈G\bar{g}\in G the nested sequence

13 Property 3: stability

Using the definition of stability given in (11), we can formulate the following theorem characterizing stability for the complex response:

The proof follows from a repeated application of the reasoning done in Theorem 9.5. See details in .

14 Comparison with stability as in [16]

Of course, by compactness of T×BT\times B and the C2\mathcal{C}_{2}-assumption, both KtK_{t} and KxK_{x} are finite. The following theorem is due to Maurer and Poggio:

The proof reveals this to be just a special case of Taylor’s theorem. Proof: Denote V(t,x)=(V1,...,Vl)(t,x)=(∂/∂t)Φ(t,x)V\left(t,x\right)=\left(V_{1},...,V_{l}\right)\left(t,x\right)=\left(\partial/\partial t\right)\Phi\left(t,x\right), V˙(t,x)=(V˙1,...,V˙l)(t,x)=(∂2/∂t2)Φ(t,x)\dot{V}\left(t,x\right)=\left(\dot{V}_{1},...,\dot{V}_{l}\right)\left(t,x\right)=\left(\partial^{2}/\partial t^{2}\right)\Phi\left(t,x\right), and set V:=V(0,0)V:=V\left(0,0\right). For s∈[0,1]s\in\left[0,1\right] we have with Cauchy-Schwartz

Theorem 9.15 and corollary 14 gives a precise mathematical motivation for the assumption that any sufficiently smooth (at least twice differentiable) transformation can be approximated in an enough small compact set with a group transformation (e.g. translation), thus allowing, based on eq. 11, stability w.r.t. small diffeomorphic transformations.

15 Approximate Factorization: hierarchy

In the first version of we conjectured that a signature invariant to a group of transformations could be obtained by factorizing in successive layers the computation of signatures invariant to a subgroup of the transformations (e.g. the subgroup of translations of the affine group) and then adding the invariance w.r.t. another subgroup (e.g. rotations). While factorization of invariance ranges is possible in a hierarchical architecture (theorem 9.13), it can be shown that in general the factorization in successive layers for instance of invariance to translation followed by invariance to rotation (by subgroups) is impossible. However, approximate factorization is possible under the same conditions of the previous section. In fact, a transformation that can be linearized piecewise can always be performed in higher layers, on top of other transformations, since the global group structure is not required but weaker smoothness properties are sufficient.

16 Why Hierarchical architectures: a summary

Optimization of local connections and optimal reuse of computational elements. Despite the high number of synapses on each neuron it would be impossible for a complex cell to pool information across all the simple cells needed to cover an entire image.

Compositionality. A hierarchical architecture provides signatures of larger and larger patches of the image in terms of lower level signatures. Because of this, it can access memory in a way that matches naturally with the linguistic ability to describe a scene as a whole and as a hierarchy of parts.

Approximate factorization. In architectures such as the network sketched in Fig. 1 in the main text, approximate invariance to transformations specific for an object class can be learned and computed in different stages. This property may provide an advantage in terms of the sample complexity of multistage learning . For instance, approximate class-specific invariance to pose (e.g. for faces) can be computed on top of a translation-and-scale-invariant representation . Thus the implementation of invariance can, in some cases, be “factorized” into different steps corresponding to different transformations. (see also for related ideas).

Probably all three properties together are the reason evolution developed hierarchies.

Synopsis of Mathematical Results

Orbits are equivalent to probability distributions, PIP_{I} and both are invariant and unique. Theorem The distribution PIP_{I} is invariant and unique i.e. I∼I′  ⇔  PI=PI′I\sim I^{\prime}\;\Leftrightarrow\;P_{I}=P_{I^{\prime}}.

PIP_{I} can be estimated within ϵ\epsilon in terms of 1D probability distributions of ⟨gI,tk⟩\left\langle{gI},{t^{k}}\right\rangle. Theorem Consider nn images Xn{\cal X}_{n} in X\cal X. Let K≥2cϵ2log⁡nδK\geq\frac{2}{c\epsilon^{2}}\log{\frac{n}{\delta}}, where cc is a universal constant. Then

with probability 1−δ21-\delta^{2}, for all I,I′∈XnI,I^{\prime}\in{\cal X}_{n}.

Invariance from a single image based on memory of template transformations. The simple property

implies (for unitary groups without any additional property) that the signature components \mu_{n}^{k}(I)=\frac{1}{|G|}\sum_{g\in{G}}\eta_{n}\Big{(}\left\langle{I},{gt^{k}}\right\rangle\Big{)}, calculated on templates transformations are invariant that is μnk(I)=μnk(gˉI)\mu_{n}^{k}(I)=\mu_{n}^{k}(\bar{g}I).

Condition in eq. (• ‣ 4. Synopsis of Mathematical Results) on the dot product between image and template implies invariance for Partially Observable Groups (observed through a window) and is equivalent to it in the case of translation and scale transformations.

Condition in Theorem 9.6 is equivalent to a localization or sparsity property of the dot product between image and template (⟨I,gt⟩=0\left\langle{I},{gt}\right\rangle=0 for g∉GLg\not\in G_{L}, where GLG_{L} is the subset of GG where the dot product is localized). In particular

Proposition Localization is necessary and sufficient for translation and scale invariance. Localization for translation (respectively scale) invariance is equivalent to the support of tt being small in space (respectively in frequency).

Optimal simultaneous invariance to translation and scale can be achieved by Gabor templates.

Theorem Assume invariants are computed from pooling within a pooling window a set of linear filters. Then the optimal templates of filters for maximum simultaneous invariance to translation and scale are Gabor functions t(x)=e−x22σ2eiω0xt(x)=e^{-\frac{x^{2}}{2\sigma^{2}}}e^{i\omega_{0}x}.

Approximate invariance can be obtained if there is approximate sparsity of the image in the dictionary of templates. Approximate localization (defined as ⟨t,gt⟩<δ\left\langle{t},{gt}\right\rangle<\delta for g∉GLg\not\in G_{L}, where δ\delta is small in the order of ≈1d\approx\frac{1}{\sqrt{d}} and ⟨t,gt⟩≈1\left\langle{t},{gt}\right\rangle\approx 1 for g∈GLg\in G_{L}) is satisfied by templates (vectors of dimensionality nn) that are similar to images in the set and are sufficiently “large” to be incoherent for “small” transformations.

Approximate invariance for smooth (non group) transformations.

Proposition μk(I)\mu^{k}(I) is locally invariant if

II and tkt^{k} transform in the same way (belong to the same class);

the transformation is sufficiently smooth.

Sparsity of II in the dictionary tkt^{k} under GG increases with size of the neural images and provides invariance to clutter.

The definition is ⟨I,gt⟩<δ\left\langle{I},{gt}\right\rangle<\delta for g∉GLg\not\in G_{L}, where δ\delta is small in the order of ≈1n\approx\frac{1}{\sqrt{n}} and ⟨I,gt⟩≈1\left\langle{I},{gt}\right\rangle\approx 1 for g∈GLg\in G_{L}. Sparsity of II in tkt^{k} under GG improves with dimensionality of the space nn and with noise-like encoding of II and tt. If n1,n2n_{1},n_{2} are additive uncorrelated spatial noisy clutter ⟨I+n1,gt+n2⟩≈⟨I,gt⟩\left\langle{I+n_{1}},{gt+n_{2}}\right\rangle\approx\left\langle{I},{gt}\right\rangle.

Factorization. Proposition Invariance to separate subgroups of affine group cannot be obtained in a sequence of layers while factorization of the ranges of invariance can (because of covariance). Invariance to a smooth (non group) transformation can always be performed in higher layers, on top of other transformations, since the global group structure is not required.

Uniqueness of signature. Conjecture:The neural image at the first layer is uniquely represented by the final signature at the top of the hierarchy and the means and norms at each layer.

General Remarks on the Theory

The second regime of localization (sparsity) can be considered as a way to deal with situations that do not fall under the general rules (group transformations) by creating a series of exceptions, one for each object class.

Whereas the first regime “predicts” Gabor tuning of neurons in the first layers of sensory systems, the second regime predicts cells that are tuned to much more complex features, perhaps similar to neurons in inferotemporal cortex.

The sparsity condition under the group is related to properties used in associative memories for instance of the holographic type (see ). If the sparsity condition holds only for I=tkI=t^{k} and for very small aa then it implies strictly memory-based recognition.

The theory is memory-based. It also view-based. Even assuming 3D images (for instance by using stereo information) the various stages will be based on the use of 3D views and on stored sequences of 3D views.

The mathematics of the class-specific modules at the top of the hierarchy – with the underlying localization condition – justifies old models of viewpoint-invariant recognition (see ).

The remark on factorization of general transformations implies that layers dealing with general transformations can be on top of each other. It is possible – as empirical results by Leibo and Li indicate – that a second layer can improve the invariance to a specific transformation of a lower layer.

The theory developed here for vision also applies to other sensory modalities, in particular speech.

The theory represents a general framework for using representations that are invariant to transformations that are learned in an unsupervised way in order to reduce the sample complexity of the supervised learning step.

Simple cells (e.g. templates) under the action of the affine group span a set of positions and scales and orientations. The size of their receptive fields therefore spans a range. The pooling window can be arbitrarily large – and this does not affect selectivity when the CDF is used for pooling. A large pooling window implies that the signature is given to large patches and the signature is invariant to uniform affine transformations of the patches within the window. A hierarchy of pooling windows provides signature to patches and subpatches and more invariance (to more complex transformations).

Connections with the Scattering Transform.

Our theorems about optimal invariance to scale and translation implying Gabor functions (first regime) may provide a justification for the use of Gabor wavelets by Mallat , that does not depend on the specific use of the modulus as a pooling mechanism.

Our theory justifies several different kinds of pooling of which Mallat’s seems to be a special case.

With the choice of the modulo as a pooling mechanisms, Mallat proves a nice property of Lipschitz continuity on diffeomorphisms. Such a property is not valid in general for our scheme where it is replaced by a hierarchical parts and wholes property which can be regarded as an approximation, as refined as desired, of the continuity w.r.t. diffeomorphisms.

Our second regime does not have an obvious corresponding notion in the scattering transform theory.

The theory characterizes under which conditions the signature provided by a HW module at some level of the hierarchy is invariant and therefore could be used for retrieving information (such as the label of the image patch) from memory. The simplest scenario is that signatures from modules at all levels of the hierarchy (possibly not the lowest ones) will be checked against the memory. Since there are of course many cases in which the signature will not be invariant (for instance when the relevant image patch is larger than the receptive field of the module) this scenario implies that the step of memory retrieval/classification is selective enough to discard efficiently the “wrong” signatures that do not have a match in memory. This is a nontrivial constraint. It probably implies that signatures at the top level should be matched first (since they are the most likely to be invariant and they are fewer) and lower level signatures will be matched next possibly constrained by the results of the top-level matches – in a way similar to reverse hierarchies ideas. It also has interesting implications for appropriate encoding of signatures to make them optimally quasi-orthogonal e.g. incoherent, in order to minimize memory interference. These properties of the representation depend on memory constraints and will be object of a future paper on memory modules for recognition.

There is psychophysical and neurophysiological evidence that the brain employs such learning rules (e.g. and references therein). A second step of Hebbian learning may be responsible for wiring a complex cells to simple cells that are activated in close temporal contiguity and thus correspond to the same patch of image undergoing a transformation in time . Simulations show that the system could be remarkably robust to violations of the learning rule’s assumption that temporally adjacent images correspond to the same object . The same simulations also suggest that the theory described here is qualitatively consistent with recent results on plasticity of single IT neurons and with experimentally-induced disruptions of their invariance .

Simple and complex units do not need to correspond to different cells: it is conceivable that a simple cell may be a cluster of synapses on a dendritic branch of a complex cell with nonlinear operations possibly implemented by active properties in the dendrites.

Unsupervised learning of the template orbit. While the templates need not be related to the test images (in the affine case), during development, the model still needs to observe the orbit of some templates. We conjectured that this could be done by unsupervised learning based on the temporal adjacency assumption . One might ask, do “errors of temporal association” happen all the time over the course of normal vision? Lights turn on and off, objects are occluded, you blink your eyes – all of these should cause errors. If temporal association is really the method by which all the images of the template orbits are associated with one another, why doesn’t the fact that its assumptions are so often violated lead to huge errors in invariance?

The full orbit is needed, at least in theory. In practice we have found that significant scrambling is possible as long as the errors are not correlated. That is, normally an HW-module would pool all the ⟨I,gitk⟩\left\langle{I},{g_{i}t^{k}}\right\rangle. We tested the effect of, for some ii, replacing tkt^{k} with a different template tk′t^{k^{\prime}}. Even scrambling 50%50\% of our model’s connections in this manner only yielded very small effects on performance. These experiments were described in more detail in for the case of translation. In that paper we modeled Li and DiCarlo’s ”invariance disruption” experiments in which they showed that a temporal association paradigm can induce individual IT neurons to change their stimulus preferences under specific transformation conditions . We also report similar results on another ”non-uniform template orbit sampling” experiment with 3D rotation-in-depth of faces in .

Empirical support for the theory

The theory presented here was inspired by a set of related computational models for visual recognition, dating from 1980 to the present day. While differing in many details, HMAX, Convolutional Networks , and related models use similar structural mechanisms to hierarchically compute translation (and sometimes scale) invariant signatures for progressively larger pieces of an input image, completely in accordance with the present theory.

With the theory in hand, and the deeper understanding of invariance it provides, we have now begun to develop a new generation of models that incorporate invariance to larger classes of transformations.

Fukushima’s Neocognitron was the first of a class of recognition models consisting of hierarchically stacked modules of simple and complex cells (a “convolutional” architecture). This class has grown to include Convolutional Networks, HMAX, and others . Many of the best performing models in computer vision are instances of this class. For scene classification with thousands of labeled examples, the best performing models are currently Convolutional Networks . A variant of HMAX scores 74% on the Caltech 101 dataset, competitive with the state-of-the-art for a single feature type. Another HMAX variant added a time dimension for action recognition , outperforming both human annotators and a state-of-the-art commercial system on a mouse behavioral phenotyping task. An HMAX model was also shown to account for human performance in rapid scene categorization. A simple illustrative empirical demonstration of the HMAX properties of invariance, stability and uniqueness is in figure 10.

All of these models work very similarly once they have been trained. They all have a convolutional architecture and compute a high-dimensional signature for an image in a single bottom-up pass. At each level, complex cells pool over sets of simple cells which have the same weights but are centered at different positions (and for HMAX, also scales). In the language of the present theory, for these models, gg is the 2D set of translations in xx and yy (3D if scaling is included), and complex cells pool over partial orbits of this group, typically outputting a single moment of the distribution, usually sum or max.

The biggest difference among these models lies in the training phase. The complex cells are fixed, always pooling only over position (and scale), but the simple cells learn their weights (templates) in a number of different ways. Some models assume the first level weights are Gabor filters, mimicking cortical area V1. Weights can also be learned via backpropagation, via sampling from training images, or even by generating random numbers. Common to all these models is the notion of automatic weight sharing: at each level ii of the hierarchy, the NiN_{i} simple cells centered at any given position (and scale) have the same set of NiN_{i} weight vectors as do the NiN_{i} simple cells for every other position (and scale). Weight sharing occurs by construction, not by learning, however, the resulting model is equivalent to one that learned by observing NiN_{i} different objects translating (and scaling) everywhere in the visual field.

One of the observations that inspired our theory is that in convolutional architectures, random features can often perform nearly as well as features learned from objects – the architecture often matters more than the particular features computed. We postulated that this was due to the paramount importance of invariance. In convolutional architectures, invariance to translation and scaling is a property of the architecture itself, and objects in images always transform and scale in the same way.

18 New models

Using the principles of invariant recognition made explicit by the present theory, we have begun to develop models that incorporate invariance to more complex transformations which, unlike translation and scaling, cannot be solved by the architecture of the network, but must be learned from examples of objects undergoing transformations. Two examples are listed here.

Faces rotating in 3D. In , we added a third H-W layer to an existing HMAX model which was already invariant to translation and scaling. This third layer modeled invariance to rotation in depth for faces. Rotation in depth is a difficult transformation due to self-occlusion. Invariance to it cannot be derived from network architecture, nor can it be learned generically for all objects. Faces are an important class for which specialized areas are known to exist in higher regions of the ventral stream. We showed that by pooling over stored views of template faces undergoing this transformation, we can recognize novel faces from a single example view, robustly to rotations in depth.

Faces undergoing unconstrained transformations. Another model inspired by the present theory recently advanced the state-of-the-art on the Labeled Faces in the Wild dataset, a challenging same-person / different-person task. Starting this time with a first layer of HOG features , the second layer of this model built invariance to translation, scaling, and limited in-plane rotation, leaving the third layer to pool over variability induced by other transformations. Performance results for this model are shown in figure 3 in the main text.

References