On the Continuity of Rotation Representations in Neural Networks

Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, Hao Li

Introduction

Recently, there has been an increasing number of applications in graphics and vision, where deep neural networks are used to perform regressions on rotations. This has been done for tasks such as pose estimation from images and from point clouds , structure from motion , and skeleton motion synthesis, which generates the rotations of joints in skeletons . Many of these works represent 3D rotations using 3D or 4D representations such as quaternions, axis-angles, or Euler angles.

However, for 3D rotations, we found that 3D and 4D representations are not ideal for network regression, when the full rotation space is required. Empirically, the converged networks still produce large errors at certain rotation angles. We believe that this actually points to deeper topological problems related to the continuity in the rotation representations. Informally, all else being equal, discontinuous representations should in many cases be “harder” to approximate by neural networks than continuous ones. Theoretical results suggest that functions that are smoother or have stronger continuity properties such as in the modulus of continuity have lower approximation error for a given number of neurons.

Based on this insight, we first present in Section 3 our definition of the continuity of representation in neural networks. We illustrate this definition based on a simple example of 2D rotations. We then connect it to key topological concepts such as homeomorphism and embedding.

Next, we present in Section 4 a theoretical analysis of the continuity of rotation representations. We first investigate in Section 4.1 some discontinuous representations, such as Euler angle and quaternion representations. We show that for 3D rotations, all representations are discontinuous in four or fewer dimensional real Euclidean space with the Euclidean topology. We then investigate in Section 4.2 some continuous rotation representations. For the nn dimensional rotation group SO(n)SO(n), we present a continuous n2−nn^{2}-n dimensional representation. We additionally present an option to reduce the dimensionality of this representation by an additional 11 to n−2n-2 dimensions in a continuous way. We show that these allow us to represent 3D rotations continuously in 6D and 5D. While we focus on rotations, we show how our continuous representations can also apply to other groups such as orthogonal groups O(n)O(n) and similarity transforms.

Finally, in Section 5 we test our ideas empirically. We conduct experiments on 3D rotations and show that our 6D and 5D continuous representations always outperform the discontinuous ones for several tasks, including a rotation autoencoder “sanity test,” rotation estimation for 3D point clouds, and 3D human pose inverse kinematics learning. We note that in our rotation autoencoder experiments, discontinuous representations can have up to 6 to 14 times higher mean errors than continuous representations. Furthermore they tend to converge much slower while still producing large errors over 170∘ at certain rotation angles even after convergence, which we believe are due to the discontinuities being harder to fit. This phenomenon can also be observed in the experiments on different rotation representations for homeomorphic variational auto-encoding in Falorsi et al. , and in practical applications, such as 6D object pose estimation in Xiang et al. .

We also show that one can perform direct regression on 3x3 rotation matrices. Empirically this approach introduces larger errors than our 6D representation as shown in Section 5.2. Additionally, for some applications such as inverse and forward kinematics, it may be important for the network itself to produce orthogonal matrices. We therefore require an orthogonalization procedure in the network. In particular, if we use a Gram-Schmidt orthogonalization, we then effectively end up with our 6D representation.

Our contributions are: 1) a definition of continuity for rotation representations, which is suitable for neural networks; 2) an analysis of discontinuous and continuous representations for 2D, 3D, and nn-D rotations; 3) new formulas for continuous representations of SO(3)SO(3) and SO(n)SO(n); 4) empirical results supporting our theoretical views and that our continuous representations are more suitable for learning.

Related Work

In this section, we will first establish some context for our work in terms of neural network approximation theory. Next, we discuss related works that investigate the continuity properties of different rotation representations. Finally, we will report the types of rotation representations used in previous learning tasks and their performance.

Neural network approximation theory. We review a brief sampling of results from neural network approximation theory. Hornik showed that neural networks can approximate functions in the LpL^{p} space to arbitrary accuracy if the LpL^{p} norm is used. Barron et al. showed that if a function has certain properties in its Fourier transform, then at most O(ϵ−2)O(\epsilon^{-2}) neurons are needed to obtain an order of approximation ϵ\epsilon. Chapter 6.4.1 of LeCun et al. provides a more thorough overview of such results. We note that results for continuous functions indicate that functions that have better smoothness properties can have lower approximation error for a particular number of neurons . For discontinuous functions, Llanas et al. showed that a real and piecewise continuous function can be approximated in an almost uniform way. However, Llanas et al. also noted that piecewise continuous functions when trained with gradient descent methods require many neurons and training iterations, and yet do not give very good results. These results suggest that continuous rotation representations might perform better in practice.

Continuity for rotations. Grassia et al. pointed out that Euler angles and quaternions are not suitable for orientation differentiation and integration operations and proposed exponential map as a more robust rotation representation. Saxena et al. observed that the Euler angles and quaternions cause learning problems due to discontinuities. However, they did not propose general rotation representations other than direct regression of 3x3 matrices, since they focus on learning representations for objects with specific symmetries.

Neural networks for 3D shape pose estimation. Deep networks have been applied to estimate the 6D poses of object instances from RGB images, depth maps or scanned point clouds. Instead of directly predicting 3x3 matrices that may not correspond to valid rotations, they typically use more compact rotation representations such as quaternion or axis-angle . In PoseCNN , the authors reported a high percentage of errors between 90∘ and 180∘, and suggested that this is mainly caused by the rotation ambiguity for some symmetric shapes in the test set. However, as illustrated in their paper, the proportion of errors between 90∘ to 180∘ is still high even for non-symmetric shapes. In this paper, we argue that discontinuity in these representations could be one cause of such errors.

Neural networks for inverse kinematics. Recently, researchers have been interested in training neural networks to solve inverse kinematics equations. This is because such networks are faster than traditional methods and differentiable so that they can be used in more complex learning tasks such as motion re-targeting and video-based human pose estimation . Most of these works represented rotations using quaternions or axis-angle . Some works also used other 3D representations such as Euler angles and Lie algebra , and penalized the joint position errors. Csiszar et al. designed networks to output the sine and cosine of the Euler angles for solving the inverse kinematics problems in robotic control. Euler angle representations are discontinuous for SO(3)SO(3) and can result in large regression errors as shown in the empirical test in Section 5. However, those authors limited the rotation angles to be within a certain range, which avoided the discontinuity points and thus achieved very low joint alignment errors in their test. However, many real-world tasks require the networks to be able to output the full range of rotations. In such cases, continuous rotation representations will be a better choice.

Definition of Continuous Representation

In this section, we begin by defining the terminology we will use in the paper. Next, we analyze a simple motivating example of 2D rotations. This allows us to develop our general definition of continuity of representation in neural networks. We then explain how this definition of continuity is related to concepts in topology.

Motivating example: 2D rotations. We now consider the representation of 2D rotations. For any 2D rotation M∈SO(2)M\in SO(2), we can also express the matrix as:

We can represent any rotation matrix M∈SO(2)M\in SO(2) by choosing θ∈R\theta\in R, where RR is a suitable set of angles, for example, R=[0,2π]R=[0,2\pi]. However, this particular representation intuitively has a problem with continuity. The problem is that if we define a mapping gg from the original space SO(2)SO(2) to the angular representation space RR, then this mapping is discontinuous. In particular, the limit of gg at the identity matrix, which represents zero rotation, is undefined: one directional limit gives an angle of and the other gives 2π2\pi. We depict this problem visually in Figure 1. On the right, we visualize a connected set of rotations C⊂SO(2)C\subset SO(2) by visualizing their first column vector [cos⁡(θ),sin⁡(θ)]T[\cos(\theta),\sin(\theta)]^{T} on the unit sphere S1S^{1}. On the left, after mapping them through gg, we see that the angles are disconnected. In particular, we say that this representation is discontinuous because the mapping gg from the original space to the representation space is discontinuous. We argue that these kind of discontinuous representations can be harder for neural networks to fit. Contrarily, if we represent the 2D rotation M∈SO(2)M\in SO(2) by its first column vector [cos⁡(θ),sin⁡(θ)]T[\cos(\theta),\sin(\theta)]^{T}, then the representation would be continuous.

Continuous representation: We can now define what we consider a continuous representation. We illustrate our definitions graphically in Figure 2. Let RR be a subset of a real vector space equipped with the Euclidean topology. We call RR the representation space: in our context, a neural network produces an intermediate representation in RR. This neural network is depicted on the left side of Figure 2. We will come back to this neural network shortly. Let XX be a compact topological space. We call XX the original space. In our context, any intermediate representation in RR produced by the network can be mapped into the original space XX. Define the mapping to the original space f:R→Xf:R\to X, and the mapping to the representation space g:X→Rg:X\to R. We say (f,g)(f,g) is a representation if for every x∈X,f(g(x))=xx\in X,f(g(x))=x, that is, ff is a left inverse of gg. We say the representation is continuous if gg is continuous.

Connection with neural networks: We now return to the neural network on the left side of Figure 2. We imagine that inference runs from left to right. Thus, the neural network accepts some input signals on its left hand side, outputs a representation in RR, and then passes this representation through the mapping ff to get an element of the original space XX. Note that in our context, the mapping f is implemented as a mathematical function that is used as part of the forward pass of the network at both training and inference time. Typically, at training time, we might impose losses on the original space XX. We now describe the intuition behind why we ask that gg be continuous. Suppose that we have some connected set CC in the original space, such as the one shown on the right side of Figure 1. Then if we map CC into representation space RR, and gg is continuous, then the set g(C)g(C) will remain connected. Thus, if we have continuous training data, then this will effectively create a continuous training signal for the neural network. Contrarily, if gg is not continuous, as shown in Figure 1, then a connected set in the original space may become disconnected in the representation space. This could create a discontinuous training signal for the network. We note that the units in the neural network are typically continuous, as defined on Euclidean topology spaces. Thus, we require the representation space RR to have Euclidean topology because this is consistent with the continuity of the network units.

Domain of the mapping ff: We additionally note that for neural networks, it is specifically beneficial for the mapping ff to be defined almost everywhere on a set where the neural network outputs are expected to lie. This enables ff to map arbitrary representations produced by the network back to the original space XX.

Connection with topology: Suppose that (f,g)(f,g) is a continuous representation. Note that gg is a continuous one-to-one function from a compact topological space to a Hausdorff space. From a theorem in topology , this implies that if we restrict the codomain of gg to g(X)g(X) (and use the subspace topology for g(X)g(X)) then the resulting mapping is a homeomorphism. A homeomorphism is a continuous bijection with a continuous inverse. For geometric intuition, a homeomorphism is often described as a continuous and invertible stretching and bending of one space to another, with also a finite number of cuts allowed if one later glues back together the same points. One says that two spaces are topologically equivalent if there is a homeomorphism between them. Additionally, gg is a topological embedding of the original space XX into the representation space RR. Note that we also have the inverse of gg: if we restrict ff to the domain g(X)g(X) then the resulting function f∣g(X)f|_{g(X)} is simply the inverse of gg. Conversely, if the original space XX is not homeomorphic to any subset of the representation space RR then there is no possible continuous representation (f,g)(f,g) on these spaces. We will return to this later when we show that there is no continuous representation for the 3D rotations in four or fewer dimensions.

Rotation Representation Analysis

Here we provide examples of rotation representations that could be used in networks. We start by looking in Section 4.1 at some discontinuous representations for 3D rotations, then look in Section 4.2 at continuous rotation representations in nn dimensions, and show that how for the 3D rotations, these become 6D and 5D continuous rotation representations. We believe this analysis can help one to choose suitable rotation representations for learning tasks.

Case 1: Euler angle representation for the 3D rotations. Let the original space X=SO(3),X=SO(3), the set of 3D rotations. Then we can easily show discontinuity in an Euler angle representation by considering the azimuth angle θ\theta and reducing this to the motivating example for 2D rotations shown in Section 3. In particular, the identity rotation II occurs at a discontinuity, where one directional limit gives θ=0\theta=0 and the other directional limit gives θ=2π\theta=2\pi. We visualize the discontinuities in this representation, and all the other representations, in the supplemental Section F.

Likewise, one can define the mapping to the original space SO(3)SO(3) as in :

Here the normalization function is defined as N(q)=q/∣∣q∣∣N(q)=q/||q||. By expanding in terms of the axis-angle representation for the matrix MM, one can verify that for every M∈SO(3),f(g(M))=M.M\in SO(3),f(g(M))=M.

2 Continuous Representations

In this section, we develop two continuous representations for the nn dimensional rotations SO(n)SO(n). We then explain how for the 3D rotations SO(3)SO(3) these become 6D and 5D continuous rotation representations.

Group operations such as multiplication: Suppose that the original space is a group such as the rotation group, and we want to multiply two representations r1,r2∈Rr_{1},r_{2}\in R. In general, we can do this by first mapping to the original space, multiplying the two elements, and then mapping back: r1r2=g(f(r1)f(r2)).r_{1}r_{2}=g(f(r_{1})f(r_{2})). However, for the proposed representation here, we can gain some computational efficiency as follows. Since the mapping to the representation space in Equation (5) drops the last column, when computing f(r2)f(r_{2}), we can simply drop the last column and compute the product representation as the product of an n×nn\times n and an n×(n−1)n\times(n-1) matrix.

Case 4: Further reducing the dimensionality for the nn dimensional rotations. For n≥3n\geq 3 dimensions, we can reduce the dimension for the representation in the previous case, while still keeping a continuous representation. Intuitively, a lower dimensional representation that is less redundant could be easier to learn. However, we found in our experiments that the dimension-reduced representation does not outperform the Gram-Schmidt-like representation from Case 3. However, we still develop this representation because it allows us to show that continuous rotation representations can outperform discontinuous ones.

Note that the un-projection is not actually back to the sphere, but in a way that coordinates 2 through mm are a unit vector. Now we can use between 11 and n−2n-2 normalized projections on the representation from the previous case, while still preserving continuity and one-to-one behavior.

For simplicity, we will first demonstrate the case of one stereographic projection. The idea is that we can flatten the representation from Case 3 to a vector and then stereographically project the last n+1n+1 components of that vector. Note that we intentionally project as few components as possible, since we found that nonlinearities introduced by the projection can make the learning process more difficult. These nonlinearities are due to the square terms and the division in Equation (9). If uu is a vector of length mm, define the slicing notation ui:j=(ui,ui+1,…,uj)u_{i:j}=(u_{i},u_{i+1},\ldots,u_{j}), and ui:=ui:mu_{i:}=u_{i:m}. Let M(i)M_{(i)} be the iith column of matrix M. Define a vectorized representation γ(M)\gamma(M) by dropping the last column of MM like in Equation (5): γ(M)=[M(1)T,…,M(n−1)T]\gamma(M)=[M_{(1)}^{T},\ldots,M_{(n-1)}^{T}]. Now we can define the mapping to the representation space as:

Here we have dropped the implicit argument MM to γ\gamma for brevity. Define the mapping to the original space as:

As a special case, for the 3D rotations, this gives us a 5D representation. This representation is made by using the 6D representation from Case 3, flattening it to a vector, and then using a normalized projection on the last 4 dimensions.

We can actually make up to n−2n-2 projections in a similar manner, while maintaining continuity of the representation, as follows. As a reminder, the length of γ\gamma, the vectorized result of the Gram-Schmidt process, is n(n−1)n(n-1): it contains n−1n-1 basis vectors each of dimension nn. Thus, we can make n−2n-2 projections, where each projection i=1,…,n−2i=1,\ldots,n-2 selects the basis vector i+1i+1 from γ(M)\gamma(M), prepends to it an appropriately selected element from the first basis vector of γ(M)\gamma(M), such as γn+1−i\gamma_{n+1-i}, and then projects the result. The resulting projections are then concatenated as a row vector along with the two unprojected entries to form the representation. Thus, after doing n−2n-2 projections, we can obtain a continuous representation for SO(n)SO(n) in n2−2n+2n^{2}-2n+2 dimensions. See Figure 4 for a visualization of the grouping of the elements that can be projected.

Empirical Results

We investigated different rotation representations and found that those with better continuity properties work better for learning. We first performed a sanity test and then experimented on two real world applications to show how continuity properties of rotation representations influence the learning process.

We first perform a sanity test using an auto-encoder structure. We use a multi-layer perceptron (MLP) network as an encoder to map SO(3)SO(3) to the chosen representation RR. We test our proposed 6D and 5D representations, quaternions, axis-angle, and Euler angles. The encoder network contains four fully-connected layers, where hidden layers have 128 neurons and Leaky ReLU activations. The fixed “decoder” mapping f:R↦SO(3)f:R\mapsto SO(3) is defined in Section 4.

For training, we compute the loss using the L2 distance between the input SO(3)SO(3) matrix MM and the output SO(3)SO(3) matrix M′M^{\prime}: note that this is invariant to the particular representation used, such as quaternions, axis-angle, etc. We use Adam optimization with batch size 64 and learning rate 10−510^{-5} for the first 10410^{4} iterations and 10−610^{-6} for the remaining iterations. For sampling the input rotation matrices during training, we uniformly sample axes and angles. We test the networks using 10510^{5} rotation matrices generated by randomly sampling axes and angles and calculate geodesic errors between the input and the output rotation matrices. The geodesic error is defined as the minimal angular difference between two rotations, written as

Figure 5(a) illustrates the mean geodesic errors for different representations as training progresses. Figure 5(b) illustrates the percentiles of the errors at 500k iterations. The results show that the 6D and 5D representations have similar performance with each other. They converge much faster than the other representations and produce smallest mean, maximum and standard deviation of errors. The Euler angle representation performs the worst, as shown in Table (c) in Figure 5. For the quaternion, axis angle and Euler angle representations, the majority of the errors fall under 25∘, but certain test samples still produce errors up to 180∘. The proposed 6D and 5D representations do not produce errors higher than 2∘. We conclude that using continuous rotation representations for network training leads to lower errors and faster convergence.

In Appendix G.2, we report additional results, where we trained using a geodesic loss, uniformly sampled SO(3)SO(3), and compared to 3D Rodriguez vectors and quaternions constrained to one hemisphere . Again, our continuous representations outperform common discontinuous ones.

2 Pose Estimation for 3D Point Clouds

We train the network with 2,290 airplane point clouds from ShapeNet , and test it with 400 held-out point clouds augmented with 100 random rotations. At each training iteration, we randomly select a reference point cloud and transform it with 10 randomly-sampled rotation matrices to get 10 target point clouds. We feed the paired reference-target point clouds into the Siamese network and minimize the L2 loss between the output and the ground-truth rotation matrices.

We trained the network with 2.6×1062.6\times 10^{6} iterations. Plot (d) in Figure 5 shows the mean geodesic errors as training progresses. Plot (e) and Table (f) show the percentile, mean, max and standard deviation of errors. Again, the 6D representation has the lowest mean and standard deviation of errors with around 95% of errors lower than 5∘, while Euler representation is the worst with around 10% of errors higher than 25∘. Unlike the sanity test, the 5D representation here performs worse than the 6D representation, but outperforms the 3D and 4D representations. We hypothesize that the distortion in the gradients caused by the stereographic projection makes it harder for the network to do the regression. Since the ground-truth rotation matrices are available, we can directly regress the 3×33\times 3 matrix using L2 loss. During testing, we use the Gram-Schmidt process to transform the predicted matrix into SO(3)SO(3) and then report the geodesic error (see the bottom row of Table (f) in Figure 5). We hypothesize that the reason for the worse performance of the 3×33\times 3 matrix compared to the 6D representation is due to the orthogonalization post-process, which introduces errors.

3 Inverse Kinematics for Human Poses

In this experiment, we train a neural network to solve human pose inverse kinematics (IK) problems. Similar to the method of Villegas et al. and Hsu et al. , our network takes the joint positions of the current pose as inputs and predicts the rotations from the T-pose to the current pose. We use a fixed forward kinematic function to transform predicted rotations back to joint positions and penalize their L2 distance from the ground truth. Previous work for this task used quaternions. We instead test on different rotation representations and compare their performance.

We train a four-layer MLP network that has 1024 neurons in hidden layers with the L2 reconstruction loss L=∣∣P−P′∣∣22L=||P-P^{\prime}||^{2}_{2}, where P′=Π(T,R)P^{\prime}=\Pi(T,R). Here Π\Pi is the forward kinematics function which takes as inputs the “T” pose of the skeleton and the predicted joints rotations, and outputs the 3D positions of the joints. Due to the recursive computational structure of forward kinematics, the accuracy of the hip orientation is critical for the overall skeleton pose prediction and thus the joints adjacent to the hip contribute more weight to the loss (10 times higher than other joints).

We use the CMU Motion Capture Database for training and testing because it contains complex motions like dancing and martial arts, which cover a wide range of joint rotations. We picked in total 865 motion clips from 37 motion categories. We randomly chose 73 clips for testing and the rest for training. We fix the global position of the hip so that we do not need to worry about predicting the global translation. The whole training set contains 1.14×1061.14\times 10^{6} frames of human poses and the test set contains 1.07×1051.07\times 10^{5} frames of human poses. We train the networks with 1,960k iterations with batch size 64. During training, we augmented the poses with random rotation along the y-axis. We augmented each instance in the test set with three random rotations along the y-axis as well. The results, as displayed in subplots (g), (h) and (i) in Figure 5, show that the 6D representation performs the best with the lowest errors and fastest convergence. The 5D representation has similar performance as the 6D one. On the contrary, the 4D and 3D representations have higher average errors and higher percentages of big errors that exceed 10 cm.

We also perform the test of using the 3×\times3 matrix without orthogonalization during training and using the Gram-Schmidt process to transform the predicted matrix into SO(3) during testing. We find this method creates huge errors as reported in the bottom line of Table (i) in Figure 5. One possible reason for this bad performance is that the 3×\times3 matrix may cause the bone lengths to scale during the forward kinematics process. In Appendix G.1, we additionally visualize some human body poses for quaternions and our 6D representation.

Conclusion

We investigated the use of neural networks to approximate the mappings between various rotation representations. We found empirically that neural networks can better fit continuous representations. For 3D rotations, the commonly used quaternion and Euler angle representations have discontinuities and can cause problems during learning. We present continuous 5D and 6D rotation representations and demonstrate their advantages using an auto-encoder sanity test, as well as real world applications, such as 3D pose estimation and human inverse kinematics.

Acknowledgements

We thank Noam Aigerman, Kee Yuen Lam, and Sitao Xiang for fruitful discussions; Fangjian Guo, Xinchen Yan, and Haoqi Li for helping with the presentation. This research was conducted at USC and Adobe and was funded by in part by the ONR YIP grant N00014-17-S-FO14, the CONIX Research Center, one of six centers in JUMP, a Semiconductor Research Corporation (SRC) program sponsored by DARPA, the Andrew and Erna Viterbi Early Career Chair, the U.S. Army Research Laboratory (ARL) under contract number W911NF-14-D-0005, Adobe, and Sony. This project was not funded by Pinscreen, nor has it been conducted at Pinscreen or by anyone else affiliated with Pinscreen. The content of the information does not necessarily reflect the position or the policy of the Government, and no official endorsement should be inferred.

References

Appendix A Overview of the Supplemental Document

In this supplemental material, in Section B, we first give explicitly our 6D representation for the 3D rotations. We then prove formally in Section C that the formula of the 5D representation as defined in Case 4 of Section 4.2 satisfies all of the properties of a continuous representation. Next, we discuss quaternions in more depth. We present in Section D a result that we elided from the main paper due to space limitations: that the unit quaternions are also a discontinuous representation for the 3D rotations. Then, in Section E, we show how the continuous 5D and 6D representations can interact with common discontinuous angle representations such as quaternions. Next, we visualize in Section F discontinuities that are present in some of the representations. We finally present in Section G some additional empirical results.

Appendix B 6D Representation for the 3D Rotations

The mapping from SO(3) to our 6D representation is:

The mapping from our 6D representation to SO(3) is:

Appendix C Proof that Case 4 gives a Continuous Representation

Here we show that the functions fP,gPf_{P},g_{P} presented in Case 4 of Section 4.2 are a continuous representation. We now prove some properties needed to show a continuous representation: that gPg_{P} is defined on its domain and continuous, and that fP(gP(M))=Mf_{P}(g_{P}(M))=M for all M∈SO(n)M\in SO(n). In these proofs, we use to denote the zero vector in the appropriate Euclidean space. We also use the same slicing notation from the main paper. That is, if uu is a vector of length mm, define the slicing notation ui:j=(ui,ui+1,…,uj)u_{i:j}=(u_{i},u_{i+1},\ldots,u_{j}), and ui:=ui:mu_{i:}=u_{i:m}.

Proof that gPg_{P} is defined on SO(n)SO(n). Suppose M∈SO(n)M\in SO(n). The same as we did for Equation (10) in the main paper, define a vectorized representation γ(M)\gamma(M) by dropping the last column of MM: γ(M)=[M(1)T,…,M(n−1)T]\gamma(M)=[M_{(1)}^{T},\ldots,M_{(n-1)}^{T}], where M(i)M_{(i)} indicates the iith column of MM. Following Equation (8), which defines the normalized projection, and Equation (10), let v=γn2−2n:/∣∣γn2−2n:∣∣v=\gamma_{n^{2}-2n:}/||\gamma_{n^{2}-2n:}||. The only way that gP(M)g_{P}(M) could not be defined is if the normalized projection P(v)P(v) is not defined, which requires v1=1v_{1}=1. However, if v1=1v_{1}=1, then because vv is unit length, it follows that γn2−2n+1:\gamma_{n^{2}-2n+1:} has length zero. But γn2−2n+1:\gamma_{n^{2}-2n+1:} is a column vector from M∈SO(n),M\in SO(n), and therefore has unit length. We conclude that v1≠1v_{1}\neq 1 and gPg_{P} is defined on SO(n)SO(n).

Proof that gPg_{P} is continuous. This case is trivial because gPg_{P} is the composition of functions that are continuous on their domains and thus is also continuous on its domain.

Now let b=Q(P(u)).b=Q(P(u)). Components 2 through mm of bb are P(u)/∣∣P(u)∣∣,P(u)/||P(u)||, but this is just u2:u_{2:}. Next, consider b1b_{1}, the first component of bb:

Appendix D The Unit Quaternions are a Discontinuous Representation for the 3D Rotations

In Case 2 of Section 4, we showed that the quaternions are not a continuous representation for the 3D rotations. We intentionally used a simpler formulation for the quaternions, which is easier to understand and also saves space in the paper, due to its quaternions in general not being unit length. However, an attentive reader might wonder what happens if we use the unit quaternions: is the discontinuity removable? However, we show here that the unit quaternions are also not a continuous representation for the 3D rotations.

By substitution, we can find the components of gu(B(θ))g_{u}(B(\theta)) as a function of θ\theta. For example, as θ→π−\theta\to\pi-, the third component is (1−cos⁡(θ))/2=1\sqrt{(1-\cos(\theta))/2}=1. Meanwhile, as θ→π+\theta\to\pi+, the third component is −(1−cos⁡(θ))/2=−1-\sqrt{(1-\cos(\theta))/2}=-1. We conclude that the unit quaternions are not a continuous representation.

A similar representation is the Cayley transformation which has a different scaling to the unit quaternion, where w=1w=1 and the vector (x,y,z)(x,y,z) is the unit axis of rotation scaled by tan(θ/2)tan(\theta/2). The limit goes to infinity when approaching 180∘180^{\circ}. Thus, it is not a representation for SO(3)SO(3).

Appendix E Interaction Between 5D and 6D Continuous Representations and Discontinuous Ones

In some cases, it may be convenient to use a common 3D or 4D angle representation, such as the quaternions or Euler angles. For example, the quaternions may be useful when interpolating between two rotations in SO(3)SO(3), or when there is an existing neural network that already accepts quaternion inputs. However, as we showed in the main paper Case 2 of Section 4.1, all 3D and 4D representations for rotations are discontinuous.

One solution for the above conundrum is to simply convert as needed from the continuous 5D and 6D representations that we presented in Case 3 and 4 of Section 4.2 to the desired representation. For concreteness, suppose the desired representation is the quaternions. Assume that any conversions done in the network are only in the direction that maps to the quaternions. Then the associated mapping in the opposite direction (i.e. from quaternions to the 5D or 6D representation) is continuous. If losses are applied only at points in the network where the representation is continuous (e.g. on the 5D or 6D representations), then the learning should not suffer from discontinuity problems. One can convert from the 5D or 6D representation to quaternions by first applying Equation (5) or Equation (10) and then using Equation (4). Of course, one could also make a similar argument for other discontinuous but popular angle representations such as Euler angles.

Appendix F Visualizing Discontinuities in 3D Rotations

Here we visualize any discontinuities that might occur in the 3D rotation representations. We do this by forming three continuous curves in SO(3)SO(3), which we call the “X, Y, and Z Rotations.” We map each of these curves to each representation, and then map the representation curve to 2D by retaining the top two components from Principal Components Analysis (PCA). We call the first curve in SO(3)SO(3) the “X Rotations:” this curve is formed by taking the X axis (1,0,0)(1,0,0), and constructing a curve consisting of all rotations around this axis as parameterized by angle. Likewise, we call the second and third curves in SO(3)SO(3) the “Y Rotations” and “Z Rotations:” these curves are formed by rotating around the Y and Z axes, respectively. We show the resulting 2D curves in Figure 6.

Appendix G Additional Empirical Results

In this section, we show some additional empirical results.

In Figure 7, we visualize the worst two frames with the highest pose errors generated by the network trained on quaternions, along with the corresponding results from the network trained with 6D representations. Likewise, we show the two frames with highest pose errors generated by the network trained on our 6D representation, along with the corresponding results from the network trained on quaternions. This shows that for the worst error frames, the quaternion representation introduces bad qualitative results while the 6D one still creates a pose that is reasonable.

G.2 Additional Sanity test

In the main paper, Section 5.1, we show the sanity test result of the network trained with L2 loss between the ground-truth and the output rotation matrices. Another option for training is using the geodesic loss. Besides, the networks in the main paper are trained and tested using a uniform sampling of the axis and the angle which is not a uniform sampling on SO(3) . We present the sanity test result of using the geodesic loss and the two sampling methods in Figure 8. They are both similar to the result in the main paper.

Additional representations. In addition to common rotation representations like Euler angles, axis-angles and quaternions, we investigated a few other rotation representations used in recent work including a 3D Rodriguez vector representation, and quaternions that are constrained to one hemisphere as given by Kendall et al. . The 3D Rodriguez vector is given as R=ωθR=\omega\theta, where ω\omega is a 3D unit vector and θ\theta is the angle . We will not provide the proofs for the discontinuity in these representations, but we show their empirical results in Figure 8. We find that the errors are significantly worse than our 5D and 6D representations.