OrcVIO: Object residual constrained Visual-Inertial Odometry
Mo Shan, Vikas Dhiman, Qiaojun Feng, Jinzhao Li, Nikolay Atanasov
Supplementary Material
Software and videos supplementing this paper are available at:
I Introduction
The foundations of visual environment understanding in robotics, machine learning, and computer vision lie in the twin technologies of inferring geometric structure and semantic content. Researchers have made significant progress in geometric structure reconstruction using Structure from Motion (SfM) and SLAM techniques. State of the art SLAM approaches work with monocular or stereo cameras , often complemented by inertial information . However, most real-time incremental SLAM techniques provide purely geometric representations, e.g., of points, lines, or planes, that lack semantic interpretation of the environment.
Recently, tremendous progress has been achieved in semantic scene understanding using deep neural networks for object detection , instance segmentation , and object tracking . Nevertheless, the literature in deep learning is sparse in techniques that provide global positioning of the detected and tracked objects to obtain an object-level map online.
This paper focuses on joint visual-inertial odometry and object mapping, bridging the gap between geometric and semantic inference in SLAM. Generating geometrically consistent and semantically meaningful maps allows compressed representation, improved loop closure (recognizing already visited locations), and robot mission specifications in terms of human-interpretable objects, e.g., for safe navigation, manipulation, multi-stage interaction . Our work introduces object states, modeling the position, orientation, and shape of object instances in the environment. We utilize a coarse category-agnostic and a fine category-specific model of object shape. The coarse model uses an ellipsoid to restrict an object’s pose variation and relate its coarse shape to bounding-box detections. The fine model uses a set of mid-level object parts (e.g., car wheels, windshield, doors), called semantic keypoints, to obtain a precise part-based shape description. Given streaming inertial measurement unit (IMU) and monocular camera measurements, we develop an algorithm to simultaneously estimate the IMU-camera trajectory and the states of the objects, detected and tracked in the camera images. The contributions of this paper are summarized as follows.
We introduce object states in the formulation of a SLAM problem, modeling position, orientation, coarse ellipsoid shape, and fine semantic-keypoint shape.
We define residuals relating object states and IMU-camera states to inertial measurements, geometric features, object semantic features, and object bounding-box detections, and explicitly derive their Jacobians.
We develop an extension of the multi-state constraint Kalman filter (MSCKF) to enable online tightly coupled estimation of object and IMU-camera states. Our innovations include closed-form mean and covariance propagation over the SE(3) pose and velocity manifold of the IMU-camera states and measurement updates based on our new semantic residuals with object states optimized over multiple views. Our algorithm is suitable for real-time incremental odometry and object mapping and is more efficient than nonlinear batch optimization.
We name our method Object residual constrained Visual Inertial Odometry (OrcVIO) to emphasize the role of the semantic residuals in the optimization process. OrcVIO is capable of producing meaningful object maps and estimating accurate sensor trajectories, as shown in Fig. 1.
II Related Work
Many visual SLAM approaches work with monocular or stereo cameras . Featureless approaches that minimize image intensity directly have been proposed, and inertial information is often used to complement the visual information . The MSCKF algorithm leverages both inertial data and visual features in an extened Kalman filter formulation. Each visual feature whose track is lost provides multi-frame constraints for a corresponding 3D landmark. The constraint residuals are linearized to perform the filter update step. A key idea is to marginalize the landmark states via projection to the null space of the visual feature Jacobians, allowing structureless estimation of the IMU-camera states only. A key extension of OrcVIO over the MSCKF and visual-inertial SLAM algorithms in general is to introduce object measurements (bounding boxes and semantic features) and object states whose residuals constrain the IMU-camera states.
Recent methods have considered learning to regress camera poses directly from raw images . For instance, monocular depth, optical flow, and ego-motion are jointly optimized from video in by relying on a view-synthesis loss. These unsupervised learning techniques have shown impressive performance in localization but do not generate global maps. In this paper, we focus on obtaining a joint geometric-semantic representations from measurements in real time, i.e., spatial perception . Prior works that utilize both spatial and semantic information include , but the spatial and semantic states are estimated independently and merged later. On the other hand, consider joint metric and semantic mapping. Recent works focus on the tightly coupled spatial and semantic estimation, and there are mainly two groups of object-based SLAM techniques: category-specific and category-agnostic.
Category-specific approaches optimize the pose and shape of object instances, using semantic keypoints or 3D shape models . For example, introduces a real-time joint 3D object pose and camera pose estimation via pose graph optimization. The objects stored in a database that are also present in the current frame are detected and optimized, using the vertex and normal map from a RGBD sensor. Object pose and shape are optimized in using semantic keypoints to provide additional error terms related to object pose in the SLAM factor graph. Visual-inertial information for object mapping is used in , relying on a database to retrieve object shapes. Hu et al. embed object priors into least-squares minimization to incrementally track and map chairs. The object shape represented by a binary voxel grid is compactly described by a latent code obtained from an auto-encoder, which is used for shape initialization and iterative residual minimization. These methods in general are computationally demanding due to iterative batch optimization and reliance on instance-specific CAD models.
Category-agnostic approaches use geometric shapes, such as spheres , cuboids , or ellipsoids , to represent objects. SSFM uses the tightest bounding cube enclosing an object to parameterize the object location and pose. The object measurements include the location, size of the object bounding box and the object pose obtained from a 3D object detector . Assuming the semantic measurements are consistent across frames, an object state is optimized via maximum likelihood estimation. CubeSLAM generates and refines 3D cuboid proposals using multi-view bundle adjustment without relying on prior models. QuadricSLAM uses an ellipsoid representation, suitable for defining a bounding-box detection model. Structural constraints based on supporting and tangent planes, commonly observed under a Manhattan assumption, have also been introduced . Using generic symmetric shapes, however, makes the orientation of object instances potentially irrecoverable. For instance, the front and back of an object become indistinguishable.
This paper extends our prior conference publication with several theoretical contributions and a large-scale experimental evaluation of OrcVIO. While used linear velocity measurements, this paper uses a complete description of all IMU states for odometry and derives a novel closed-form expression for covariance propagation. A new bounding-box residual is developed to ensure that it scales equivalently as the keypoint residual. In , the bounding-box residual was quadratic instead of linear as a function of the bounding-box lines. This version also implements a zero-velocity residual to reduce VIO drift when the sensor is completely static. Extensive evaluation is conducted on the KITTI dataset, a photo-realistic Unity dataset, and in indoor and outdoor real-time experiments. We also extend OrcVIO to handle multiple object classes, including cars, doors, barriers, monitors, chairs.
III Background and Notation
We denote the IMU, camera, object, and global reference frames as , , , , respectively. The transformation from frame to is specified by a matrix:
We define an infinitesimal change of pose using a right perturbation (see [73, Ch.7]).
where is the identity matrix. A quadric shape [74, Ch.3] is a set , where is a symmetric matrix. Consider an axis-aligned ellipsoid centered at :
IV Problem Formulation
The shape of an object in the global frame is obtained by deforming and transforming the semantic landmark positions, , and the dual ellipsoid, , using the instance pose and deformations , . Fig. 2 shows an illustration for a car model with semantic landmarks.
Let indicate whether the -th geometric keypoint observed at time is associated with the -th geometric landmark. Similarly, let indicate whether the -th object detection at time is associated with the -th object instance. This data association information can be obtained by keypoint and object tracking as described in Sec. V-A. Given the associations, we introduce error functions:
for the inertial, geometric, semantic keypoint and bounding-box measurements, respectively, defined precisely in Sec. V. We also introduce an object shape regularization error term to ensure that the instance deformations remain small, and consider the following problem.
Determine the sensor trajectory , geometric landmarks , and object states that minimize the weighted sum of squared errors:
where are positive constants determining the relative importance of the error terms and are positive-definite matrices specifying the covariances associated with the inertial, geometric, semantic, and bounding-box measurements. A measurement covariance defines a quadratic (Mahalanobis) norm .
V Landmark and Object Reconstruction
Our approach consists for a front-end measurement generation stage and a back-end landmark state and sensor pose optimization stage. This section discusses the detection and tracking of geometric keypoint measurements and object class , bounding-box , and semantic keypoint measurements in the front-end. It also defines the error functions in (Problem) and their Jacobians needed for the back-end optimization. Finally, it presents the back-end optimization over the geometric landmarks and the object states for a given sensor sensor trajectory .
Geometric keypoints are detected in the camera images using the FAST detector and are tracked temporally using the Lucas-Kanade (LK) algorithm . Keypoint-based tracking has lower accuracy but higher efficiency than descriptor-based methods, allowing our method to use a high frame-rate camera and process more keypoints. Outliers are eliminated by estimating the essential matrix between consecutive views and removing those keypoints that do not fit the estimated model. Assuming that the time between consecutive images is short, the relative orientation is obtained by integrating the gyroscope measurements as described in Sec. VI-A and only the unit translation vector is estimated using two-point RANSAC .
The YOLO detector is used to detect object classes and bounding-box lines . Semantic keypoints are extracted within each bounding box using the StarMap stacked hourglass convolutional neural network . We augment the original StarMap network with dropout layers as shown in Fig. 3(b). Several stochastic forward passes may be preformed using Monte Carlo dropout to obtain semantic keypoint covariances , illustrated in Fig. 3(c).
The bounding boxes are tracked temporally using the SORT algorithm , which performs intersection over union (IoU) matching via the Hungarian algorithm. The semantic keypoints within each bounding box are tracked via a Kalman filter, which uses the LK algorithm for prediction and the StarMap keypoint detections for update.
V-B Landmark and Object Error Functions
where emphasizes that some additions are over the manifold, defined as follows:
where we use right perturbations , , and for the IMU orientation , camera pose , and object pose , respectively.
where is defined in (3), is the Jacobian of and:
The Jacobians with respect to other perturbations in (7) are .
The semantic-keypoint error is defined as the difference between the projection of a semantic landmark from the object frame to the image plane, using instance pose and camera pose , and its corresponding semantic-keypoint observation :
The Jacobians with respect to other perturbations in (7) are .
We define the bounding-box error as the distance between the hyperplane induced by projecting the bounding-box line to the object frame and the closest hyperplane that is tangent to the quadric surface of object :
where is the signum function.
where and with :
The Jacobians with respect to other perturbations in (7) are .
Finally, the object shape regularization error is defined as:
V-C Landmark and Object State Optimization
for all and all such that , where the unknowns are and the semantic keypoint depths .
The least squares problem for semantic keypoints is a generalization of the pose from point correspondences (PnP) problem . While this system may be solved using polynomial equations , we perform a more efficient initialization by defining and solving the first set of (now linear) equations in (19) for and . We recover via the Kabsch algorithm between and . This approach works well as long as there is a sufficient number of semantic keypoints (at least two per landmark across time for at least three semantic landmarks ) associated with the object. If fewer semantic keypoints are available, can be recovered from the second equation in (19). Let , , then a linear system can be formed:
for all and all such that . The operator vectorizes a matrix by stacking its columns into a single column vector and is the Kronecker product. The solution can be found by:
where the equality constraint avoids the trivial solution. The minimization in (21) can done via SVD of the matrix . The object pose can be recovered from by relating the estimated ellipsoid in global coordinates to the ellipsoid in the object frame:
The translation can be recovered from the last column of . To recover the rotation, note that is a positive semidefinite matrix. Let its eigen-decomposition be , where is a diagonal matrix containing the eigenvalues of . Since is diagonal, it follows that .
VI The OrcVIO Algorithm
Given time discretization and assuming and remain constant over the interval with values , we can compute the predicted IMU-camera state mean and covariance from (VI-A) and (24), respectively. Let and be the prior mean and covariance.
The nominal dynamics (VI-A) can be integrated in closed-form to obtain the predicted mean :
where is the left Jacobian of and . Both and admit closed-form (Rodrigues) expressions:
where is the top-left block of corresponding to the IMU state, is white Gaussian noise with power spectral density and is:
where is the transition matrixNote that in is time varying. Time-invariant approximations of the transition matrix are presented in [86, App. B, E] using right-perturbation error dynamics and in using left-perturbation error dynamics. of (27).
The LTV SDE in (27) has a closed-form transition matrix:
where , and the blocks are:
where .
Using (28) and Proposition 5, we can approximate the discrete-time covariance propagation for the IMU state as . The full covariance propagation, accounting for the addition of the latest IMU pose to the sliding window and the dropping of the oldest IMU pose , is:
VI-B Update Step
where is the associated noise term with covariance . Stacking the errors for all camera poses in associated with , leads to:
In Sec. IX-F, we also introduce a zero-velocity residual to account for complete stops of the system, which is common for autonomous ground and some aerial robots.
Finally, let , , be the stacked errors, Jacobians, and noise covariances (after null-space projection) across all geometric landmarks and object instances, whose tracks are lost at . The updated IMU-camera mean and covariance are:
Note that the dimension of can be reduced in the computation above via QR factorization as described in .
VII Evaluation
This section presents results from large-scale evaluation of OrcVIO using indoor and outdoor simulated and real data. We present experiments in photo-realistic Unity simulation (Fig. 5), on raw and odometry data sequences from the KITTI dataset , and on real data collected on UCSD’s campus.
We use two standard metrics for quantitative VIO evaluation: position Root Mean Square Error (RMSE) (also referred to as position ATE ) and KITTI’s translation error (TE) metric . Let return the position component of a pose . Let be the ground-truth pose trajectory and be the estimated pose trajectory. To measure error, the estimated trajectory is first aligned to the initial frame of the ground-truth trajectory via . After alignment, RMSE (m) and TE (%) are measured as:
where is a set of frames with fixed distances over the values .
The object estimates are evaluated using 3D Intersection over Union (IoU). A 3D bounding box is obtained from each estimated object , and IoU is defined as the ratio of the intersection volume over the union volume with respect to the bounding box of the closest ground truth object:
To understand the distribution of the object orientation and translation errors, we define an estimate as true positive if the closest ground-truth object pose is within a specific rotation or translation error threshold. Specifically, a rotation error of means , and translation error of means . We define precision as the fraction of true positives over all estimated objects and recall is the fraction of true positives over all ground-truth objects.
VII-B Unity Dataset Results
OrcVIO is evaluated in a Unity simulation containing 50 car, road barrier, and door object instances, shown in Fig. 5. A ROS bridge between Unity and Gazebo is used to simulate a quadrotor robot, navigating in the environment and providing IMU and camera measurements. The object map reconstructed by OrcVIO is shown in Fig. 5. Additional qualitative results and a video demo are available in the Supplementary Material. The estimated objects are generally quite close to the ground-truth ones. The object poses near the starting position are less accurate due to insufficient motion parallax, since the quadrotor performs a pure rotation in the beginning. The trajectory RMSE and TE are and , respectively. The odometry drift is mainly due to pure rotation maneuvers at planned path corners executed by the quadrotor controller.
The 3D IoU of the object estimates is . The Precision and Recall of the object reconstruction is shown in Table I. Despite that doors and barriers have a thin structure, causing even small pose estimation drift to reduce the overlap with the ground-truth object instances, OrcVIO is able to produce an accurate object map with good 3D IoU.
VII-C KITTI Dataset Results
We also evaluate OrcVIO on the KITTI dataset , both qualitatively and quantitatively. We use the raw data sequences with object annotations to evaluate the object state estimation and the odometry sequences without object annotations for trajectory accuracy evaluation. Since the original inertial data from KITTI is not reliable, we use a VIO simulator to generate realistic noisy high-frequency IMU measurements at 250 Hz from the ground-truth poses.
Fig. 6 shows the IMU-camera trajectory and object states estimated on KITTI odometry sequence 07, and a video demo is provided in the Supplementary Material. The results demonstrate that OrcVIO produces meaningful object maps. We obtained ground-truth 3D annotations from the KITTI tracklets and the KITTI detection benchmark labels for quantiative evaluation of the object reconstruction. Table II reports 3D IoU results comparing OrcVIO against state-of-the-art methods, including a deep learning approach for single-view 3-D bounding-box estimation (SingleView ), and a multi-view bundle-adjustment algorithm that uses cuboids to represent objects (CubeSLAM ). OrcVIO has better 3D IoU than SingleView for the majority of the sequences (23, 36, 39, 64, 96) because, unlike SingleView, OrcVIO use multi-view optimization over the object states. The performances of OrcVIO and CubeSLAM are similar since both rely on point and bounding-box measurements to optimize the object states. CubeSLAM is better than OrcVIO in terms of 3D IoU, possibly because OrcVIO does not model dynamic objects. Table II also shows that OrcVIO is slightly better than CubeSLAM according to the TE odometry metric. In contrast with CubeSLAM, OrcVIO uses inertial residuals in addition to the geometric and object residuals.
In Table III, we compare the Precision and Recall of OrcVIO on the KITTI raw sequences (2011_09_26_00XX, XX = ) against a single-view, end-to-end object estimation approach (SubCNN ), and a visual-inertial object detector (VIS-FNL ). The first six rows are the Precision and Recall associated with different translation error (row) and rotation error (column) thresholds, whereas the last 3 rows ignore the rotation error. The results demonstrate that OrcVIO retrieves a reasonable amount of the ground-truth objects and provides a high-quality object map. When both rotation and translation errors are considered (first six rows), OrcVIO is better than SubCNN, since the latter does not rely on temporal object tracking. OrcVIO is comparable with VIS-FNL, even though VIS-FNL uses multiple object hypotheses while OrcVIO only keeps one object state. OrcVIO outperforms SubCNN and VIS-FNL when only translation error is taken into account (last three rows), which suggests that the object position estimates are accurate but the orientation estimates could be improved.
We evaluate the RMSE of the IMU-camera trajectory estimation in Table IV. OrcVIO is compared with two visual object SLAM methods: a monocular visual SLAM that integrates spherical object models to estimate the scale via bundle-adjustment (Object BA ), and CubeSLAM . Table IV shows that OrcVIO outperforms Object BA because spheres are a very coarse shape representation, compared to our ellipsoid and semantic keypoint representation, and thus Object BA cannot maintain the object scales as accurately as OrcVIO. The use of inertial data in OrcVIO leads to superior results as well. OrcVIO also outperforms CubeSLAM, possibly due to the incorporation of a zero-velocity update (Sec. IX-F), which is critical in driving datasets with frequent stops.
VII-D UCSD Campus Dataset Results
We also evaluated OrcVIO using real data collected with two commercial VIO sensors on UCSD’s campus. The results are qualitative due to the lack of ground truth information.
First, we used Intel RealSense D435i with image frequency of 30 Hz, image resolution of , IMU frequency of 200 Hz in an indoor lab scene with chairs and monitors as shown in Fig. 7. The estimated sensor trajectory and reconstructed object map by OrcVIO are shown in the figure. A video demo can be found in the Supplementary Material. The results demonstrate that OrcVIO can map object instances from different categories and operates at real-time speed in a cluttered indoor scene. Since OrcVIO does not currently have a loop-closing mechanism for object re-identification, objects reappearing after getting lost will be mapped twice. Thus, there are more reconstructed chairs in Fig. 7 than in the reality.
Semantic keypoint detection is challenging due to occlusion, viewpoint change, and the lack of distinctive features on the monitors as shown in Fig. 7 (b). To handle the monitor class succesfully, we decreased the weight of the semantic keypoint residual in the object Levenberg-Marquardt optimization (17). Although removing the reliance on semantic keypoints leads to worse object orientation estimation, it allows OrcVIO to work with bounding-box detections only. This simplifies the front-end to an object detector and tracker and makes the algorithm more efficient and easier to deploy on resource-constrained robots. We release code for both the full OrcVIO algorithm and the OrcVIO-Lite version in the Supplementary Material.
Finally, we used an INDEMIND Binocular Visual-Inertial Camera to run OrcVIO outdoors with images at 25 Hz with resolution of and IMU measurements at 200 Hz. The sensor initially observes bikes and chairs, and then makes a transition into a parking structure, as shown in Fig. 8 (b), (c). A video demonstrating the performance of OrcVIO on this dataset can be found in the Supplementary Material. The resulting object map is shown in Fig. 8 (a), demonstrating that OrcVIO is able to estimate object states from different categories in both indoor and outdoor scenes. This experiment also shows that OrcVIO can handle large illumination changes, transitioning from direct outdoor sunlight to dim lighting inside the parking structure.
VIII Conclusion
This paper presented a unified formulation of ego-motion and object pose and shape estimation and developed a real-time simultaneous localization and object mapping algorithm. OrcVIO provides computationally constrained robot platforms with the ability to understand their surroundings at geometric and semantic levels, which may enable further capabilities such as object-level loop closure in challenging environments, collaborative semantic SLAM, and object interaction. Estimating object motion and performing collaborative mapping and loop closure are promising directions for future research.
IX Appendix
Similarly, approximating using and leads to:
Equating the two expression above, we get:
Next, we derive the expressions in (9). Note that:
where by discarding high-order terms in , and we used (7.159) in [73, Ch.7] in the last equality. Applying the chain rule to (8):
IX-B Proof of Proposition 2
IX-C Proof of Proposition 3
Given a hyperplane and an origin-centered ellispoid as a dual quadric , there are two hyperplanes parallel to that are tangent to the ellipsoid: and the signed distance between and the closest tangent hyperplane is:
Given a hyperplane and an origin-centered ellispoid as a dual quadric , the Jacobian of the distance in (44) with respect to is:
We rewrite explicitly in terms of by replacing :
Taking the derivative using the product rule leads to:
The derivative of is:
The derivative of is:
Putting the derivatives in (45) leads to:
We proceed with the proof of Proposition 3. Note that in (44) with . The expression for , evaluated at and , follows directly from Lemma 2. The expressions for and are obtained using the perturbations and , (7.159) in [73, Ch.7], and dropping second-order perturbation terms:
IX-D Proof of Proposition 4
Using that for :
The derivation for is equivalent. ∎
The result follows by integrating the terms of the Taylor series of and since and are well defined by Lemma 3. ∎
To obtain (4), we compute the solutions to the linear time-invariant (LTI) ordinary differential equations (ODEs) in (VI-A). With and , the solution to:
is and, hence:
Similarly, with , , and initial condition , the solution of the LTI ODE for in (VI-A) is:
where the second equality uses Lemma 4. Hence, satisfies the second expression in (4). Also, by Lemma 4, for with initial condition , the solution of the LTI ODE for in (VI-A) is:
Hence, satisfies the third expression in (4). Finally, and for all . The IMU pose history in (4) is updated by adding the pose to the sliding window and dropping the oldest pose .
IX-E Proof of Proposition 5
To integrate (55), observe that from :
By integrating the terms of the Taylor series of and using , we can verify that:
where is defined in (29). The second integral in (59) can be computed using , , , and Lemma 4, leading to:
where (63) and (65) follow from Lemma 4. To integrate (64), we use that:
IX-F Zero-Velocity Update
Zero-velocity conditions are frequently encountered in autonomous driving and autonomous flight applications, and thus the ability to determine whether the robot is static is important for reducing drift in VIO estimation . In OrcVIO zero-velocity is detected similarly as in by using pseudo zero inertial measurements and then checking the velocity magnitude. For the zero-velocity update, we need to compute the residuals and the Jacobians of the inertial measurements with respect to the state. Based on (VI-A) the residuals are:
The corresponding Jacobians are presented as follows: