Simple multi-dataset detection

Xingyi Zhou, Vladlen Koltun, Philipp Krähenbühl

Introduction

Computer vision aims to produce broad, general-purpose perception systems that work in the wild. Yet object detection is fragmented into datasets and our models are locked into the corresponding domains. This fragmentation brought rapid progress in object detection and instance segmentation , but comes with a drawback. Single datasets are limited in both image domains and label vocabularies and do not yield general-purpose recognition systems. Can we alleviate these limitations by unifying diverse detection datasets?

In this paper, we first make training an object detector on a collection of disparate datasets as straightforward as training on a single one. Different datasets are usually trained under different training losses, data sampling strategies, and schedules. We show that we can train a single detector with separate outputs for each dataset, and apply dataset-specific supervision to each. Our training mimics training parallel dataset-specific models with a common network. As a result, our single detector takes full advantages of all training data, performs well on training domains, and generalizes better to new unseen domains. However, this detector produces duplicate outputs for classes that occur in multiple datasets.

A core challenge is integrating different datasets into a common taxonomy, and training a detector that reasons about general objects instead of dataset-specific classes. Traditional approaches create this taxonomy by hand , which is both time-consuming and error-prone. We present a fully automatic way to unify the output space of a multi-dataset detection system using visual data only. We use the fact that object detectors for similar concepts from different datasets fire on similar novel objects. This allows us to define the cost of merging concepts across datasets, and optimize for a common taxonomy fully automatically. Our optimization jointly finds a unified taxonomy, a mapping from this taxonomy to each dataset, and a detector over the unified taxonomy using a novel 0-1 integer programming formulation. An object detector trained on this unified taxonomy has a large, automatically constructed vocabulary of concepts from all training datasets.

We evaluate our unified object detector at an unprecedented scale. We train a unified detector on 3 large and diverse datasets: COCO , Objects365 , and OpenImages . For the first time, we show that a single detector performs as well as dataset-specific models on each individual dataset. A unified taxonomy further improves this detector. Crucially, we show that models trained on diverse training sets generalize to new domains without retraining, and outperform single-dataset models.

Related Work

Training on multiple datasets. In recent years, training on multiple diverse datasets has emerged as an effective tool to improve model robustness for depth estimation , stereo matching , and person detection . In these domains, unifying the output space involves modeling different camera transformations or depth ambiguities. In contrast, for recognition, dataset unification involves merging different semantic concepts. MSeg manually unified the taxonomies of 7 semantic segmentation datasets and used Amazon Mechanical Turk to resolve inconsistent annotations between datasets. In contrast, we propose to learn a label space from visual data automatically, without requiring any manual effort.

Wang et al. train a universal object detector on multiple datasets, and gain robustness by joining diverse sources of supervision. This is similar to our partitioned detector, while they work on small datasets and didn’t model the training differences between different datasets. Universal-RCNN trains an partitioned detector on three large datasets and models the class relations with a inter-dataset attention module. However again they use the same training recipe for all datasets, and produce duplicated outputs for the same object if it occurs in more one dataset. Both Wang et al. and MSeg observe a performance drop in a single unified model. With our dedicated training framework, this is not the case: our unified model performs as well as single-dataset models on the training datasets. Also, these multi-headed models produce a dataset-specific prediction for each input image. When evaluated in-domain, they require knowledge of the test domain. When evaluated out-of-domain, they produce multiple outputs for a single concept. This limits their generality and usability. Our approach, on the other hand, unifies visual concepts in a single label space and yields a single consistent model that does not require knowledge of the test domain and can be deployed cleanly in new domains.

Zhao et al. trains a universal detector on multiple datasets: COCO , Pascal VOC , and SUN-RGBD , with under 100 classes in total. They manually merge the taxonomies and then train with cross-dataset pseudo-labels generated by dataset-specific models. The pseudo-label idea is complementary to our work. Our unified label space learning removes the manual labor, and works on a much larger scale: we unify COCO, Objects365, and OpenImages, with more complex label spaces and 900+900+ classes. YOLO9000 combines detection and classification datasets to expand the detection vocabulary. LVIS extents COCO annotations to > ⁣1000>\!1000 classes in a federated way. Our approach of fusing multiple annotated datasets is complementary and can be operationalized with no manual effort to unify disparate object detection datasets.

Zero-shot classification and detection reasons about novel object categories outside the training set . This is often realized by representing a novel class by a semantic embedding or auxiliary attribute annotations . In zero-shot detection, Bansal et al. proposed a statically assigned background model to avoid novel classes being detected as background. Rahman et al. used test-time training to progressively generate new class labels based on word embeddings. Li et al. leveraged external text descriptions for novel objects. Our program is complementary: we aim to build a sufficiently large label space by merging diverse datasets during training, such that the trained detector transfers well across domains even without machinery such as word embeddings or attributes. Such machinery can be added, if desired, to further expand our model’s vocabulary.

Preliminaries

Let’s now consider training a detector on multiple datasets D1,D2,…\mathcal{D}_{1},\mathcal{D}_{2},\ldots, each with their own label space L1,L2,…L_{1},L_{2},\ldots. A natural way to train on multiple datasets is to simply combine all annotations of all datasets into a much larger dataset D=D1∪D2∪…\mathcal{D}=\mathcal{D}_{1}\cup\mathcal{D}_{2}\cup\ldots, and merge their label spaces L=L1∪L2∪…L=L_{1}\cup L_{2}\cup\ldots. Labels that repeat across datasets are merged. We then optimize the same loss with more data:

This has shown promise on smaller, evenly distributed datasets . It has the advantage that shared classes between the datasets train on a larger set of annotations. However, modern large-scale detection datasets feature more natural class distributions that are imbalanced. Objects365 contains 5×5\times more images than COCO and OpenImages is 18×18\times larger than COCO. While the top 20%20\% of classes in Objects365 and OpenImages contain 19×19\times and 20×20\times more images than COCO, respectively, the bottom 20%20\% classes actually have fewer images than COCO. This imbalance in class distributions and dataset sizes all but guarantees that a simple concatenation of datasets will not work. In fact, not even the same loss (1) works for all datasets. Most successful Objects365 models employ class-aware sampling . OpenImages models treat rare classes differently and model the hierarchy of classes in the loss .

No single loss generalizes to all datasets. In the next section, we present a different view of multi-dataset training and show how to train a model that performs well on all datasets.

Training a multi-dataset detector

Here, evenly sampling datasets, i.e. showing the partitioned detector the same number of images from each dataset, works best empirically, as we will show in Section 5.

While the partitioned detector learns to detect all classes, it still produces different dataset-specific outputs. For example, it predicts a COCO-person separately from an Objects365-Person, etc. Next we show how to convert this partitioned model into a joint detector that reasons about a unified set of output labels L=L1∪L2∪…L=L_{1}\cup L_{2}\cup\ldots.

Consider multiple datasets, each with its own label space L1,L2,…L_{1},L_{2},\ldots. Our goal is to jointly learn a common label space LL for all datasets, and define a mapping between this common label space and dataset-specific labels Tk:L→Lk{\mathcal{T}_{k}:L\to L_{k}}. Mathematically, Tk∈{0,1}∣Lk∣×∣L∣\mathcal{T}_{k}\in\{0,1\}^{|L_{k}|\times|L|} is a Boolean linear transformation. In this work, we only consider direct mappings. Each joint label c∈Lc\in L maps to at most one dataset-specific label c^∈Lk\hat{c}\in L_{k}: Tk⊤1≤1\mathcal{T}_{k}^{\top}\boldsymbol{1}\leq\boldsymbol{1}. I.e., no dataset contains duplicated classes itself. Also, each dataset-specific label matches to exactly one joint label: Tk1=1\mathcal{T}_{k}\boldsymbol{1}=\boldsymbol{1}. In particular, we do not hierarchically relate concepts across datasets. When there are different label granularities, we keep them all in our label-space, and expect to predict all of themThis follows the official evaluation protocol of OpenImages ..

Simple baselines include hand-designed mappings T\mathcal{T} and label spaces LL , or language-based merging. One issue with these techniques is that word labels are ambiguous. Instead, we let the data speak and optimize a label space automatically based on correlations in the firings of a pre-trained partitioned detector on different images, which is a proxy for perceptual similarity.

The cardinality penalty λ∣L∣\lambda|L| encourages a small and compact label space. A factorization of the loss Lc\mathcal{L}_{c} over the output space c∈Lkc\in L_{k} may seem restrictive. However, it does include the most common loss functions in detection: score distortion and Average Precision (AP). Section 4.2 discusses the exact loss functions used in our optimization.

Objective 6 mixes combinatorial optimization over LL with a 0-1 integer program over T\mathcal{T}. However, there is a simple reparametrization that lends itself to efficient optimization.

Crucially, the merge cost ctc_{\boldsymbol{t}} can be precomputed for any subset of labels t\boldsymbol{t}. This leads to a compact integer linear programming formulation of objective 6:

For two datasets, the above objective is equivalent to a weighted bipartite matching. For a higher number of datasets, it reduces to weighted graph matching and is NP-hard, but is practically solvable with integer linear programming .

2 Loss functions

The loss function in our constrained objective 6 is quite general and captures a wide range of commonly used losses. We highlight two: an unsupervised objective based on the distortion between partitioned and unified outputs, and Average Precision (AP) on a validation set.

measures the difference in detection scores between partitioned and unified detectors:

A drawback of this distortion measure is that it does not take task performance into consideration when optimizing the joint label space.

Average Precision.

The AP computation is computationally quite expensive. We will provide an optimized joint evaluation in our code.

These two loss functions allow us to train a partitioned detector and merge its output space after training, either maximizing the original evaluation metric (AP) or minimizing the change incurred by the unification.

Experiments

Our goal is to facilitate the training of a single model that performs well across datasets. In this section, we first introduce our dataset setup and implementation details. In Section 5.1, we analyze our key design choices for a partitioned detector baseline. In Section 5.2, we evaluate our unified detector and our unified label space learning algorithm. We further evaluate the unified detector in new test datasets in a cross-dataset evaluation (Section 5.3) without any training on the test domain.

Datasets. Our main training datasets are adopted from the Robust Vision Challenge (RVC)http://www.robustvision.net. These are four large datasets for object detection: COCO , OpenImages , Objects365 , and optionally Mapillary . To evaluate the generalization ability of the models, we follow MSeg to set up a ross-dataset evaluation protocol: we evaluate models on new test dataset without training on them. Specifically, we test on VIPER , Cityscapes , ScanNet , WildDash , KITTI , Pascal VOC , and CrowdHuman . A detailed description of all datasets is contained in the supplement. In our main evaluation, we use large and general datasets: COCO, Objects365, and OpenImages. Mapillary is relatively small and is specific to traffic scenes; we only add it for the RVC and cross-dataset experiments.

For each dataset, we use its official evaluation metric: for COCO, Objects365, and Mapillary, we use mAP at IoU thresholds 0.5 to 0.95. For OpenImages, we use the official modified mAP@0.5 that excludes unlabeled classes and enforces hierarchical labels . For the small datasets in cross-dataset evaluation, we use mAP at IoU threshold 0.5 for consistency with PascalVOC .

Implementation details. We use the CascadeRCNN detector with a shared region proposal network (RPN) across datasets. We evaluate two models in our experiments: a partitioned detector (i.e., detector with dataset-specific output heads) and a unified detector. For the partitioned detector, the last classification layers of all cascade stages are split between datasets. The unified detector uses CascadeRCNN as is.

Our implementation is based on Detectron2 . We adopt most of the default hyper-parameters for training. We use the standard data augmentation, including random flip and scaling of the short edge in the range $.WeuseSGDwithbaselearningrate0.01andbatchsize16over8GPUs.WeuseResNet50asthebackboneinourcontrolledexperimentsunlessspecifiedotherwise.Weusea. We use SGD with base learning rate 0.01 and batch size 16 over 8 GPUs. We use ResNet50 as the backbone in our controlled experiments unless specified otherwise. We use a2\times$ training schedule (180k iterations with learning rate dropped at the 120k and 160k iterations) in most experiments unless specified otherwise, regardless of the training data size.

We first evaluate the partitioned detector. We use dataset-specific outputs and do not merge classes between different datasets. During evaluation, we assume the target dataset is known and only look at the corresponding output head. As discussed in Section 4, our baseline highlights two basic components: uniform sampling of images between datasets and dataset-specific training objective. For these experiments we distinguish between modifications of the objective that merely sample data differently within each dataset (e.g. class-aware sampling), and changes to the loss functions (e.g. hierarchical losses).

We start from the baseline of . They simply collect all data from all datasets and train with a common loss. As is shown in Table 1, this biases the model to large datasets (OpenImages) and yields low performance for relatively small datasets (COCO). Sampling datasets uniformly (second row) trades the performance on smaller datasets with large datasets, and overall improves performance. On the other hand, both OpenImages and Objects365 are long-tailed and best train with advanced inter-dataset sampling strategy , namely class-aware sampling. Class-aware sampling significantly improves accuracy on OpenImages and Objects365. Combining the uniform dataset sampling and the intra-dataset class-aware sampling gives a further boost. Finally, OpenImages requires predicting a label hierarchy. For example, it requires predicting “vehicle” and “car” for all cars. This breaks the default cross-entropy loss that assumes exclusive class labels per object. We instead use a dedicated hierarchy-aware sigmoid cross-entropy loss for OpenImages . Specifically, for an annotated class label in OpenImages, we set all its parent classes as positives and ignore the losses over its descendant classes. Our partitioned detector combines both sampling strategies and the dataset-specific loss. The hierarchy-aware loss yields a significant +2.7+2.7mAP improvement on OpenImages alone, and does not degrades other datasets.

Dateset-specific vs. partitioned detectors. In our partitioned detector, training on multiple datasets resembles training separate individual models but with a shared detector. Table 2 compares training a partitioned detector on all datasets with dataset-specific models. We compare detectors under different training schedules (n×n\times the COCO default schedule). Each of the three dataset-specific models sees the same number of gradient updates as our partitioned detector. In a 2×2\times training schedule (180k iterations), single-dataset models generally perform better than a partitioned model, as each dataset is only trained for a 13×\frac{1}{3}\times schedule in the partitioned model. At a 6×6\times schedule, the partitioned detector starts to match dataset-specific models, and outperforms 2×2\times dataset-specific models under the same total iterations. In a 8×8\times schedule, all models converge. The partitioned detector surpasses the single-dataset model on COCO, and matches OpenImages and Objects365 models.

2 Unified multi-dataset detection

Next, we evaluate different ways to unify the label space.

Unified label space We run our label space learning algorithm from Section 2 based on the output of a partitioned detector with a ResNeSt backbone trained on COCO, Objects365, and OpenImages, with a total of 945945 disjoint classes. The hyperparameters are λ=0.5\lambda=0.5 and τ=0.25\tau=0.25. The optimization ends up with a unified label space with cardinality ∣L∣=701|L|=701. we compare our automated data-driven unification to human and language-based baselines. We use the official manually-crafted RVC taxonomy as the human expert baselinehttps://github.com/ozendelait/rvc_devkit/blob/master/objdet/obj_det_mapping.csv.

Over two-thirds of our learned label space agrees with the human expert. Figure 3 highlights some of the differences. Our unification successfully groups similar concepts with different descriptions (“Cow” and “Cattle”), and is not distracted by spurious linguistic matches (“American football” and “football”). Interestingly, the learned label space splits the “oven” classes from COCO, Objects365, and OpenImages, even though they share the same word. A visual examination reveals that they are visually dissimilar due to different underlying definitions of the “oven” concept in the different datasets: COCO ovens include the cooktop, OpenImages ovens include the control panel, and Objects365 ovens are just the front door. Our data-driven taxonomy reconciliation is able to detect such distinctions, which are missed by word-level approaches.

We next quantitatively compare our learned label space with alternatives. For each label space, we retrain a multi-dataset detector with that label space. During training, as with our partitioned model, we only apply training losses to the classes that are annotated in the source dataset. We compare our learned label space to a “best effort” human baseline and a language-based baseline. For the language-based baseline, we replace the cost measurement defined in Section 4.2 with the cosine distance between the GloVe word embeddings , and run the same integer linear program. Table 3 shows the results. We repeat the training for three runs with different random seeds and report the mean and standard deviation. The four label spaces agree on most classes and the overall mAP is thus close. Our automatically constructed label space consistently outperforms the human expert baseline, with a healthy 0.30.3 mAP margin on average. The improvement appears statistically stable under multiple training runs. Notably, the relative improvement of our model over the expert is larger than the expert’s improvement over the language-based baseline.

Hyper-parameter choices. Table 4 ablates the hyper-parameters λ\lambda and τ\tau of the label space learning algorithm (Section 2). Our algorithm is robust to the cardinality penalty factor λ\lambda. Varying the cardinality penalty λ\lambda from 0.10.1 to 1.01.0 only affects the size of the label space by 33. The pruning threshold τ\tau has a larger impact on the label space size, but not the final performance. We use λ=0.5\lambda=0.5 and τ=0.25\tau=0.25 for a good balance between the label space size and overage performance.

Unified vs. partitioned detectors. We next compare unified detectors with and without retraining using the joint taxonomy, a partitioned detector, and an ensemble of dataset-specific detectors. The partitioned detector and the ensemble need to know the target domain at test time, while the unified models do not. This means that the unified models can be deployed without any modification in new domains, while the alternatives must know which domain they are in. Table 5 shows the results. A partitioned detector outperforms a dataset-specific ensemble under the same conditions (Table 5 bottom), especially on the “small” COCO dataset. An offline unification loses some accuracy, but this is regained when retraining the model under the unified taxonomy (Table 5 top). Crucially, the unified models do not need to know what domain they are in at test time.

3 Cross-dataset evaluation

We evaluate the generalization ability of object detectors by evaluating them in new test domains not seen during training. In this setting, we do not assume to know the test classes ahead of time. To allow for a fair and unbiased evaluation, we use a simple language-based matching to find the test-to-train label correspondence. Specifically, we calculate the GloVe word embedding distances between each test label and the training label, and match the test label to its closest training label. If multiple training labels match, we break ties in a fixed order: COCO, Objects365, OpenImages, and MapillaryWe also tried evaluating under different orders, and find the listed order to perform best for all methods..

We compare both our multi-dataset models (partitioned or unified) to single-dataset models. We use all four RVC training sets to train the multi-dataset models. Specifically, we start from a 6×6\times schedule model trained on the three large datasets, and add Mapillary in a 2×2\times fine-tuning schedule with 10×10\times smaller learning rate. We compare all models under the same schedule except for the Mapillary model, for which a 2×2\times schedule performs better than longer schedules., hyperparameters, and detection models. In addition, we also compare to the ensemble of the four single-dataset models trained analogously to the partitioned model. For reference, we also show the performance of detectors trained on the training set of each test dataset. This serves as an oracle “upper bound” that has seen the test domain and label space. Note that KITTI and WildDash are small and do not have a validation set. We thus evaluate on the training set and do not provide the oracle model.

Table 6 shows the results. The COCO model exhibits reasonable performances of some test datasets, such as Pascal VOC and CrowdHuman. However, its performance is less than satisfactory on datasets such as ScanNet, whose label space differs significantly from COCO. Training on the more diverse Objects365 dataset yields higher accuracy in the indoor domain, but loses ground on VOC and CrowdHuman, which are more similar to COCO. Training on all datasets, either with a partitioned detector (row 6) or a unified one (row 7) yields generally good performance on all test datasets. Notably, both our detectors perform better than the ensemble of the 4 single dataset models (row 5), showing that the multi-dataset models learned more general features. On Pascal VOC, both multi-dataset models outperform the VOC-trained upper-bound without seeing VOC training images. Our unified model outperforms the partitioned detector overall and operates on a unified taxonomy.

4 Scale up to large models

Next, we scale up our unified detector with a large backbone to develop a ready-to-deploy object detector. We used a ResNeSt200 backbone and followed the same training procedure as in Section 4 with an 8×8\times schedule. The training took ∼\sim16 days on a server with 8 Quadro RTX 6000 GPUs. Table 7 shows our single model achives 52.952.9 mAP on COCO, 60.660.6 mAP on OpenImages, and 33.733.7 mAP on Objects 365. We compare to state-of-the-art results with comparable baselines on each individual dataset. On COCO, our result improves the COCO-only ResNeSt200 model, by 22 mAP with the same detector, thanks to our ability to train with more data. On OpenImages, our result matches the best single model in the OpenImages 2019 Challenge, TSD , with a comparable backbone (SENet154-DCN of TSD). On Objects365, we outperform the 2019 Object365 detection challenge winner by 22 mAP points.

Conclusion

We presented a simple recipe for training a single object detector across multiple datasets and a formulation to automatically construct a unified taxonomy. Our resulting detector can be deployed in new domains without additional knowledge. We hope our model makes object detection more accessible to general users.

Limitations. Our label space learning algorithm currently uses only visual cues, integrating language cues as auxiliary information may further improve the performance. Our formulation currently does not consider label hierarchies, and the resulting label space treats COCO person and OpenImages boy as two independent classes. We leave incorporating label hierarchies as exciting future work.

Acknowledgments. This material is based upon work supported by the National Science Foundation under Grant No. IIS-1845485 and IIS-2006820. Xingyi is supported by a Facebook Fellowship.

References

Appendix A Dataset details

Table 8 lists the datasets we used in our experiments. We use the Robust Vision Challengehttp://www.robustvision.net official release of each dataset. Specifically, we use the standard 2017 train/ validation split for COCO , the Challenge-2019 release of OpenImages , and the default version of Objects365 and Mapillary . For ScanNet , as there is no standard train/ validation split, we use the first 80%80\% scenes (sorted by scene ID) as training and the last 20%20\% scene as validation. For KITTI , we used the RVC challenge version that has instance-segmentation version, which contains 200 images. For WildDash , we use the public version for evaluation, and report standard mAP performance. We don’t consider the negative label metric in the official website. For CrowdHuman , we use the visible bounding box annotation, and report mAP instead of the missing rate as the official metric. We use the official train/ validation split and the official evaluation metrics for VIPER , Cityscapes , and Pascal VOC .

Appendix B Computation of label space learning algorithm and pruning

For an aggressive enough threshold τ\tau, the number of potential merges ∣T′∣|\mathcal{T}^{\prime}| remains manageable. We greedily grow T′\mathcal{T}^{\prime} by first enumerating all feasible two-class merges (∣t∣=2|\boldsymbol{t}|=2), then three-class merges, and so on. The detailed algorithm diagram is shown in Algorithm 1. The runtime of this greedy algorithm is O(∣T′∣max⁡i∣L^i∣)O(|\mathcal{T}^{\prime}|\max_{i}|\hat{L}^{i}|). In practice, the cost computation took a few seconds for the distortion loss function and about 10 minutes for the AP loss (due to the need to repeatedly recompute AP). The integer programming solver finds the optimal solution within one second in both cases.

Appendix C Adding new datasets to a label space

While we tend to keep the training domains and label space large and comprehensive, it is inevitable in practice that more fine-grained labels or specific testing domains are needed. Given a learned a unified label space on an existing set of training datasets, we use a simple label space expansion algorithm to allow adding more datasets and labels after the unified detector is trained.

Similar to our unified label space learning algorithm, we run the unified detector on the new training data. We evaluate the AP between each class in the new dataset annotation and each class in the unified label space. We merge the new class into the existing class that gives the lowest merge cost (Section. 4.2). In our experiments, add Mapillary dataset to our label space we using the AP loss. If the cost is lower than a threshold (AP change <5<5 AP in our implementation). Otherwise, we append the new class to the unified label space as a single class.

Appendix D Discussion on label hierarchy

Different datasets may contain different label granularities for the same concept, and there exists label hierarchies inter or intra datasets. For example, Objects365 does not have a “bird” category, but has more fine-grained bird species like parrot, pigeon, and swan, while most other datasets only annotate “bird”. Our label space optimization algorithm automatically handles the hierarchical label space issue: the fine-grained birds in Objects365 will not merge with COCO birds because this merge introduces many false-positives for the fine-grained birds in Objects365 and yields a large cost. Our unified label space will contain both the general “bird” class and each fine-grained class. The model trained on the unified label space is expected to predict both the coarse “bird” label and the fine-grained label in testing.

Appendix E Instance segmentation

We further evaluate our label space learning algorithm and unified training framework on instance segmentation. We follow the Robust vision challenge set up to use 8 datasets: COCO, OpenImages, Mapillary, ScanNet, VIPER, CityScapes, WildDash and KITTI (the same as Table 8, except OpenImages segmentation set has 300 instead of 500 classes.). Again, we leave WildDash and KITTI as testing only as they are small and similar to CityScapes and Mapillary. We run our label space learning algorithm (Section. 4) on the remaining six datasets, resulting a unified label space of 358358 classes. We use CascadeRCNN with a standard mask head as the detector, and train a 2×2\times schedule with ResNet50. The dataset-specific models are trained with 1×1\times or 2×2\times schedule depending on their size.

Table. 9 compares the unified detector to dataset specific models. As expected, no single dataset-specific model performs well on all test domains. Our unified model performs consistently good on all training datasets. More importantly, it generalizes the best to the new test datasets (KITTI and WildDash) than any single dataset model.