SUDS: Scalable Urban Dynamic Scenes
Haithem Turki, Jason Y. Zhang, Francesco Ferroni, Deva Ramanan
Introduction
Scalable geometric reconstructions of cities have transformed our daily lives, with tools such as Google Maps and Streetview becoming fundamental to how we navigate and interact with our environments. A watershed moment in the development of such technology was the ability to scale structure-from-motion (SfM) algorithms to city-scale footprints . Since then, the advent of Neural Radiance Fields (NeRFs) has transformed this domain by allowing for photorealistic interaction with a reconstructed scene via view synthesis.
Recent works have attempted to scale such representations to neighborhood-scale reconstructions for virtual drive-throughs and photorealistic fly-throughs . However, these maps remain static and frozen in time. This makes capturing bustling human environments—complete with moving vehicles, pedestrians, and objects—impossible, limiting the usefulness of the representation.
Challenges. One possible solution is a dynamic NeRF that conditions on time or warps a canonical space with a time-dependent deformation . However, reconstructing dynamic scenes is notoriously challenging because the problem is inherently under-constrained, particularly when input data is constrained to limited viewpoints, as is typical from egocentric video capture . One attractive solution is to scale up reconstructions to many videos, perhaps collected at different days (e.g., by an autonomous vehicle fleet). However, this creates additional challenges in jointly modeling fixed geometry that holds for all time (such as buildings), geometry that is locally static but transient across the videos (such as a parked car), and geometry that is truly dynamic (such as a moving person).
SUDS. In this paper, we propose SUDS: Scalable Urban Dynamic Scenes, a 4D representation that targets both scale and dynamism. Our key insight is twofold; (1) SUDS makes use of a rich suite of informative but freely available input signals, such as LiDAR depth measurements and optical flow. Other dynamic scene representations require supervised inputs such as panoptic segmentation labels or bounding boxes, which are difficult to acquire with high accuracy for our in-the-wild captures. (2) SUDS decomposes the world into 3 components: a static branch that models stationary topography that is consistent across videos, a dynamic branch that handles both transient (e.g., parked cars) and truly dynamic objects (e.g., pedestrians), and an environment map that handles far-field objects and sky. We model each branch using a multi-resolution hash table with scene partitioning, allowing SUDS to scale to an entire city spanning over 100 .
Contributions. We make the following contributions: (1) to our knowledge, we build the first large-scale dynamic NeRF, (2) we introduce a scalable three-branch hash table representation for 4D reconstruction, (3) we present state-of-the-art reconstruction on 3 different datasets. Finally, (4) we showcase a variety of downstream tasks enabled by our representation, including free-viewpoint synthesis, 3D scene flow estimation, and even unsupervised instance segmentation and 3D cuboid detection.
Related Work
The original Neural Radiance Fields (NeRF) paper inspired a wide body of follow-up work based on the original approach. Below, we describe a non-exhaustive list of such approaches along axes relevant to our work.
Scale. The original NeRF operated with bounded scenes. NeRF++ and mip-NeRF 360 use non-linear scene parameterization to model unbounded scenes. However, scaling up the size of the scene with a fixed size MLP leads to blurry details and training instability while the cost of naively increasing the size of the MLP quickly becomes intractable. BungeeNeRF introduced a coarse-to-fine approach that progressively adds more capacity to the network representation. Block-NeRF and Mega-NeRF partition the scene spatially and train separate NeRFs for each partition. To model appearance variation, they incorporate per-image embeddings like NeRF-W . Our approach similarly partitions the scene into sub-NeRFs, making use of depth to improve partition efficiency and scaling over an area 200x larger than Block-NeRF’s Alamo Square Dataset. Both of these methods work only on static scenes.
Dynamics. Neural 3D Video Synthesis and Space-time Neural Irradiance Fields add time as an input to handle dynamic scenes. Similar to our work, NSFF , NeRFlow , and DyNeRF incorporate 2D optical flow input and warping-based regularization losses to enforce plausible transitions between observed frames. Multiple methods instead disentangle scenes into a canonical template and per-frame deformation field. BANMo further incorporates deformable shape models and canonical embeddings to train articulated 3D models from multiple videos. These methods focus on single-object scenes, and all but and use single video sequences.
While many of the previous works use segmentation data to factorize dynamic from static objects, D2NeRF does this automatically through regularization and explicitly handling shadows. Neural Groundplans uses synthetic data to do this decomposition from a single image. We borrow some of these ideas and scale beyond synthetic and indoor scenes.
Object-centric approaches. Several approaches represent scenes as the composition of per-object NeRF models and a background model. NSG is most similar to us as it also targets automotive data but cannot handle ego-motion as our approach can. None of these methods target multi-video representations and are fundamentally constrained by the memory required to represent each object, with NSG needing over 1TB of memory to represent a 30 second video in our experience.
Semantics. Follow-up works have explored additional semantic outputs in addition to predicting color. Semantic-NeRF adds an extra head to NeRF that predicts extra semantic category logits for any 3D position. Panoptic-NeRF and Panoptic Neural Fields extend this to produce panoptic segmentations and the latter uses a similar bounding-box based object and background decomposition as NSG. NeSF generalizes the notion of a semantic field to unobserved scenes. As these methods are highly reliant on accurate annotations which are difficult to reliably obtain in the wild at our scale, we instead use a similar approach to recent works that distill the outputs of 2D self-supervised feature descriptors into 3D radiance fields to enable semantic understanding without the use of human labels and extend them to larger dynamic settings.
Fast training. The original NeRF took 1-2 days to train. Plenoxels and DVGO directly optimize a voxel representation instead of an MLP to train in minutes or even seconds. TensoRF stores its representation as the outer product of low-rank tensors, reducing memory usage. Instant-NGP takes this further by encoding features in a multi-resolution hash table, allowing training and rendering to happen in real-time. We use these tables as the base block of our three-branch representation and use our own hashing method to support dynamics across multiple videos.
Depth. Depth provides a valuable supervisory signal for learning high-quality geometry. DS-NeRF and Dense Depth Priors incorporate noisy point clouds obtained by structure from motion (SfM) in the loss function during optimization. Urban Radiance Fields supervises with collected LiDAR data. We also use LiDAR but demonstrate results on dynamic environments.
Approach
Our goal is to learn a global representation that facilitates free-viewpoint rendering, semantic decomposition, and 3D scene flow at arbitrary poses and time steps. Our method takes as input ordered RGB images from videos (taken at different days with diverse weather and lighting conditions) and their associated camera poses. Crucially, we make use of additional data as “free” sources of supervision given contemporary sensor rigs and feature descriptors. Specifically, we use (1) aligned sparse LiDAR depth measurements, (2) 2D self-supervised pixel (DINO ) descriptors to enable semantic manipulation, and (3) 2D optical flow predictions to model scene dynamics. All model inputs are generated without any human labeling or intervention.
2 Representation
Preliminaries. We build upon NeRF , which represents a scene within a continuous volumetric radiance field that captures both geometry and view-dependent appearance. It encodes the scene within the weights of a multilayer perceptron (MLP). At render time, NeRF projects a camera ray r for each image pixel and samples along the ray, querying the MLP at sample position and ray viewing direction d to obtain opacity and color values and . It then composites a color prediction for the ray using numerical quadrature , where and is the distance between samples. The training process optimizes the model by sampling batches of image pixels and minimizing the loss function \sum_{\textbf{r}\in\mathcal{R}}\big{\lVert}{C(\textbf{r})-\hat{C}(\textbf{r})}\big{\rVert}^{2}. NeRF samples rays through a two-stage hierarchical sampling process and uses frequency encoding to capture high-frequency details. We refer the reader to for more details.
Scene composition. To model large-scale dynamic environments, SUDS factorizes the scene into three branches: (a) a static branch containing non-moving topography consistent across videos, (b) a dynamic branch to disentangle video-specific objects , moving or otherwise, and (c) a far-field environment map to represent far-away objects and the sky, which we found important to separately model in large-scale urban scenes .
However, conventional NeRF training with MLPs is computationally prohibitive at our target scales. Inspired by Instant-NGP , we implement each branch using multiresolution hash tables of -dimensional feature vectors followed by a small MLP, along with our own hash functions to index across videos.
Static branch. We generate RGB images by combining the outputs of our three branches. The static branch maps the feature vector obtained from the hash table into a view-dependent color and a view-independent density . To model lighting variations which could be dramatic across videos but smooth within a video, we condition on a latent embedding computed as a product of a video-specific matrix and a fourier-encoded time index (as in ):
Dynamic branch. While the static branch assumes the density is static, the dynamic branch allows both the density and color to depend on time (and video). We therefore omit the latent code when computing the dynamic radiance. Because we find shadows to play a crucial role in the appearance of urban scenes (Fig. 3), we explicitly model a shadow field of scalar values , used to scale down the static color (as done in ):
Rendering. We derive a single density and radiance value for any position by computing the weighted sum of the static and dynamic components, combined with the pointwise shadow reduction:
We then calculate the color for a camera ray r with direction d at a given frame t and video vid by accumulating the transmittance along sampled points along the ray, forcing the ray to intersect the far-field environment map if it does not hit geometry within the foreground:
which are combined into a single value at any 3D location and rendered into per camera ray, following the equations for color (10, 3.2).
3 Optimization
We jointly optimize all three of our model branches along with the per-video weight matrices by sampling random batches of rays across our input videos and minimizing the following loss:
Reconstruction losses. We minimize the L2 photometric loss \mathcal{L}_{c}(\mathbf{r})=\big{\lVert}{C(\textbf{r})-\hat{C}(\textbf{r})}\big{\rVert}^{2} as in the original NeRF equation . We similarly minimize the L1 difference \mathcal{L}_{f}(\textbf{r})=\big{\lVert}{F(\textbf{r})-\hat{F}(\textbf{r})}\big{\rVert}_{1} between the feature outputs of the teacher model and that of our network.
To make use of our depth measurements, we project the LiDAR sweeps onto the camera plane and compare the expected depth with the measurement :
Flow. We supervise our 3D scene flow predictions based on 2D optical flow (Sec. 4.1). We generate a 2D displacement vector for each camera ray by first predicting its position in 3D space as the weighted sum of the scene flow neighbors along the ray:
which we then “render” into 2D using the camera matrix of the neighboring frame index. We minimize its distance from the observed optical flow via \mathcal{L}_{o}(\textbf{r})=\sum_{t^{\prime}\in}\big{\lVert}{X(\textbf{o})-\hat{X}_{t^{\prime}}(\textbf{r})}\big{\rVert}_{1}. We anneal over time as these estimates are noisy.
3D warping. The above loss ensures that rendered 3D flow will be consistent with the observed 2D flow. We also found it useful to enforce 3D color (and feature) constancy; i.e., colors remain constant even when moving. To do so, we use the predicted forward and backward 3D flow and to advect each sample along the ray into the next/previous frame:
The warped radiance and density are rendered into warped color and feature (10, 3.2). We add a loss to ensure that the warped color (and feature) match the ground-truth input for the current frame, similar to . As in NSFF , we found it important to downweight this loss in ambiguous regions that may contain occlusions. However, instead of learning explicit occlusion weights, we take inspiration from Kwea’s method and use the difference between the dynamic geometry and the warped dynamic geometry to downweight the loss:
resulting in the following warping loss terms:
Flow regularization. As in prior work we use a 3D scene flow cycle term to encourage consistency between forward and backward scene flow predictions, down-weighing the loss in areas ambiguous due to occlusions:
with vid omitted for brevity. We also encourage spatial and temporal smoothness through as described in Sec. C. We finally regularize the magnitude of predicted scene flow vectors to encourage the scene to be static through \mathcal{L}_{slo}(\textbf{r})=\sum_{t^{\prime}\in[\textbf{t}-1,\textbf{t}+1]}\sum_{\textbf{x}}\big{\lVert}{s_{t^{\prime}}(\textbf{x},\textbf{t})}\big{\rVert}_{1}.
Static-dynamic factorization. As physically plausible solutions should have any point in space occupied by either a static or dynamic object, we encourage the spatial ratio of static vs dynamic density to either be 0 or 1 through a skewed binary entropy loss that favors static explanations of the scene :
and with set to 1.75, and further penalize the maximum dynamic ratio along each ray.
Shadow loss. We penalize the squared magnitude of the shadow ratio along each ray to prevent it from over-explaining dark regions .
Experiments
We demonstrate SUDS’s city-scale reconstruction capabilities by presenting quantitative results against baseline methods (Table 1). We also show initial qualitative results for a variety of downstream tasks (Sec. 4.2). Even though we focus on reconstructing dynamic scenes at city scale, to faciliate comparisons with prior work, we also show results on small-scale but highly-benchmarked datasets such as KITTI and Virtual KITTI 2 (Sec. 4.3). We evaluate the various components of our method in Sec. 4.4.
2D feature extraction. We use Amir et al’s feature extractor implementation based on the dino_vits8 model. We downsample our images to fit into GPU memory and then upsample with nearest neighbor interpolation. We L2-normalize the features at the 11th layer of the model and reduce the dimensionality to 64 through incremental PCA .
Flow supervision. We explored using an estimator trained on synthetic data in addition to directly computing 2D correspondences from DINO itself . Although the correspondences are sparse (less than 5% of pixels) and expensive to compute, we found its estimates more robust and use it for our experiments unless otherwise stated.
Training. We train SUDS for 250,000 iterations with 4098 rays per batch and use a proposal sampling strategy similar to Mip-NeRF 360 (Sec. B). We use Adam with a learning rate of decaying to .
Metrics. We report quantitative results based on PSNR, SSIM , and the AlexNet implementation of LPIPS .
2 City-Scale Reconstruction
City-1M dataset. We evaluate SUDS’s large-scale reconstruction abilities on our collection of 1.28 million images across 1700 videos gathered across a 105 urban area using a vehicle-mounted platform with seven ring cameras and two LiDAR sensors. Due to the scale, we supervise optical flow with an off-the-shelf estimator trained on synthetic data instead of DINO for efficiency.
Baselines. We compare SUDS to the official Mega-NeRF implementation alongside two variants: Mega-NeRF-T which directly adds time as an input parameter to compute density and radiance, and Mega-NeRF-A which instead uses the latent embedding used by SUDS.
Results. We train both SUDS and the baselines using 48 cells and summarize our results in Table 1. SUDS outperforms all Mega-NeRF variants by a large margin. We provide qualitative results on view synthesis, static/dynamic factorization, unsupervised 3D instance segmentation and unsupervised 3D cuboid detection in Fig. 5. We present additional qualititive tracking results in Fig. 7.
Instance segmentation. We derive the instance count as in prior work by sampling dynamic density values , projecting those above a given threshold onto a discretized ground plane before applying connected component labeling. We apply k-means to obtain 3D centroids and volume render instance predictions as for semantic segmentation.
3D cuboid detection. After computing point-wise instance assignments in 3D, we derive oriented bounding boxes based on the PCA of the convex hull of points belonging to each instance .
Semantic segmentation. Note the above tasks of instance segmentation and 3D cuboid detection do not require any additional labels as they make use of geometric clustering. We now show that the representation learned by SUDS can also enable downstream semantic tasks, by making use of a small number of 2D segmentation labels provided on a held-out video sequence. We compute the average 2D DINO descriptor for each semantic class from the held out frames and derive 3D semantic labels for all reconstructions by matching each 3D descriptor to the closest class centroid. This allows to produce 3D semantic label fields that can then be rendered in 2D as shown in Fig. 5.
3 KITTI Benchmarks
Baselines. We compare SUDS to SRN , the original NeRF implementation , a variant of NeRF taking time as an additional input, NSG , and PNF . Both NSG and PNF are trained and evaluated using ground truth object bounding box and category-level annotations.
Image reconstruction. We compare SUDS’s reconstruction capabilities using the same KITTI subsequences and experimental setup as prior work . We present results in Table 3. As PNF’s implementation is not publicly available, we rely on their reported numbers. SUDS surpasses the state-of-the-art in PSNR and SSIM.
Novel view synthesis. We demonstrate SUDS’s capabilities to generate plausible renderings at time steps unseen during training. As NSG does not handle scenes with ego-motion, we use subsequences of KITTI and Virtual KITTI 2 with little camera movement. We evaluate the methods using different train/test splits, holding out every 4th time step, every other time step, and finally training with only one in every four time steps. We summarize our findings in Table 2 along with qualitative results in Fig. 6. SUDS achieves the best results across all splits and metrics. Both NeRF variants fail to properly represent the scene, especially in dynamic areas. Although we provide NSG with the ground truth object poses at render time, it fails to learn a clean decomposition between objects and the background, especially as the number of training view decreases, and generates ghosting artifacts near areas of movement.
4 Diagnostics
We ablate the importance of major SUDS components by removing their respective loss terms along with occlusion weights, the latent embedding used to compute static color , and separate model branches (Sec. D). We run all approaches for 125,000 iterations across our datasets and summarize the results in Table 4. Although all components help performance, flow-based warping is by far the single most important input. Interestingly, depth is the least crucial input, suggesting that SUDS can generalize to settings where depth measurements are not available.
Conclusion
We present a modular approach towards building dynamic neural representations at previously unexplored scale. Our multi-branch hash table structure enables us to disentangle and efficiently encode static geometry and transient objects across thousands of videos. SUDS makes use of unlabeled inputs to learn semantic awareness and scene flow, allowing it to perform several downstream tasks while surpassing state-of-the-art methods that rely on human labeling. Although we present a first attempt at building city-scale dynamic environments, many open challenges remain ahead of building truly photorealistic representations.
Acknowledgments
This research was supported by the CMU Argo AI Center for Autonomous Vehicle Research.
References
Supplemental Materials
Appendix A Tracking
We can compute mask and keypoint-level correspondences across frames after detecting instances (Sec. 4.2) by using Best-Buddies similarity on features within or between instances. As a 3D representation, SUDS can track correspondences through 2D occluders. We show an example in Fig. 7.
Appendix B Proposal Sampling
We use a proposal sampling strategy similar to Mip-NeRF 360 that first queries a lightweight occupancy proposal network at uniform intervals along each camera ray and then picks additional samples based on the initial samples. We model our proposal network with separate hash table-backed static and dynamic branches as in Sec. 3.2. We train each branch of the proposal network with histogram loss using the weights of the respective branch of our main model and regularize the resulting sample distances and weights using distortion loss . We find that proposal sampling gives a 2-4x speedup.
Appendix C Smoothness Priors
We use the same spatial and temporal smoothness priors as NSFF to regularize our scene flow. We specifically denote:
where x and indicate neighboring points along the camera ray r.
Appendix D Ablation Details
w/o Depth loss. We remove depth from the reconstruction loss term:
w/o Optical flow loss. We remove optical flow from the reconstruction loss term:
w/o Warping loss. We remove all warping and flow-related loss terms:
w/o Appearance embedding. We compute static color without the latent embedding vector :
w/o Occlusion weights. We do not use occlusion weights (24) to downweight the warping loss terms (25, 26):
w/o Separate branches. We generate all model outputs using a single time-dependent branch:
We accordingly remove factorization-related loss terms:
Appendix E Additional Training Details
We divide City-1M into 48 cells using camera-based k-means clustering. Each cell covers 2.9 and 32k frames across 98 videos on average. We evaluate the effect of geographic coverage and number of frames/videos on cell quality in Table 5. We train with 1 A100 (40 GB) GPU per cell for 2 days (same for each KITTI scene). We can fit all cells on a single A100 at inference time.
Appendix F Assets
City-1M. Our dataset is constructed from street-level videos collected across a vehicle fleet with seven ring cameras that collect 2048x1550 resolution images at 20 Hz with a combined 360° field of view. Both VLP-32C LiDAR sensors are synchronized with the cameras and produce point clouds with 100,000 points at 10 Hz on average. We localize camera poses using a combination of GPS-based and sensor-based methods.
Third-party assets. We primarily base the SUDS implementation on Nerfstudio and tiny-cuda-nn along with various utilities from OpenCV , Scikit , and Amir et al’s feature extractor implementation , all of which are freely available for noncommercial use. KITTI is similarly available under an Apache license, whereas VKITTI2 uses the noncommercial CC BY-NC-SA 3.0 license.
Appendix G Limitations
Video boundaries. Although our global representation of static geometry is consistent across all videos used for reconstruction, all dynamic objects are video-specific. Put otherwise, our method does not allow us to extrapolate the movement of objects outside of the boundaries of videos from which they were captured, nor does it provide a straightforward way of rendering dynamic visuals at boundaries where camera rays intersect regions with training data originating from disjoint video sequences.
Camera accuracy. Accurate camera extrinsics and intrinsics are arguably the largest contributors to high NeRF rendering quality. Although multiple efforts attempt to jointly optimize camera parameters during NeRF optimization, we found the results lacking relative to using offline structure-from-motion based approaches as a preprocessing step.
Flow quality. Although our method tolerates some degree of noisiness in the supervisory optical flow input, high-quality flow still has a measurable impact on model performance (and completely incorrect supervision degrades quality). We also assume that flow is linear between observed timestamps to simplify our scene flow representation.
Resources. Modeling city scale requires a large amount of dataset preprocessing, including, but not limited to: extracting DINO features, computing optical flow, deriving normalized coordinate bounds, and storing randomized batches of training data to disk. Collectively, our intermediate representation required more than 20TB of storage even after compression.
Shadows. SUDS attempts to disentangle shadows underneath transient objects. However, if a shadow is present in all observations for a given location (eg: a parking spot that is always occupied, even by different cars), SUDS may attribute the darkness to the static topology, as evidenced in several of our videos, even if the origin of the shadow is correctly assigned to the dynamic branch.
Instance-level tasks. Although we provide initial qualitative results on instance-level tasks as a first step towards true 3D segmentation backed by neural radiance field, SUDSis not competitive with conventional approaches.
Appendix H Societal Impact
As SUDS attempts to model dynamic urban scenes with pedestrians and vehicles, our approach carries surveillance and privacy concerns related to the intentional or inadvertent capture or privacy-sensitive information such as human faces and vehicle license plate numbers. As we distill semantic knowledge into SUDS, we are able to (imperfectly) filter out either entire categories (people) or components (faces) at render time. However this information would still reside in the model itself. This could in turn be mitigated by preprocessing the input data used to train the model.