Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, Jason Baldridge

Introduction

Vision-and-Language Navigation (VLN) tasks require computational agents to mediate the relationship between language, visual scenes and movement. Datasets have been collected for both indoor Anderson et al. (2018b); Thomason et al. (2019b); Qi et al. (2020) and outdoor Chen et al. (2019); Mehta et al. (2020) environments; success in these is based on clearly-defined, objective task completion rather than language or vision specific annotations. These VLN tasks fall in the Goldilocks zone: they can be tackled – but not solved – with current methods, and progress on them makes headway on real world grounded language understanding.

We introduce Room-across-Room (RxR), a VLN dataset that addresses gaps in existing ones by (1) including more paths that (2) counter known biases in existing datasets, and (3) collecting an order of magnitude more instructions for (4) three languages (English, Hindi and Telugu) while (5) capturing annotators’ 3D pose sequences. As such, RxR includes dense spatiotemporal grounding for every instruction, as illustrated in Figure 1.

We provide monolingual and multilingual baseline experiments using a variant of the Reinforced Cross-Modal Matching agent Wang et al. (2019). Performance generally improves by using monolingual learning, and by using RxR’s follower paths as well as its guide paths. We also concatenate R2R and RxR annotations as a simple multitask strategy Wang et al. (2020): the agent trained on both datasets obtains across the board improvements.

RxR contains 126K instructions covering 16.5K sampled guide paths and 126K human follower demonstration paths. The dataset is available.https://github.com/google-research-datasets/RxR We plan to release a test evaluation server, our annotation tool, and code for all experiments.

Motivation

A number of VLN datasets situated in photo-realistic 3D reconstructions of real locations contain human instructions or dialogue: R2R Anderson et al. (2018b), Touchdown Chen et al. (2019); Mehta et al. (2020), CVDN Thomason et al. (2019b) and REVERIE Qi et al. (2020). RxR addresses shortcomings of these datasets—in particular, multilinguality, scale, fine-grained word grounding, and human follower demonstrations (Table 1). It also addresses path biases in R2R. More broadly, our work is also related to instruction-guided household task benchmarks such as ALFRED Shridhar et al. (2020) and CHAI Misra et al. (2018). These synthetic environments provide interactivity but are generally less diverse, less visually realistic and less faithful to real world structures than the 3D reconstructions used in VLN.

Multilinguality. The dominance of high resource languages is a pervasive problem as it is unclear that research findings generalize to other languages Bender (2009). The issue is particularly severe for VLN. Chen and Mooney (2011) translated(∼\sim1K) English navigation instructions into Chinese for a game-like simulated 3D environment. Otherwise, all publicly available VLN datasets we are aware of have English instructions.

To enable multilingual progress on VLN, RxR includes instructions for three typologically diverse languages: English (en), Hindi (hi), and Telugu (te). The English portion includes instructions by speakers in the USA (en-US) and India (en-IN). Unlike Chen and Mooney (2011) and like the TyDi-QA multilingual question answering dataset Clark et al. (2020), RxR’s instructions are not translations: all instructions are created from scratch by native speakers. This especially matters for VLN, as different languages encode spatial and temporal information in idiosyncratic ways–e.g., how contact/support relationships are expressed Munnich et al. (2001), frame of reference Haun et al. (2011), and how temporal accounts are expressed Bender and Beller (2014).

Scale. Embodied language tasks suffer from a relative paucity of training data; for VLN, this has led to a focus on data augmentation Fried et al. (2018); Tan et al. (2019), pre-training Wang et al. (2019); Huang et al. (2019); Li et al. (2019), multi-task learning Wang et al. (2020) and better generalization through piece-wise curriculum design Zhu et al. (2020). To address this shortage, for each language RxR contains 14K paths with 3 instructions per path, for a total of 126K instructions and 10M words (based on whitespace tokenization). As illustrated in Table 1, this is an order of magnitude larger than previous datasets.

Fine-Grained Grounding. Like R2R, RxR’s instructions are collected by immersing Guide annotators in a simulated first-person environment backed by the Matterport3D dataset Chang et al. (2017) and asking them to describe predefined paths. RxR also enhances each instruction with dense spatiotemporal groundings. Guides speak as they move and later transcribe their audio; our annotation tool records their 3D poses and time-aligns the entire pose trace with words in the transcription. Instructions and pose traces can thus be aligned with any Matterport data including surface reconstructions (Figure 1), RGB-D panoramas (Figure 4), and 2D and 3D semantic segmentations.

Follower Demonstrations. Annotators also act as Followers who listen to a Guide’s instructions and attempt to follow the path. In addition to verifying instruction quality, this allows us to collect a play-by-play account of how a human interpreted the instructions, represented as a pose trace. Guide and Follower pose traces provide dense spatiotemporal alignments between instructions, visual percepts and actions – and both perspectives are useful for agent training.

Path Desiderata. R2R paths span 4–6 edges and are the shortest paths from start to goal. Thomason et al. (2019a) showed that agents can exploit effective priors over R2R paths, and Jain et al. (2019) showed that R2R paths encourage goal seeking over path adherence. These matter both for generalization to new environments and fidelity to the descriptions given in the instruction—otherwise, strong performance might be achieved by agents that mostly ignore the language. RxR addresses these biases by satisfying four path desiderata:

High variance in path length, such that agents cannot simply exploit a strong length prior.

Paths may approach their goal indirectly, so agents cannot simply go straight to the goal.

Naturalness: paths should not enter cycles or make continual direction changes that would be difficult for people to describe and follow.

Uniform coverage of environment viewpoints, to maximize the diversity of references to visual landmarks and objects over all paths.

This increases RxR’s utility for testing agents’ ability to ground language. It also makes RxR a more challenging VLN dataset—but one for which human followers still achieve a 93.9% success rate.

Two-Level Path Sampling

We satisfy desiderata 1-3 using a two-level procedure. At a high-level, each path visits a sequence of rooms; these are simple paths with no repeated (room) vertices. Such paths are not necessarily shortest paths. The low-level sequence is then the shortest panorama path, constrained by the room sequence. Given the set of all such paths across all houses, the fourth desiderata is satisfied by iteratively selecting the path that most improves coverage while maintaining a bias against shortest paths.

Movement in the simulator is based on a navigation graph. Vertices correspond to 360-degree panoramic images, captured at approximately 2.2m intervals throughout 90 indoor environments. Edges are navigable links between panoramas. Chang et al. (2017) also partition panoramas via human-defined room annotations.

Let PP be an undirected graph of interconnected panoramas, with vertices pi∈V(P)p_{i}\in\mathcal{V}(P) and edges (pi,pj)∈E(P)(p_{i},p_{j}){\in}\mathcal{E}(P). Let ARA_{R} be a set of disjoint room annotations; each room ri∈ARr_{i}{\in}A_{R} is a non-overlapping subset of panoramas ri⊆V(P)r_{i}\subseteq\mathcal{V}(P), as shown in Figure 2(a). We abbreviate (p1,⋯ ,pm)(p_{1},\cdots,p_{m}) as p1:mp_{1:m}.

We create RR, an undirected room graph with vertices V(R)={⋃C(P[ri])∣ri∈AR}\mathcal{V}(R)=\{\bigcup\mathcal{C}(P[r_{i}])\mid r_{i}{\in}A_{R}\}. P[ri]P[r_{i}] is the subgraph of PP induced by room annotation rir_{i} and C\mathcal{C} returns a graph’s connected components. Simply put, each vertex in RR encompasses a subgraph of PP. An edge (ri,rj)∈E(R)(r_{i},r_{j})\in\mathcal{E}(R) exists if the subgraph of PP induced by V(ri)∪V(rj)\mathcal{V}(r_{i})\cup\mathcal{V}(r_{j}) is connected.

Path Generation

We generate the set of all simple paths in RR that traverse at most 5 rooms and two building levels. Let rpi∈V(R)r_{p_{i}}{\in}\mathcal{V}(R) be the room containing panorama pip_{i}. As shown in Figure 2(b), for each room path r1:nr_{1:n}, we construct a directed graph P[r1:n]P[r_{1:n}] in which an edge (pi,pj)(p_{i},p_{j}) exists if rpi=rpjr_{p_{i}}{=}r_{p_{j}} (pip_{i} and pjp_{j} are in the same room) or (rpi,rpj)(r_{p_{i}},r_{p_{j}}) is an edge in the room path. Given P[r1:n]P[r_{1:n}], we sample the start p1p_{1} and goal pmp_{m} uniformly from r1r_{1} and rnr_{n}, respectively. The full panorama path p1:mp_{1:m} is then the shortest path between p1p_{1} and pmp_{m} in P[r1:n]P[r_{1:n}].

Room size varies greatly, so this approach produces high path length variance. It also satisfies naturalness because people tend to ground instructions at the room level (e.g., Exit through the carved wooden door on the other side of the room). We find such paths easy to describe even with as many as 20 edges. Finally, these paths can approach their goal indirectly, as exemplified in Figure 2(b).

Greedy Selection for Coverage

The final path dataset DD is constructed by repeatedly selecting a panorama path p1:mp_{1:m} from all sampled paths (without replacement) until a desired size is reached. After selecting kk paths, let O(pi,Dk){\cal O}(p_{i},D_{k}) be the number of occurrences of panorama pip_{i} in the paths in DkD_{k}. At step k+1k+1, we select the path with the minimum value for d(p1,pm)L(p1:m)+1m∑pi∈p1:mO(pi,Dk)\frac{d(p_{1},p_{m})}{L(p_{1:m})}+\frac{1}{m}\sum_{p_{i}\in p_{1:m}}{\cal O}(p_{i},D_{k}), where LL is path length in PP and d(p1,pm)d(p_{1},p_{m}) is the shortest path distance between p1p_{1} and pmp_{m} in PP. The first term prefers non-shortest paths while the second encourages selection of paths that cover panoramas with low coverage in DkD_{k}. This selection step is also subject to a maximum path length of 40m, and a maximum of 500 paths per building environment.

Path Statistics

In total, we sample 16522 paths, which are split: 11089 train, 1232 val-seen (train environments), 1517 val-unseen (val environments), and 2684 test, following the same environment splits as Matterport3D and R2R. Compared to R2R, RxR paths are longer, spanning 8 edges and 14.9m on average, vs. 5 edges and 9.4m in R2R. More importantly, as shown in Figure 3, RxR paths exhibit much greater variation in length while also achieving more uniform coverage of the panoramas (and edges). Furthermore unlike R2R, 44.5% of RxR paths are not the shortest path from the start to the goal location. RxR paths are on average 27.4% longer than the shortest path.

Data Collection and Metrics

We immerse annotators in our own web-based version of the Matterport3D simulator using the panoramic images and the navigation graph. Compared to Anderson et al. (2018b), our annotation tool has additional capabilities including speech collection, virtual pose tracking, and time-alignment between transcript and pose. Figure 4 gives an example instruction with accompanying Guide and Follower pose traces. Here, we describe our collection process, analysis of the data, path evaluation metrics and simple baselines.

Like R2R, our simulator has camera controls allowing continuous heading and elevation changes and movement between panoramas. Guides look around and move to explore a provided path and attempt to create an instruction others can follow. R2R’s Guides create written instructions. In contrast, RxR’s Guides speak and the tool logs their entire virtual camera pose sequence. We use a 640 ×\times 480 pixel viewing canvas and a camera vertical field of view of 75 degrees. This process is inspired by Localized Narratives Pont-Tuset et al. (2020), an image captioning dataset for which annotators move mouse pointers around images while talking about them.

As with Localized Narratives, RxR Guides transcribe their own recordings; this produces high quality text versions of the instructions. To align text and pose traces, we generate a time-stamped transcription using automatic speech recognition.https://cloud.google.com/speech-to-text The transcription and ASR output are aligned using dynamic time warping. The output of the Guide task is an audio file, a tokenized, timestamped, manually-transcribed instruction, and a pose trace (a series of timestamped 6-DOF camera poses). On average, Guide task annotations (including both steps, performed back-to-back) take 458 seconds.

For each language (English, Hindi and Telugu) we annotate 14K paths with three instructions each. In the English dataset, each path gets one US English instruction and two Indian English instructions. Of the 14K paths per language, 12.8K paths are common across all three languages, and 1.2K paths in each language are unique (equaling 16.5K paths in total). The fact that most paths are annotated 9 times (3 per language) creates interesting opportunities to study aligned instructions across languages. Unique paths add variety and coverage.

Follower Task

As Followers, annotators begin at the start of an unknown path and try to follow the Guide’s instruction. They observe the environment and navigate in the simulator as the Guide’s audio plays. They can pause, rewind and skip forward in the instruction. If they believe they have reached the the end of the path, or give up, they indicate they are done and rate the instruction’s clarity and their confidence in their own navigation. On average, Follower tasks take 132 seconds.

The Follower tasks objectively validate the quality of Guide instructions based on whether the Follower can succeed (i.e., reaching within 3m of the last panorama in the path). If the Follower doesn’t succeed, the Guide instruction is paired with a second Follower. If the second Follower succeeds, the first Follower annotation is discarded and replaced. If the second Follower also fails, then the path is re-enqueued to generate another Guide and Follower annotation. The most successful of the three resulting Guide-Follower pairs is selected for inclusion in RxR and the others are discarded.

In addition to validating data quality, the Follower task also trains annotators to be better Guides—following bad instructions often helps one see how to produce better instructions. Most importantly, we collect the pose trace of the Follower as they execute the instruction. This provides an alternative path with dense grounding that we can compare to the Guide’s pose trace and use as an additional training signal.

Dataset Analysis

Table 2 provides summary statistics for RxR. The average words per instruction (using whitespace tokenization) is 78 vs R2R’s 29. US English instructions are the longest on average. We attribute this to conventions developed by each annotator pool rather than language specific properties. On average Guide tasks take much longer than Follower tasks (458 vs. 132 seconds). Most of the Guide’s time is spent transcribing audio (Guide audio recordings average 60 seconds).

Following a similar analysis as Chen et al. (2019), Table 3 gives examples and statistics for linguistic phenomena, based on manual analysis of instructions for 25 paths. All RxR subsets produce a higher rate of entity references compared to R2R. This is consistent with the extra challenge of RxR’s paths and our annotation guidance that instructions should help followers stay on the path as well as reach the goal. Doing so requires more extensive use of objects in the environment. RxR’s higher rate of both coreference and sequencing indicates that its instructions have greater discourse coherence and connection than R2R’s. RxR also includes a far higher proportion of allocentric relations and state verification compared to R2R, and matches Touchdown (navigation instructions). Hindi contains less coreference, sequencing, and temporal conditions than the other languages. That said, it is not clear how much the differences within RxR exhibited in Table 3 can be attributed to language, dialect, annotator pools, or other factors.

Figure 5 (top) illustrates the close alignment between instruction progress (measured in words) and path progress (measured in steps). Figure 5 (bottom) indicates that both Guide and Tourist annotators orient themselves by looking around at the first panoramic viewpoint, after which they maintain a narrower focus. On average, Guides / Tourists observe 43% / 44% of the available spherical visual signal at the first viewpoint, and 27% / 28% at subsequent viewpoints. These findings stand in contrast to standard VLN agents that routinely consume the entire panoramic image and attend over the entire instruction sequence at each step. Inputs that the Guide / Tourist have not observed cannot influence their utterances / actions, so pose traces offer rich opportunities for agent supervision.

Evaluation

We use the following standard evaluation metrics (with arrows indicating improvement): Path Length (PL), Navigation Error (NE ↓\downarrow) Success Rate (SR ↑\uparrow), Success weighted by inverse Path Length (SPL ↑\uparrow), Normalized Dynamic Time Warping (NDTW ↑\uparrow), and Success weighted by normalized Dynamic Time Warping (SDTW ↑\uparrow). See Anderson et al. (2018a) and Ilharco et al. (2019) for discussion of VLN metrics. Since RxR was designed to include paths that approach their goal indirectly, we focus primarily on NDTW and SDTW which explicitly capture path adherence. See Table 4 for a comparison of the performance of several simple baselines on R2R and RxR. Each simple baseline requires a stopping criteria; we choose to stop after NN steps where NN is the average number of steps in the train set paths (5 in R2R and 8 in RxR). Consistent with our motivation to reduce biases in paths, these simple baselines show that going straight is far less effective in RxR than R2R.

Experiments

For the image features we use an EfficientNet-B4 CNN Tan and Le (2019). Following Parekh et al. (2020), we pretrain the CNN in an image-text dual encoder setting using the Conceptual Captions dataset Sharma et al. (2018). In preliminary experiments, we found that pretraining the CNN in this way gave noticeable improvements over the same CNN pretrained for image classification on ImageNet Russakovsky et al. (2015).

Grounding Supervision

To incorporate spatiotemporal groundings into agent training, for each Guide path (G-path) and Follower path (F-path) we convert the corresponding pose trace into: (1) a sequence of text masks bt∈{0,1}lb_{t}\in\{0,1\}^{l} indicating which words in instruction xx the Guide spoke / Follower heard at or prior to step tt, and (2) a sequence of visual masks Mt∈{0,1}h×wM_{t}\in\{0,1\}^{h\times w} indicating which pixels were observed in the panoramic image at tt (like Figure 5 bottom). We then project and max-pool MtM_{t} to a vector mask mt∈{0,1}km_{t}\in\{0,1\}^{k} aligning to the agent’s visual input features vtv_{t}. Zeros in btb_{t} and mtm_{t} indicate irrelevant textual and visual inputs that were not observed by the annotators, and are therefore not related to their utterances and actions.

To help prevent the agent from overfitting to superficial correlations in the training data, we use btb_{t} and mtm_{t} to supervise the normalized textual and visual attention weights in the model. Specifically, during training whenever the agent is on the gold path we apply a cross-entropy loss to the visual attention weights given by L(z,mt)=log⁡∑i=1kexp⁡(zi)−log⁡∑i=1kmt,iexp⁡(zi)\mathcal{L}(z,m_{t})=\log\sum_{i=1}^{k}\exp(z_{i})-\log\sum_{i=1}^{k}m_{t,i}\exp(z_{i}), where zz is the vector of unnormalized logits determining attention weights via a softmax. This loss forces the attention weights on irrelevant input features towards zero. The textual version is analogous.

Implementation Details

Agents are implemented in VALAN Lansing et al. (2019), a distributed reinforcement learning framework designed for VLN. We use a mix of supervised learning and policy gradients. Each minibatch is constructed from 50% behavioural cloning roll-outs (following the gold paths while minimizing cross-entropy loss), and 50% policy gradient rollouts with reward (following paths sampled from the agent’s policy). As in Ilharco et al. (2019), the reward at each step is the incremental difference in NDTW, plus a linear function of navigation error after stopping. All agents are trained with Adam Kingma and Ba (2014) to convergence (100K iterations with batch size of 32 and initial learning rate of 1e-4).

Monolingual Results

Table 5 provides results on the val-unseen split for several training settings, as well as human performance from Follower annotations. We report en-US and en-IN results together as en. Experiments 1–3 compare agents trained (1) only on G-paths, (2) only on F-paths, and (3) on both. In contrast to algorithmically generated G-paths, each F-path reflects a grounded human interpretation of an instruction, which may deviate from the G-path because multiple correct interpretations are possible (e.g., Figure 4). For training, we do not differentiate F-paths from G-paths, and each instruction-path pair is treated as an independent example. Experiment (3) shows that including both G- and F-paths in training benefits every metric. Given the overall positive impact of F-paths, we use both path types in our further experiments.

Multilinguality

For experiment (4) in Table 5, we train a single multilingual agent on all three languages simultaneously. While the multilingual agent sees substantially more instructions than each monolingual agent, performance is worse across all metrics. This is consistent with results in multilingual machine translation (MT) and automatic speech recognition (ASR) where adding more languages can also lead to degradation for high-resource languages Aharoni et al. (2019); Pratap et al. (2020). Experiment (5) takes this one step further by obtaining translations from every instruction into the two other languages (e.g., en →\rightarrow hi, te) using a MT service.https://cloud.google.com/translate These translations are included in the RxR data release. Including these translations hurts performance for all languages. The fact that most G-paths are shared across languages may limit the value of automatic cross-translations. Notwithstanding the higher performance of the monolingual approaches, in the remaining experiments we focus on multilingual agents for greater scalability.

Spatiotemporal Grounding Supervision

Table 5 experiment (6) incorporates a loss for spatiotemporal grounding over visual attention which gives mixed results on val-unseen (better on NDTW, NE and worse on success-based metrics) compared to (4). Applying the same approach to textual attention did not improve performance. However, we stress that this is only a preliminary investigation. Using human demonstrations to supervise visual groundings is an active area of research Wu and Mooney (2019); Selvaraju et al. (2019). As one of the first large-scale spatially-temporally aligned language datasets, RxR offers new opportunities to extend this work from images to environments.

Multitask and Transfer Learning

Table 6 reports the performance of the multilingual agent under multitask and transfer learning settings. For simplicity, the R2R model (exp. 7) is trained without data augmentation from model-generated instructions Fried et al. (2018); Tan et al. (2019) and with hyperparameters tuned for RxR. Under these settings, the multitask model (exp. 8) performs best on both datasets. However, transfer learning performance (RxR →\rightarrow R2R and vice-versa) is much weaker than the in-domain results. Although RxR and R2R share the same underlying environments, we note that RxR →\rightarrow R2R cannot exploit R2R’s path bias, and for R2R →\rightarrow RxR, the much longer paths and richer language are out-of-domain.

Unimodal Ablations

Table 7 reports the performance of the multilingual agent under settings in which we ablate either the vision or the language inputs during both training and evaluation, as advocated by Thomason et al. (2019a). The multimodal agent (4) outperforms both the language-only agent (9) and the vision-only agent (10), indicating that both modalities contribute to performance. The language-only agent performs better than the vision-only agent. This is likely because even without vision, parts of the instructions such as ‘turn left‘ and ‘go upstairs‘ still have meaning in the context of the navigation graph. In contrast, the vision-only model has no access to the instructions, without which the paths are highly random.

Test Set

RxR includes a heldout test set, which we divide into two splits: test-standard and test-challenge. These splits will remain sequestered to support a public leaderboard and a challenge so the community can track progress and evaluate agents fairly. Table 8 provides test-standard performance of the mono and multilingual agents using Guide and Follower paths, along with random and human Follower scores. While the learned agent is clearly much better than a random agent, there is a great deal of headroom to reach human performance.

Conclusion

RxR represents a significant evolution in the scale, scope and possibilities for research on embodied language agents in simulated, photo-realistic 3D environments. RxR’s paths better ensure that language itself will play a fundamental role in better agents. Evaluating on three typologically diverse languages will help the community avoid overfitting to a particular language and dataset.

We have only begun to explore the possibilities opened up by pose traces. Whereas others have retro-actively refined R2R’s annotations to get alignments between sub-instructions and panorama sequences Hong et al. (2020), RxR provides word-level alignments to specific pixels in panoramas. This is obtained as a by-product of significant work on the annotation tooling itself and designing the process to be more natural for Guides. Finally, every instruction is accompanied by a Follower demonstration, including a perspective camera pose trace that shows a play-by-play account of how a human interpreted the instructions given their position and progress through the path. We have shown that these can help with agent training, but they also open up new possibilities for studying grounded language pragmatics in the VLN setting, and for training VLN agents with perspective cameras – either in the graph-based simulator or by lifting RxR into a continuous simulator Krantz et al. (2020).

Acknowledgments

We thank Sneha Kudugunta for analyzing the Telugu annotations, and the Google Data Compute team, especially Igor Karpov, Ashwin Kakarla and Christina Liu, for their tooling and annotation support for this project, Austin Waters and Su Wang for help with image features, and Daphne Luong for executive support for the data collection.

References

Appendix A Supplementary Material

In total, 247 annotators contributed to RxR, with 97 based in the USA and the remainder based in India and contributing to the Indian English, Hindi and Telugu annotations. The annotators were paid hourly wages that are competitive for their locale. They have standard rights as contractors. They were fluent in the language they were tasked with.

We ensure that a Guide does not annotate the same path twice. As Followers, annotators do not follow their own Guide instructions. Furthermore, we have provided annotators multiple forms of feedback as they complete tasks. After a round of pilot instructions were collected, we provided detailed analysis of common patterns that produced poor instructions and clear guidelines for producing better instructions. Annotators provided UI suggestions and interesting corner cases to us that allowed us to refine the simulator and annotation process before kicking off the full annotation process. Throughout the process, annotators have had access to a dashboard that shows them their success rate as both Guide and Follower. We indicated that their success as Guide and a Follower should be above 80%. Any annotator whose success is lower is either given further training or is taken off the task.

Unfortunately, we cannot release the audio instructions yet due to the impact of COVID-19: our annotators had to complete the tasks from home, so we need to review all recordings for safety and privacy. We hope to include the audio in a future release.

Appendix B DATASHEET: ROOM-ACROSS-ROOM (RxR)

This document is based on Datasheets for Datasets Gebru et al. (2020). Please see the most updated version here.

COMPOSITION

COLLECTION

PREPROCESSING / CLEANING / LABELING

USES

DISTRIBUTION

MAINTENANCE