StructFormer: Learning Spatial Structure for Language-Guided Semantic Rearrangement of Novel Objects

Weiyu Liu, Chris Paxton, Tucker Hermans, Dieter Fox

I Introduction

Organizing objects into complex, semantically meaningful arrangements has many real-world applications, such as setting the table, organizing books, and loading the dishwasher. Indeed, the organizing task received the second most internet searches on the comprehensive chore list from suggesting its practical importance for future domestic robots. More generally rearrangement was recently proposed as a benchmark for embodied AI . However, to be broadly deployed these robots should ideally receive instructions in a way that doesn’t overly burden the operator assigning the task. In this work, we focus on the problem of semantic rearrangement, where a robot must move a set of novel objects to form a spatial structure that satisfies a high-level language instruction such as put the mugs in a row, build a circle, and set the table as shown in Fig. 1.

Successful rearrangement manipulation in unstructured environments requires a robot to jointly reason about geometric and semantic properties of these novel objects and the objects’ physical interactions . Realistic settings additionally present objects not required to accomplish the task requiring the robot to reason about additional relations between structures of interest and distractor objects. Finally, the robot needs an encoding of its goal compatible with efficient planning, while also being easily provided by the operator. Language provides an obvious input modality for untrained users to specify goals ; however, it brings with it challenges in inferring the implied object configuration and generating representations interpretable by the robot.

Previously, researchers have worked on modeling, interpreting, and grounding relations between scene elements . Such reasoning can enable robots to move a query object with respect to an anchor object to satisfy a desired spatial relations . More complex spatial structures from language instructions have be generated in a blocks world environment by chaining these binary relation changing manipulations . Others have leveraged spatio-semantic relations between pairs of objects as an abstraction for planning . Alternatively, robots can perform multi-object manipulation by directly learning to transform object groups without associated semantic reasoning .

In contrast to the pairwise reasoning of prior work, our model explicitly reasons about multi-object semantic relations given a structured language instruction to specify the goal. We treat the semantic rearrangement planning problem as a sequential prediction task using a novel transformer-based architecture named StructFormer. We use transformer encoders to build a contextualized representation of abstract concepts expressed in language instructions as well as semantic and geometric multi-object relations. This enables StructFormer to reason over a variety of objects of varying number without access to object models. The encoder can directly predict what objects to move and also provide a context for an autoregressive transformer decoder to predict where the objects should go.

To train and validate our approach, we procedurally generate a dataset of four different structures—circles, lines, towers, and table settings—using models from 335 real-world objects. We tested our method both on novel objects and arrangements in simulation, and on a physical robot. Our experiments show that StructFormer’s ability to directly reason over multiple objects enables higher manipulation success compared to a model that only reasons about pairwise relations between objects. For further videos, gifs, and supplemental material please visit our website https://sites.google.com/view/structformer.

II Related Work

Grounding spatial relations defines prior work on spatial understanding focusing on modeling relations between pairs of objects, such as in, on, and left of . These spatial relations have been used to specify action goals (e.g., moving the cup to the right of the bowl) . More complex structures such as towers or letters can also be created by chaining these action goals, as shown in , while grounds such spatial abstractions. Other works examine classifying spatial structures . Besides facilitating communicating language goals, spatial relations have also been used as an abstraction for planning . A recent work segments video demonstrations of object manipulations into sequences of spatial relations, demonstrating the benefit of this abstract representation for imitation learning . Different from existing methods, we model spatial relations between multiple objects, extending from binary spatial relations to multi-object relations.

Visual question answering and understanding reasons about spatial and compositional structures of visual elements–an active research area in the vision community . The CLEVR visual question answering (VQA) dataset and subsequent extensions have provided useful benchmarks for developing systems that perform object-centric reasoning. More recent efforts leverage RAVEN’s Progressive Matrices to test models’ abilities to discover structure among objects . The NLVR dataset also helps develop methods to explicitly reason about sets of objects . In contrast to passively parsing objects in the scene, our method actively manipulates objects to achieve desired structures.

3D structure synthesis methods can be categorized into holistic generative models , which directly generate the whole 3D structure, and structure-aware models , which leverages substructures (e.g., object parts or furniture in a room). Our work is most closely related to the latter. These models decompose generations into synthesizing object parts and creating spatial arrangements of parts. Instead of leveraging structure prior of object parts (e.g., a table has four legs on each corner), we condition spatial arrangements of objects based on high-level language instruction.

III Transformer Preliminary

Transformers were proposed in for modeling sequential data. At the heart of the Transformer architecture is the scaled dot-product attention function, which allows elements in a sequence to attend to other elements. Specifically, an attention function takes in an input sequence {x1,...,xn}\{x_{1},...,x_{n}\} and outputs a sequence of the same length {y1,...,yn}\{y_{1},...,y_{n}\}. Each input xix_{i} is linearly projected to a query qiq_{i}, key kik_{i}, and value viv_{i}. The output yiy_{i} is computed as a weighted sum of the values, where the weight assigned to each value is based on the compatibility of the query with the corresponding key. The function is computed on a set of queries simultaneous with matrix multiplication.

IV StructFormer for Object Rearrangement

Given a single view of a scene ss containing objects {o1,...,oN}\{o_{1},...,o_{N}\} and a structured language instruction ll containing word tokens {w1,...,wM}\{w_{1},...,w_{M}\}, our goal is to rearrange a subset of the objects {o1,...,oNq}\{o_{1},...,o_{N_{q}}\}, which we call query objects, to reach a goal scene s∗s^{*}. The rearranged scene should satisfy the spatial and semantic constraints encoded in the instruction ll while being physically valid. We assume we are given a partial-view point cloud of the scene ZZ with segment labels for points to identify the different objects. Given this point cloud ZZ and the language instruction ll, the robot must find pose offsets {δ1,...,δNq}\{\delta_{1},...,\delta_{N_{q}}\}, which can transform the query objects {o1,...,oNq}\{o_{1},...,o_{N_{q}}\} in the initial scene to new poses that satisfy the goal arrangement implied by ll.

Our model consists of an object selection network and a pose generator network, as shown in Fig. 2. Both networks jointly reason about the object point clouds and language instructions using transformers . For a given scene, latent representations of words and objects are used to construct the input sequence {w1,..,wN,o1,...,oN}\{w_{1},..,w_{N},o_{1},...,o_{N}\} for the two networks. The object selection network uses a transformer encoder to predict output sequence {κ1,...κN}\{\kappa_{1},...\kappa_{N}\} in a single forward inference, where κi\kappa_{i} is a binary variable indicating if the robot should rearrange object oio_{i}. The pose generator network uses a transformer decoder to autoregressively generate a sequence of pose offsets {δ1,...,δNq}\{\delta_{1},...,\delta_{N_{q}}\}, which we achieve by feeding in representations of the selected objects {o1,...,oNq}={oi∣κi=1}i=1N\{o_{1},...,o_{N_{q}}\}=\{o_{i}|\kappa_{i}=1\}_{i=1}^{N} as targets for decoding. Below, we discuss how we build our latent representations of objects and language instructions. Then we describe the two transformer networks in detail. Finally, we discuss how to train our system and use it for rearrangement planning.

IV-B Object Selection Network

The object selection network kΦ({ei},{ci})→{κi}k_{\Phi}(\{e_{i}\},\{c_{i}\})\rightarrow\{\kappa_{i}\} predicts objects that need to be rearranged based on the language instruction. We use a transformer encoder to perform relational inference over all objects in a scene and the given language instruction (e.g., identifying “objects that are smaller than a pan”). The object and sentence encoders encode the words {wi}\{w_{i}\} and objects {oi}\{o_{i}\} to create the input sequence {c1,..,cM,e1,...eN}\{c_{1},..,c_{M},e_{1},...e_{N}\} to the transformer encoder. We feed encoder’s output at each object’s position into a linear layer to predict {κ1,...κN}\{\kappa_{1},...\kappa_{N}\}. Formally, the object selection transformer models the distribution

IV-C Language Conditioned Pose Generator

We learn a generative distribution πΩ({ei},{ci})→{δi}\pi_{\Omega}(\{e_{i}\},\{c_{i}\})\rightarrow\{\delta_{i}\} over possible pose offsets for objects that might satisfy the language instruction and are physically valid. We use a transformer encoder-decoder model. The encoder has the same architecture as the object selection network and encodes the sequence {c1,..,cM,e1,...eN}\{c_{1},..,c_{M},e_{1},...e_{N}\} to build a contextualized representation of the language instruction and objects in the scene, including objects that need to be moved and objects that will remain stationary. The decoder autoregressively predicts each object’s pose offset, conditioning on the global context and the pose offsets of previously predicted objects. Formally, the decoder takes as input the sequence {e0,[δ0;e1],[δ1;e2],...,[δNq−1;eNq]}\{e_{0},[\delta_{0};e_{1}],[\delta_{1};e_{2}],...,[\delta_{N_{q}-1};e_{N_{q}}]\} and predicts {δ0,δ1,...,δNq}\{\delta_{0},\delta_{1},...,\delta_{N_{q}}\}. We ensure the input object poses are not used by the decoder by shifting the input poses by one position and using a causal attention mask. We model the following distribution with the encoder-decoder model

The network obtains its stochasticity by using a dropout layer with probability p∈p\in during training and inference.

The order of the query objects in the input sequence is predefined for each spatial structure. For example, the rearranged objects will build a circle structure clockwise. We find empirically that imposing an order on objects and using a virtual frame help create precise spatial structures.

IV-D Inference and Training

During inference, we select objects to rearrange based on prediction from the object selection network kΦk_{\Phi}. We sample a batch of BB rearrangements from our pose generator πΩ\pi_{\Omega} for the query objects. For each sample, we autoregressively predict the target pose of the structure frame and each object, conditioned on the previous predictions in the sequence.

We train the object selection network kΦk_{\Phi} and pose generator πΩ\pi_{\Omega} with data from rearrangement sequences. The object selection network is trained on initial scenes and groundtruth query objects using a binary cross entropy loss. The generator is trained with an L2L2-loss minimizing the distance between groundtruth and predicted placement poses.

V Data Generation

We introduce a dataset containing more than 100,000 rearrangement sequences. We pair each rearrangement with a high level language instruction specifying the target spatial rearrangement for a set of objects. The language instructions involve many different semantic and geometric properties for both grounding objects and specifying the spatial structures, as shown in Table I. We procedurally generate stable and collision-free object arrangements in the PyBullet physics simulator and render with the photo-realistic image render NVISII .

With the goal of generalization in mind, we adopt 335 everyday household objects from the acronym dataset . Figure 3 shows the diversity of objects used from 35 distinct classes.

We generate a language-conditioned rearrangement sequence in three steps: (1) sampling a referring expression for query objects, (2) arranging query objects into a physically realistic spatial structure, and (3) creating an action sequence with time reversal. We discuss each step in detail below.

We functionally generate referring expressions for sets of objects that need to be rearranged. A referring expression can indicate the query objects explicitly with a discrete feature (e.g., metal objects) or by relating to an anchor object, which in turn can be described by one to three discrete features. Using anchor objects allows us to create referring expressions that require relational reasoning of abstract semantic properties of objects (e.g., objects that have the same material as the blue bottle) and continuous geometric properties (e.g., objects that are shorter than the glass cup). After sampling a referring expression, we add query objects, an optional anchor object, and additional distractor objects, which do not match the referring expression, to the scene. Since table settings involve specific objects, we do no create referring expressions for this structure type.

We rearrange query objects into physically correct instances of one of the four defined spatial structures according to different geometric parameters (e.g., radius, size, position, rotation) in PyBullet. By discretizing the parameters of the structure according to a pre-specified vocabulary, we generate a sentence describing the structure (e.g., place query objects into a large circle on the top right of the table).

Finally, we move objects out of the structure to random, collision-free poses in the scene. We obtain an action sequence that rearranges objects from random poses into the goal configuration described by the language instruction by reversing the random action and associated image sequences. We use NVISII to render color and depth images and instance segmentation masks of all objects in the sequence.

VI Experiments

In this section, we provide rigorous experimental validation of our approach. We first evaluate the individual components of StructFormer on the collected dataset. Following this we show planning performance for our entire system. We then provide results for using our system to generate and execute rearrangement plans on a physical robot with real-world objects and sensing.

We split the dataset into 80% training, 10% validation, and 10% testing. The test data consist of new target spatial structures and novel object combinations.

We compare our pose generator to the following baselines.

No Encoder: This variant of our pose generator does not use the transformer encoder to extract global context for decoding. When predicting the pose offset for an object, it only has information about the language instruction and previously predicted objects. This baseline is similar to the previous transformer models used in floor plan generation and clip-art generation work , where the placements of objects are less spatially constrained.

No Structure: This variant of our pose generator directly predicts 6-DoF pose offset of each object in the world frame without predicting and using the virtual structure frame.

As shown in Table II, our model outperforms all baselines at predicting precise placements of objects for four different structures. Comparison with No Encoder validates that the transformer encoder in our pose generator is crucial for precise rearrangements. The larger errors produced by the No Struct baseline indicate that using a virtual structure frame helps anchor placements of objects for better structure generation. This is analogous to leveraging one object as the spatial anchor for placing another object when manipulating pairwise spatial relations . Additionally, we hypothesize that separately predicting the placements of the structures and objects also helps deal with spatial ambiguities embedded in language instructions (e.g., arrange a circle in the middle of the table). Finally, we highlight the performance difference between Binary and our model, which confirms that modeling multi-object spatial relations is beneficial not only for creating complex spatial structures but also for generating simpler structures such as lines and towers, which can be described by binary relations. The ”Struct” results in Table II show that we have comparable accuracy in locating the target structure relative to anchor objects in the scene with or with out the use of the encoder.

Besides better quantitative performance achieved by our method, we also see qualitative improvement. In Fig. 4, we visualize predicted rearrangements by transforming point clouds of objects. We highlight that the No Encoder and Binary baselines are inadequate at producing circular structures because these two methods are not able to condition placement of an object based on future objects (e.g., how many more objects to be placed and what are their dimensions). We also note that No Struct fails to predict precise placement for a large number of objects (O) and is inconsistent at producing accurate alignments of objects (P).

VI-A2 Object Selection Network

We test our object selection network on novel combinations of objects. Fig. 5 shows that our model can retrieve objects based on both direct references and relational references and can also reason about continuous-valued properties including height and volume. Fig. 5 further demonstrates that our model maintains high accuracy when given an increasing number of target objects. We visualize identifying objects to move in Fig. 6.

VI-B Evaluating Full System in Simulation

We evaluated our entire system in the simulation environment using 138 novel object models from 23 known object classes. We preserve physical interactions of objects while using ground-truth instance segmentation and omitting low-level control of the robot. We simulate object placement by dropping any moved object from 33 cm above the predicted target zz value. For each scene, we use the same procedure as our data generation to sample a referring expression and corresponding query, anchor, and distractor objects. After putting objects into the scene, we randomly select a structure and its parameters to create a high-level language instruction.

We first tested the combined system. In our experiments, the object selection network was able to identify all query objects given a referring expression in 95/156 (61%) of the tested scenes and identify 70% of the specified objects in 129/156 (83%) of the scenes. With the selected objects, the pose generator successfully rearranged objects into circles, lines, and towers in 40/61 (66%), 40/61 (65%) and 10/34 (29%) of the scenes. The table setting structure was not included because it does not use the same object referring expressions as the other three structures. The overall success rate of the whole pipeline on building spatial structures based on high-level language instructions was 58/156 (37%). Constructing tower structures was especially challenging because many tested objects have irregular shapes and inherently cannot be stacked (e.g., apples, teapots, candle stands). Another major failure mode was incompatible structure parameters and objects (e.g., arranging large pans into a small circle).

We also compared our pose generator to the Binary baseline using groundtruth object selections. As shown in Fig. 7, our model consistently outperformed Binary at building all four structures. Our model successfully built 58/112 (52%) more structures that requires modeling complex spatial relations (i.e., circles and table settings) and 14/95 (15%) structures that can be described by pairwise spatial relations (i.e., lines and towers). In scenes where both methods were successful, our method moved on average 0.17±0.380.17\pm 0.38 distractor objects in the scenes while Binary moved 0.21±0.410.21\pm 0.41. This result suggests that our method can more effectively reason about the relevance of the objects in the scene, a necessary ability for rearrangements in clutter. Figure 1 show our model can generate rearrangements of different structures and size at different parts of the table conditioned on the language instructions.

VI-C Physical Robot Experiments

We deploy our system on a Franka Panda Robot with an arm-mounted RGB-D camera to evaluate real-world object manipulation. We generate grasps using the method from and use RRT-connect for motion planning. We visualize successful rearrangements in Fig. 1.

Failures of the system are driven by oddly perceived objects. When a large portion of an object is occluded, the system is prone to place the object such that it intersects with other objects. Since our work focus on generating 3D structures, we do not use any sophisticated planning method. As a result, motion planning fails sometimes due to unreachable objects. However, the candidate goal scenes generated by our method could be combined with the vision-based planners in to find feasible motion plans.

VII Conclusions

We presented an learning-based approach for robot planning and manipulation of multi-objects semantic arrangements. Our method leverages a transformer architecture to generate plans as sequence output from an input scene point cloud and language command. Our results show the benefit of our specific architecture over alternative networks.

While our method outperforms the baselines, we leave open the problem of operating directly from natural language, instead using structured language to specify parameters of each structure. We also don’t currently address placement in clutter or finding the optimal order of rearranging actions. Instead, we always build in a predefined order. In the future, we will incorporate StructFormer into a full-fledged task-and-motion planner to examine solving rearrangement problems in a variety of environments, not just on tables.

References