iGibson 2.0: Object-Centric Simulation for Robot Learning of Everyday Household Tasks

Chengshu Li, Fei Xia, Roberto Martín-Martín, Michael Lingelbach, Sanjana Srivastava, Bokui Shen, Kent Vainio, Cem Gokmen, Gokul Dharan, Tanish Jain, Andrey Kurenkov, C. Karen Liu, Hyowon Gweon, Jiajun Wu, Li Fei-Fei, Silvio Savarese

Introduction

Recent years, we have seen the emergence of many simulation environments for robotics and embodied AI research . The main function of these simulators is to compute the motion resulting from the physical contact-interaction between (rigid) bodies, as this is the main process that allows robots to navigate and manipulate the environment. This kinodynamic simulation is sufficient for pick-and-place and rearrangement tasks ; however, as the field advances, researchers are taking on more diverse and complex tasks that cannot be performed in these simulators, e.g., household activities that involve changing the temperature of objects, their dirtiness and wetness levels. There is a need for new simulation environments that can maintain and update new types of object states to broaden the diversity of activities that can be studied.

We present iGibson 2.0, an open-source extension of the kinodynamic simulator iGibson with several novel functionalities. First and foremost, iGibson 2.0 maintains and updates new extended physical states resulting from the approximation of additional physical processes. These states include not only kinodynamics (pose, motion, forces), but also object’s temperature, wetness level, cleanliness level, toggled and sliced state (functional states). These states have a direct effect on the appearance of the objects, captured by the high-quality virtual sensor signals rendered by the simulator.

Second, iGibson 2.0 provides a set of logical predicates that can be evaluated with a single object (e.g. Cooked) or a pair of objects (e.g. InsideOf). These logical predicates discriminate the continuous physical state maintained by the simulator into semantically meaningful logical states (e.g. Cooked is True if the temperature is above a certain threshold). Complementary to the discriminative functions, iGibson 2.0 implements generative functions that sample valid simulated physical states based on logical states. Scene initialization can then be described as a set of logical states that the simulator can translate into valid physical instances. This enables faster prototyping and specification of scenes in iGibson 2.0, facilitating the training of embodied AI agents in diverse instances of the same tasks. We demonstrate the potential of this generative functionality with a new dataset of home scenes densely populated with small objects. We generate this new dataset by applying a set of hand-designed semantic-logic rules to the original scenes of iGibson 1.0.

Third, to facilitate the development of new embodied AI solutions to new tasks in these new scenes, iGibson 2.0 includes a new virtual reality interface (VR) compatible with the two main commercially available VR systems. All states are logged during execution and can be replayed deterministically in the simulator, enabling the generation a posteriori of additional virtual sensor signals or visualizations of the interactions and the development of imitation learning solutions.

We evaluate the new functionalities of iGibson 2.0 on six novel tasks for embodied AI agents, and apply state-of-the-art robot learning algorithms to solve them. These tasks were not possible before in iGibson or in alternative simulation environments. Additionally, we evaluate the use of the new iGibson 2.0 VR interface to collect human demonstrations to train an imitation learning policy for bimanual operations. While the previous version of iGibson and other simulators provide interfaces to control an agent with a keyboard and/or a mouse, these interfaces are insufficient for bimanual manipulation.

In summary, iGibson 2.0 presents the following contributions:

A set of new physical properties, e.g. temperature, wetness and cleanliness level, maintained and updated by the simulator; and a set of unary and binary logical predicates that map simulated states to a logical state that have a direct connection to semantics and language,

A set of generative functions associated with the logical predicates to sample valid simulated states from a given logical state, and a new rule-based mechanism exploiting these functions to populate the iGibson scenes with small objects placed at semantically meaningful locations to increase realism,

A novel virtual reality interface that allows humans to collect demonstrations for robot learning,

We hope that iGibson 2.0 will open new avenues of research and development in embodied AI, enabling solutions to household activities that have been under-explored before.

Related Work

Simulation environments with (mostly) kinodynamic simulation: In the last years, the robotics and AI communities have presented several impressive simulation environments and benchmarks: iGibson , Habitat AI , AI2Thor (and variants) , ThreeDWorld , Sapien , Robosuite , VirtualHome , RLBench , MetaWorld , and more. They are based in physics engines such as (py)bullet , MuJoCo , and Nvidia Physx , combined with rendering capabilities and usually enriched with a dataset of objects and/or scenes to use to develop embodied AI solutions. While these simulators have fueled research with new possibilities for training, testing and developing robotic solutions, they have skewed, with few exceptions, the exploration towards activities related to what they can simulate accurately: changes in kinematic states of (rigid or flexible) objects, i.e. Rearrangement tasks . However, many everyday activities require the simulation of other physical states that can be modified by the agents, like the temperature of objects and their level of wetness or cleanliness. Recent simulators have attempted realistic simulation of fluids and flexible materials or extended rigid body simulators to approximate the dynamics of soft materials . Compared with these simulators, iGibson 2.0 provides a simple but effective mechanism to simulate the temperature of objects that change based on proximity to heat sources. It also simulates fluids through a system of droplets that can be absorbed by objects and change their wetness level. Despite simpler than the accurate simulation of heat transfer, fluids and soft materials used by a few other simulators, iGibson 2.0 object-centric solution leads to realistic robotic behavior and motion in tasks involving changing temperature, handling liquids, or soaking objects.

Simulation environments with object-centric representation: In robotics, some simulators have adopted an object-centric representation with extended physical states, e.g. AI2Thor and VirtualHome . Both are based on Unity and share a common set of functionalities. Actions in these simulators are predefined, discrete and symbolic, and characterized by preconditions (“what conditions need to be fulfilled for this action to be executable?”) and postconditions (“what conditions will change after the execution of this action?”), similar to actions in the planning domain definition language (PDDL ) or STRIPS but with the additional link to visual rendering of the predefined action execution and outcome. In AI2Thor and VirtualHome, some of the actions change the temperature of an object between two or more discrete values pre- and post-execution of an action, e.g. raw and cooked. This is fundamentally different to our approach in iGibson 2.0: instead of maintaining only a symbolic state (raw/cooked), we provide a simple simulation of the underlying physical process (e.g. heat transfer) leading to continuously varying values of temperature and other extended states. These states are then mapped into a symbolic representation through predicates (see Sec. 4). This provides a new level of detail in the execution of actions, where the agent can and should control the specific value of the object’s extended states (temperature, wetness, cleanliness) to achieve a task, leading to more complex activities and more realistic execution. In the Appendix, we include a detailed comparison between iGibson 2.0 and other simulation environments in Table A.7.

Simulation environments with virtual reality interfaces: Researchers have used virtual reality interfaces before to develop robotic solutions with real robots . VR has also been used in simulation environments to collect demos. VRKitchen collected demos for five cooking tasks in simulation. While realistic looking, activities in VRKitchen are performed with primitive actions similar to the pre-condition/post-condition system of AI2Thor and VirtualHome, falling short of realistic motion. More physically realistic are the VR interactions with UnrealROX/RobotriX and ThreeDWorld , based on Unreal and Nvidia Physx . Our interface also enables realistic manipulation of objects, with additional features such as gaze tracking and assistive grasping to bridge the differences between simulation and the real-world.

Extended Physical States for Simulation of Everyday Household Tasks

To perform household tasks, an agent needs to change objects’ states beyond their poses. iGibson 2.0 extends objects with five additional states: temperature, TT, wetness level, ww, cleanliness level (dustiness level, dd, or stain level, ss), toggled state, TS, and sliced state, SS. While some of these states could be different for different parts of an object, in iGibson 2.0 we simplify their simulation and adopt an object-centric representation: the simulator maintains a single value of each extended state for every simulated object (rigid, flexible, or articulated). This simplification is sufficient to simulate realistically household tasks such as cooking or cleaning. We assume that the extended properties are latent: agents are not able to observe them directly. Therefore, iGibson 2.0 implements a mechanism to change objects’ appearance based on their latent extended states (see Fig. 2), so that visually-guided agents can infer the latent states from sensor signals.

We further impose in iGibson 2.0 that every simulated object should be an instance of an existing object category in WordNet . This semantic structure allows us to associate characteristics to all instances of the same category . For example, we further simplify the simulation of extended states by annotating what extended states each category need. Not all object categories need all five extended states (e.g., the temperature is not necessary/relevant for non-food categories for most tasks of interest). The extended states required by each object category are determined by a crowdsourced annotation procedure in the WordNet hierarchy.

In the following, we explain the details of the five extended states and the way they are updated in iGibson 2.0. For a full list of object states (kinematics and extended), see Table A.4.

iGibson 2.0 maintains also a historical value of the maximum temperature that each object has reached in the past, Tomax=max⁡TotT_{o}^{\text{max}}=\max T_{o}^{t} for t∈[0,…,tnow]t\in[0,\ldots,t_{\text{now}}]. This value dictates the appearance of an object: if the object reached cooking or burning temperature in the past, it will look cooked or burned, even if its current temperature is low. Fig. 2(a) depicts the temperature system in action.

Wetness Level: Similar to temperature, iGibson 2.0 maintains the level of wetness for each object that can get soaked. This level corresponds to the number of droplets that have been absorbed by the object. In iGibson 2.0, the system of droplets approximates liquid/fluid simulation. Specifically, droplets are small particles of liquid that are created in droplet sources (e.g. faucets), destroyed by droplet sinks (e.g. sinks), and absorbed by soakable objects (e.g. towels). They can also be contained in receptacles (e.g. cups) and poured later, leading to realistic behavior for the simulation of several household activities involving liquids, illustrated in Fig. 2(b).

Cleanliness – Dustiness and Stain Level: A common task for robots in homes and offices is to clean dirt. This dirt commonly appears in the form of dust or stains. In iGibson 2.0, the main difference between dust and stains is the way they get cleaned: while dust can be cleaned with a dry cleaning tool like a cleaning cloth, stains can only be cleaned with a soaked cleaning tool like a scrubber. To clean a particle of dirt (dust or stain), the right part of a cleaning tool should get in physical contact with the particle. Once a dirt particle is cleaned, it disappears from objects’ surface.

In iGibson 2.0, objects can be initialized with visible dust or stain particles on its surface. The number of particles at initialization corresponds to a 100% level of dustiness, dd, or stains, ss, as we assume that dust/stain particles cannot be generated after initialization. As particles are cleaned, the level decreases proportionally to the number of particles removed, reaching a level of 0% dustiness, dd, or stains when the object is completely clean (no particles left). This extended state allows simulating multiple cleaning tasks in our simulator: the agent needs to exhibit a behavior (motion, use of tools) similar to the one necessary in the real world. Fig. 3(a) depicts an example of the cleanliness level simulation in iGibson 2.0, for both dust and stain particles.

Toggled State: Some object categories in iGibson 2.0 can be toggled on and off. iGibson 2.0 maintains and updates an internal binary functional state for those objects. The functional state can affect the appearance of an object, but also activate/deactivate other processes, e.g. heating food inside a microwave requires the microwave to be toggled on. To toggle the object, a certain area needs to be touched. For object models of categories that can be toggled on/off, we annotate a TogglingLink, an additional virtual fixed link that needs to be touched by the agent to change the toggled state. Fig. 3(b) depicts an example of an object that can change its toggled state, an oven.

Sliced State: Many cooking activities require the agent to slice objects, e.g. food items. Slicing is challenging in simulation environments where objects are assumed to have a fixed (rigid or flexible) 3D structure of vertices and faces. To approximate the effect of slicing, iGibson 2.0 maintains and updates a sliced state in instances of object categories that are annotated as sliceable. When the sliced state transitions to True, the simulator replaces the whole object with two halves. The two halves will be placed at the same location and inherit the extended states from the whole object (e.g. temperature). The transition is not reversible: the object will remain sliced for all upcoming simulated time steps. Objects can only be sliced into two halves, with no further division. The sliced state changes when the object is contacted with enough force (over a slicing force threshold for the object) by a slicing tool, e.g. a knife. Objects of these categories are annotated as SlicingTool. If an object is a slicing tool, it will undergo a second annotation process to obtain a new virtual fixed link that acts as SlicingLink, the part of the slicing tool that can slice an object, e.g., the sharp edge of a knife. Fig. 3(c) depicts an example of a peach being sliced by a knife. For more information about update rules for all extended physical states, please refer to see Sec. A.2.

Logical Representation of Physical States

The new extended object states from iGibson 2.0 are sufficient to simulate a new set of household activities in indoor environments. However, there is a semantic gap between the extended states (e.g. temperature or wetness level) and the natural description of activities in a household setup (e.g. cooking apples). To bridge this gap, we define a set of functions that map the extended object states to logical states for single objects and pairs of objects. The logical states are semantically grounded on common natural language representing properties such as cooked or dusty.

The list of logical predicates covers kinematic states between pair of objects (InsideOf, OnTopOf, NextTo, InContactWith, Under, OnFloor), states related to the internal degrees of freedom of articulated objects (Open), states based on the object temperature (Cooked, Burnt, Frozen), wetness and cleanliness level (Soaked, Dusty, Stained), and functional state (ToggledOn, Sliced). A complete list of the logic predicates with detailed explanation is included in Table A.1. They allow iGibson 2.0 to map a physical simulated state into an corresponding logical state.

Logical predicates map multiple physically simulated states to the same logical state, e.g., all relative poses between to objects that correspond to being onTop. In addition to this discriminative role, we include functionalities in iGibson 2.0 to use logical predicates in a generative manner, to describe initial states symbolically that can be used to initialize the simulator. iGibson 2.0 includes a sampling mechanism to create valid instances of tasks described with logical predicates. This mechanism facilitates the creation of multiple semantically meaningful initial states, without the laborious process of manually annotating the initial distributions per scene.

The process of sampling valid object states is different depending on the nature of the logical predicate. For predicates based on objects’ extended states such as Cooked, Frozen or ToggledOn, we just sample values of the extended states that satisfy the predicate’s requirements, e.g., a temperature below the annotated freezing point of an object to fullfil the predicate Frozen. Particles for Dusty and Stained are sampled on the surface of an object following a pseudo-random procedure. Generating initial states to fulfill kinematic predicates such as OnTopOf or Inside is a more complex procedure as the underlying physical state (the object pose) must lead to a stationary state (e.g., not falling) that does not cause penetration between objects. Each kinematic predicate is implemented differently, combining mechanisms that include ray-casting and analytical methods to verify the validity of sampled poses. For example, to sample a state that satisfies Inside(A, B), we implement a procedure that generates 6D poses inside of object B and we evaluate that 1) a bounding box of the the size of object A does not penetrate object B and rays cast from A intersect object B from evaluated by casting rays from the the internal box. For more details on the sampling mechanism, see Sec. A.2.

2 iGibson 2.0 Scenes with Realistic Object Distribution Created by Generative System

One common issue that limits the realism of indoor scenes in simulation is that they are less densely populated than those in the real world. Creating highly populated simulated houses is usually a laborious process that requires manually selecting and placing models of small objects in different locations. Thanks to the generative system in iGibson 2.0, this process can be extremely simplified. The users only need to specify a list of logical predicates that represent a realistic distribution of objects in a house.

We provide as part of iGibson 2.0 a set of semantic rules to generate more populated scenes, and a new version of the iGibson original 15 fully interactive scenes, populated with additional small objects as a result of the application of the rules. Given an indoor scene with multiple types of rooms (kitchen, bathroom, …) that are populated with furniture containers and appliances (fridge, cabinet, …), the semantic rules define the probabilities for object instances of diverse categories to be sampled in certain container type in a given room type, e.g. p(p(InsideOf(Beer, Fridge)∧\landInsideOf(Fridge, Kitchen))). To generate the more densely populated versions of the 15 scenes, we first collect a large number of small 3D object models created by artists and we annotate them with semantic categories (e.g. cereal, apple, bowl) and realistic dimensions. Then, we apply in a random sequence the logic predicates using the generative system explained in Sec. 4.1, increasing the number of objects in the scene by over 100. The result is depicted in Fig. 4.

Virtual Reality Interface

iGibson 2.0’s new functionalities enable modeling new household activities and generating multiple instances in more densely populated scenes. To facilitate research in these new, complex tasks, iGibson 2.0 includes a novel virtual reality (VR) interface compatible with major commercially available VR headset through OpenVR . One of the goals is to allow researchers to collect human demonstrations, and use them to develop new solutions via imitation (see Sec. 6).

iGibson 2.0’s VR interface creates an immersive experience: humans embody an avatar in the same scene and for the same task as the AI agents. The virtual reality avatar (see Fig. 1) is composed of a main body, two hands and a head. The human controls the motion of the head and the two hands via the VR headset and hand controllers with an optional, additional tracker for control of the main body. Humans receive stereo images as generated from the point of view of the head of the virtual avatar, at at least 30 fps (up to 90 fps) using the PBR rendering functionalities of iGibson .

Grasping in VR: Grasping in the real-world, while natural to adult humans, is a complex experience that proves difficult to reproduce with virtual reality controllers. Our empirical observations revealed that while using solely physical grasping, user dexterity was significantly impaired relative to the real-world resulting in unnatural behavior when manipulating in VR. To provide a more natural grasping experience, we implement an assistive grasp (AG) mechanism that enables an additional constraint between the palm and a target object after the user passes a grasp threshold (50% actuation) and provided the object is in contact with the hand, between the fingers and the palm. This facilitates grasping of small objects, and prevents object slippage. To not render grasping artificially trivial, the AG connection can break if the constraint is violated beyond a set threshold, such as while lifting heavy objects or during intense acceleration, encouraging natural task execution that leverages careful motions and bimanual manipulation. Please refer to Sec. A.1 for additional details about AG.

Navigating in VR: Navigation of the avatar is controlled by the locomotion of the human. However, the VR space is much smaller than the typical size of iGibson 2.0 scenes. To navigate between rooms, we configured a touchpad in the hand controller that humans can use to translate the avatar.

Evaluation

In our evaluation, we test the new functionalities of iGibson 2.0 explained above and that sets it apart from other existing environments. First, we create a set of tasks that showcase the new extended states (see Fig. 5(a)) as their modification is required to achieve the tasks. In our experiments, we make use of our discriminative and generative logical engine to detect task completion and create multiple instances of each task for training. Then, we use our novel VR interface to collect human demonstrations to train an imitation learning policy for a bimanual task.

RL experiments: With the simplified grasping mechanism, the agents trained with SAC for both the bimanual humanoid and the Fetch embodiment achieve 100%100\% success rate for Grasping Book, Soaking Towel, Cleaning Stained Shelf, Cooking Meat tasks. For Slicing Fruit, the agents achieve only 15%15\% and 0%0\% for bimanual humanoid and Fetch robot respectively, due to the increased accuracy necessary to align the knife blade with the fruit. The agent using bimanual humanoid achieves 0%0\% success in the Bimanual Pick and Place task because of the difficulties of controlling and coordinating both hands. Additionally, we evaluate the performance with the Fetch robot in more realistic conditions, without any simplification for grasping, and observe a significant drop in performance, achieving 25%25\% success rate for 2 tasks and 0%0\% for the other 3 (see Fig. A.2 in the Appendix). This indicates that successful grasping for diverse objects is a significant challenge in these manipulation tasks. To test generalization, we conducted an ablation study in which we train policies with three different levels of variability in Soaking Towel –no variations, different poses, different objects and poses– and evaluate them on an unseen setup (an unseen object with randomized initial pose). The policies achieve success rate of 19%19\%, 79%79\%, and 87%87\%, respectively. This study shows that it’s essential to train with diverse object models and initial states to obtain robust policies, and the generative system of iGibson 2.0 facilitates it by specifying a few logical states that describe the initial scene. The attached video shows the policies performing the tasks trained using iGibson 2.0’s new extended physical states and logical predicates that help generating task instances and discriminating their completion.

Conclusion

We presented iGibson 2.0, an open-source simulation environment for household tasks with several key novel features: 1) an object-centric representation and extended object states (e.g. temperature, wetness and cleanliness level), 2) logical predicates mapping simulation states to logical states, and generative mechanism to create simulated worlds based on a given logical description, and 3) a virtual reality interface to easily collect human demonstrations for imitation. We demonstrate in multiple experiments the new avenues for research enabled by iGibson 2.0. We hope iGibson 2.0 becomes a useful tool for the community, and facilitates the development of novel embodied AI solutions.

This work is in part supported by ARMY MURI grant W911NF-15-1-0479 and Stanford Institute for Human-Centered AI (SUHAI). S. S. is supported by the National Science Foundation Graduate Research Fellowship Program (NSF GRFP) and Department of Navy award (N00014-16-1-2127) issued by the Office of Naval Research. S. S. and C. L. are supported by SUHAI Award # 202521. R. M-M is supported by SAIL TRI Center – Award # S-2018-28-Savarese-Robot-Learn. F. X. and B. S. are supported by Qualcomm Innovation Fellowship. M. L. is supported by Regina Casper Stanford Graduate Fellowship.

References

Appendix for iGibson 2.0: Object-Centric Simulation for Robot Learning of Everyday Household Tasks

In this section, we provide additional information about the implementation of our virtual reality (VR) interface in iGibson 2.0.

Haptic feedback: To approximate the real-world experience, it is important to create haptic feedback to human subjects when interacting with the scene. To that end, in our VR interface collisions of the body and the hands trigger haptic vibrations in the controllers. The body will trigger a strong vibration in both controllers facilitating navigation in the scene. The hands also generate low-strength vibration when they are in contact with an object, to notify users that they are in contact with a virtual object. These mechanisms create a multi-modal stream to the humans (vision and haptics) that help them interact more dexterously and realistically.

Assistive Grasping: Creating realistic, robust and dexterous grasping in virtual reality is challenging. Grasping objects in real-world involves generating multiple frictional contact points and surfaces between the hand and the object: simulating physically this complex process in a realistic manner is non-trivial. Additionally, real-world grasping involves rich multimodal signals that include tactile and haptic information, that is not available in common virtual reality interfaces. To compensate for these differences and generate a natural experience in simulation, we implement an assistive grasping mechanism.

In our VR interface, grasping is performed by pressing the left or right trigger, a one degree of freedom (DoF) actuation. This single DoF is mapped to a closing motion of the avatar’s hand where all fingers move synchronously. If the trigger is pressed more than 50% of its range, we activate the assistive grasping (AG) mode. The AG mode facilitates grasping by creating an additional fixed or point-to-point joint between the hand and the movable objects inside the hand. We use three criteria to decide what object inside the hand should be assisted for grasping: First, the object has to be inside the hand. We evaluate this we a ray casting mechanism. Rays are shot between the following two sets of points: (thumb tip, thumb middle, palm middle, palm base) and (4 non-thumb finger tips). These 16 rays return a list of all objects that are in the hand. Second, the object has to be close to the palm. This is defined as the distance between the center of mass of the palm link and the object is the smallest among all objects in hand. And third, the object needs to be in contact with the hand and the hand has to be applying force on it. We found that, defined with these three criteria, the AG mechanism is realistic as humans can grasp objects reliably with motions that are close to the ones used in real-world for the same objects. While a user is grasping an object using AG, collision of the hand with that object is also disabled, to avoid any recurring collisions or simulation instability.

Logging and replaying demonstrations: All information of the runs can be logged and replay deterministically (same output for the same actions). The logged information includes kinematics and extended iGibson 2.0 states, and VR agent actions. We believe the logged information acquired with the iGibson 2.0 VR interface will facilitate research: the information can be analyzed to understand human strategies, or used with modern robot learning techniques, e.g. with imitation learning, to train embodied AI solutions.

A.2 Extended Object States, Logical Predicates and Generative System in iGibson 2.0

In this section, we provide additional information on iGibson 2.0’s new extended physical states, logical predicate system that map physical states to logical states, and the generative system that sample valid simulated physical states based on logical states.

Extended states associated to object categories: In iGibson 2.0, not all extended states need to be maintained for instances of all object categories, e.g., most objects cannot be Sliced and only food objects could be Cooked. We assume that any object instance added to iGibson 2.0 belongs to a category annotated with properties. The properties indicate what extended states should be updated for object instances of that category. The exhaustive list of all possible object category properties is shown in Table A.2. Some properties will require additional annotation for each object model to estimate logical states.

Object model annotations: Each model needs to be annotated with additional physical and semantic information to simulate correctly interactions and their associated logical states. The exhaustive list of all possible object model properties are included in Table A.3. Some properties directly come from the 3D assets, such as Shape and KinematicStructure. We compute Weight based on query results from Amazon Product API, and compute CenterOfMass and MomentOfInertia accordingly for each link based on the assumption of uniform density. Note that if an object category is annotated with a certain property, e.g., stove is annotated as HeatSourceSink, we need to annotate HeatSourceSinkLink, a virtual (non-colliding) fixed link that heats or cools objects, for all object models of stove. For SlicingTool and CleaningTool, we additionally annotate SlicingToolLink and CleaningToolLink, which are colliding fixed links that can slice objects (the blade of a knife) and remove dirt particles from objects (the bottom of a vacuum), respectively.

Updating object state: During simulation, our simulator maintains and updates not only the kinematic states of the objects, such as Pose and InContactObjs, using our underlying physics engine Bullet , but also the non-kinematic states, such as Temperature and WetnessLevel with custom rules. These update rules are explained in Sec. 3 and summarized in Table A.4.

Logical predicates as discriminative functions: As explained in Sec. 4, we define a set of discriminative functions that map the extended physical states to logical states that are semantically grounded on natural language, such as Cooked and Sliced. These logical states can used for symbolic planning and checking intermediate success for sub-tasks for reinforcement learning. The details of the discriminative functions of all the logical predicates can be found in Table A.1.

Logical predicates as generative functions: We also define a set of sampling functions that can generate valid physical states that satisfy the given logical states. For example, if the initial conditions of the task require a book placed OnTopOf a table or a shelf being Stained, our system can automatically sample concrete physical states that satisfy the requirements: sampling a random position on the table and place the book there, and sample stain particles on random locations on the shelf (see Fig. 5(a)). The details of the generative functions of all the logical predicates can be found in Table A.5. Please also refer to our supplementary video for more details.

Sampling extended states based on the given logical predicates is relatively simple, e.g., sampling a temperature that corresponds to an object being Frozen. However, sampling object poses to fulfill the given kinematic predicates is more involved as it requires sampling values in the Special Euclidean group SE(3) with additional constraints such as placing objects in stable configurations and not causing penetrations between objects. In the following, we describe our algorithm to sample valid poses based on kinematic logical predicates.

Say we are sampling a valid pose for an object o1o_{1} to be OnTopOf object o2o_{2}. First, we query the set of stable orientations allowed for object o1o_{1}. We assume these orientations are provided per object model, e.g., for a book the orientations to place the book on its cover and last page or upright. Each stable orientation is linked to an axis-aligned bounding box with an associated bounding-box base area. The next step would be to find areas on the surface of object o2o_{2} that can hold the bounding-box area and that are flat, unobstructed, and accessible. To find these areas of o2o_{2} surface we use a ray-casting mechanism conditioned on the specific kinematic logical predicate. For example, for our case of OnTopOf we will generate rays starting immediately above o2o_{2} by sampling points from the top face of its axis-aligned bounding box, and marching downwards in the vertical direction. The points where the rays intersect o2o_{2} surface will be used to define planes where we can attempt to sample object o1o_{1} if they fulfill some criteria such as providing stable support. We can repeat the procedure for different stable orientations of o1o_{1}. Other logical predicates use a similar generative procedure but with variations in the ray-tracing step. For example, for InsideOf, we start our rays at different points inside the o2o_{2} bounding box rather than above it. In addition, for particle-based states such as Dusty and Stained, we additionally allow casting rays in horizontal directions.

A.3 Experimental Setup and Additional Results

In this section we provide the experimental setup for the reinforcement learning and imitation learning experiments described in Sec. 6.

The observation space include 128×128128\times 128 RGB-D images from the onboard sensor on the agent’s head, and proprioceptive information (hand poses in agent’s local frame, and a fraction indicating how much each hand is closed). The action space is 6-dimensional representing the desired linear and angular velocities of the right hand, where the rest of the agent is stationary. For grasping, we adopt the “sticky mitten” simplification from other works : we create a fixed constraint between the hand and the object as soon as they get in contact.

The agent receives a one-time success reward if it satisfies the single predicate (e.g. Cooked(meat)). Additionally, we provide distance-based reward shaping for each experiment to encourage the hand to approach activity-relevant objects, e.g. encourage the hand to approach the meat and the meat to approach to stove. Finally, for the Cleaning Stained Shelf task, we provide partial progress reward for each stain particle that has been cleaned. The episode terminates if the agent achieves success or times out (200 timesteps, or equivalently 20 seconds). We train for 10K episodes, evaluate on the same setups, and report the results. The training reward curves can be found in Fig. A.1.

We use Soft Actor-Critic for training. The policy network has two encoders for RGB-D images and proprioceptive information. With RGB-D images as input, we use a 3-layer convolutional neural network to encode the image into a 256 dimensional vector. The proprioceptive information is encoded into a 256 dimensional vector with an MLP. The features are concatenated and pass through additional MLP layers to generate the action.

We also conducted the same RL experiments with a Fetch robot. The observation space include 128×128128\times 128 RGB-D images from the onboard sensor on the agent’s head, and proprioceptive information (the end effector pose in agent’s local frame, joint configurations, and whether the end effector is currently grasping something). The action space is 6-dimensional representing the desired linear and angular velocities of the end effector, where the rest of the agent is stationary. We experimented with both the “sticky mitten” grasping simplification (the same as Bimanual Humanoid) and without such simplification. For the later setup, the Fetch robot has to rely on the friction between the gripper fingers and the objects to grasp them with realistic physics simulation. Its action space also includes one additional DoF for closing the gripper. With this later setup, we hope to minimize the sim2real gap as much as possible. Due to the additional complexity of grasping, we add one more reward shaping terms to encourage the gripper to grasp the task-relevant object and penalize the agent for dropping it. We use different reward scaling. The termination conditions remain the same. We also use the same policy network architecture and training schema as before. The training reward curves can be found in Fig. A.2.

: The observation space includes the ground truth poses of the task-relevant objects, and proprioceptive information, the same as the RL setup. In the Bimanual Pick and Place experiment, task relevant objects include the cauldron, the table and the agent itself. The agent can control both of its hands with the desired linear and angular velocities, which result in 12 degree of freedom. The hand closing action is not learned, but replayed from the human demonstrations.

We collected 6500+6500+ state-action pairs from 3030 human demos and used behavior cloning to predict the action based on the state. The policy network has two MLP encoders for proprioceptive information and ground truth object poses. The features are concatenated and pass through additional MLP layers to generate the action.

The network is trained until validation loss plateaus, and evaluated on a test set of demonstrations. As discussed in the main paper, the entire task is long horizon (>300 steps) and the policy diverges due to covariate shift . We then evaluate if the policy can successfully perform the task if we initialize the simulator a few seconds before task completion of a successful human demo. The success rate with respect to different policy starting time can be found in Fig. 5(b). We show a successful sequence after rewinding 2 seconds in Fig. A.3.

A.4 Performance Benchmark of iGibson 2.0

To evaluate whether iGibson 2.0 can be used in computationally expensive embodied AI research, we benchmarked the performance (simulation time) and compared with the previous version. The benchmark setup is the same as in Shen et al. , which considered an “idle” setup, in which we place a robot (a TurtleBot model) in the scene and run the physics simulation and extended physical state simulation loop. The benchmark runs on 15 scenes, and statistics are collected. The agent applies zero actions and stays still. We use action time step of ta=130st_{a}=\frac{1}{30}\text{s} and physics time step of ts=1120st_{s}=\frac{1}{120}\text{s} to be consistent with Shen et al. . Both settings are benchmarked on a computer with Intel 5930k CPU and Nvidia GTX 1080 Ti GPU, in a single process setting, rendering 512×512512\times 512 RGB-D images.

The simulator speed is shown in Table A.6. Although we added many extended physical states, we still achieved a 25%25\% increase in average performance compared with iGibson 1.0 . In iGibson 2.0, the main source of speed up with respect to the previous version of iGibson is obtained from better usage of the object sleeping mechanism, and lazy update of object poses in the renderer. This allows us to simulate much larger scenes with many more objects with extended physical states tracked, and as a result, more diverse everyday household activities.

A.5 Feature Comparisons of Simulators

In Table A.7, we provide a detailed comparison across multiple simulation environments. The table is adapted from Table I of . We include more recent simulation environments as columns and more feature comparisons as rows.

A.6 Limitations and Future Work

Although iGibson 2.0 has made several significant contributions towards simulating complex, everyday household tasks for robot learning, it is not without limitation. First of all, iGibson 2.0 doesn’t support soft bodies / flexible material in a scalable way at the moment, due to the limitation of our underlying physics engine. This prevents us from simulating tasks like folding laundry and making bed in large, interactive scenes. Also, iGibson 2.0 doesn’t support accurate human behavior modeling (other than goal-oriented navigation), and thus prevent us from simulating tasks that are inherently rich in human-robot interaction (e.g. elderly care). With the recent advancement of physics engines, and human behavior modeling and motion synthesis, we plan to overcome these limitations in the future. In addition, we also plan to support a more diverse set of extended object states (e.g. Filled, Hung, Assembled, etc) as well as bi-directional transitions for some of our existing states (e.g. Soaked and Stained/Dusty), which can unlock even more household tasks. Finally, we plan to transfer mobile manipulation policies trained in iGibson 2.0 to the real world.