Situated and Interactive Multimodal Conversations
Seungwhan Moon, Satwik Kottur, Paul A. Crook, Ankita De, Shivani Poddar, Theodore Levin, David Whitney, Daniel Difranco, Ahmad Beirami, Eunjoon Cho, Rajen Subba, Alborz Geramifard
Introduction
As virtual digital assistants become increasingly ubiquitous, they are expected to be embedded in the day-to-day life of users the same way a human assistant would. We thus envision that the next generation of virtual assistants will be able to process multimodal inputs and provide multimodal outputs beyond the traditional NLP stack. To this end, we present Situated Interactive MultiModal Conversations (SIMMC) tasks and datasets as a starting point in this new research direction. Specifically, SIMMC focuses on task-oriented dialogs that encompass a situated multimodal context, where situated implies that the user and assistant are continually co-observing the same context, and that context can be updated on each turn. We provide two new SIMMC datasets in the domain of interactive shopping, collected using the SIMMC Platform [Crook et al., 2019]: (1) Furniture and (2) Fashion. Moreover, we provide fine-grained annotations to allow for both end-to-end and component-level modelling. The annotation includes natural language understanding (NLU), multimodal-coreference, multimodal state tracking, assistant actions, natural language generation (NLG), and item appearance logs.
Fig. 1 illustrates an exemplary dialog from our SIMMC-Furniture dataset, where a user is interacting with an assistant with the goal of browsing for furniture. In our setting, the assistant can update the co-observed environment to create a new multimodal context based on the preceding dialog with the user, e.g., visually presenting recommended chairs in a virtual reality (VR) environment, or responding to the request “I like the brown one. Show me the back of it.” by executing the actions of focusing on, and rotating the indicated item. These actions change the shared multimodal context, which grounds the next part of the dialog. The example also highlights novel challenges such as multimodal action prediction (italics above) and multimodal coreference resolution (underlined elements above).
Novelty & Related Work
Tab. 1 presents main distinctions of SIMMC compared to the the existing multimodal dialog datasets.
Multimodal datasets for co-observed, real-world assistant. With the ultimate goal of laying the foundations for the real-world assistant scenarios, we assume a co-observed multimodal context between a user and an assistant. This shifts the primary focus onto the core problem of grounding conversations in the co-observed multimodal context. In contrast, the existing literature [Das et al., 2017, Kottur et al., 2019, De Vries et al., 2017, de Vries et al., 2018], drawing motivation from the Visual Question Answering [Antol et al., 2015], posits the roles of a primary and secondary observer, i.e., questioner and answerer, who do not co-observe the same multimodal context.
In addition, we study scenarios in which the situated multimodal context is dynamically updated, reflecting the agent actions. In our settings, agent actions can be enacted on both the object-level – changing the view of a specific object within a scene, and the scene-level – introducing a new scene or an image. While the dialog-based image retrieval tasks [Guo et al., 2018, Saha et al., 2018] and the visual navigation tasks [Thomason et al., 2019, de Vries et al., 2018] do comprise context updates, they are limited to the introduction of new visual scenes, e.g., new images or locations.
Focus on task-oriented dialogs. We frame the problem as a task-oriented, multimodal dialog system, with the aim of extending the capabilities of digital assistants to real-world multimodal settings. While one focus area of the dialog community is on task-oriented dialog, which has practical applicability consumer-facing virtual assistants [Henderson et al., 2014, Budzianowski et al., 2018, Eric et al., 2019, Rastogi et al., 2019, Chen et al., 2020], this form of dialog is often neglected in many existing multimodal ‘dialog’ datasets (both in terms of the task design and the annotations), where the primary focus lies in visual grounding of language. Our work aims to bring important challenges actively studied in the dialog community to the multimodal setting. Specifically, we study the multimodal extension of the traditional dialog state tracking (DST) and the assistant API prediction tasks, which have been the key focus of dialog literature [Wu et al., 2019, Gao et al., 2019, Chao and Lane, 2019].
Compared to the conventional task-oriented dialog datasets (e.g., MultiWoZ [Budzianowski et al., 2018]), the agent actions in SIMMC span across a diverse multimodal action space (e.g., rotate, search, add_to_cart). Our study thus shifts the focus of the visual dialog research from the token-level grounding of visual scenes to the task-level understanding of dialogs given multimodal context.
Semantic annotations for multimodal dialogs. Finally, we present a novel flexible schema for semantic annotations that we developed specifically for the natural multimodal conversations. The proposed SIMMC annotation schema allows for a more systematic and structural approach for visual grounding of conversations, which is essential for solving this challenging problem in the real-world scenarios. To the best of our knowledge, our dataset is the first among the related multimodal dialog corpora to provide fine-grained semantic annotations.
SIMMC Datasets
For SIMMC datasets, we focused on the shopping domain as it often induces rich multimodal interactions around browsing visually grounded items. As shown in Fig. 1, the setup consists of two human workers, a user and an assistant, conversing around a shopping scenario. The goal of the user is to interactively browse through an inventory of items while that of the assistant is to facilitate this conversation. In addition to having an interactive dialog, the assistant manipulates the co-observed environment to show off items from the shopping inventory. A conversational assistant model for the SIMMC datasets would need to (i) understand the user’s utterance using both the dialog history and the state of the environment – the latter provided as multimodal context, and (ii) produce a multimodal response to the user utterance, including updates to the co-observed environment to convey meaningful information as part of the user’s shopping experience. We provide two SIMMC datasets with slightly different setups and modalities. See Tab. 3 for overall statistics.
The SIMMC-Furniture (VR) Dataset captures a scenario where a user is interacting with an assistant whilst browsing for furniture, e.g., couch, or side table. Grounded in a VR environment [Unity Technologies, 2019] the assistant can manipulate items in the scene while engaging in conversation. We seed the conversation by presenting the user with either a high-level directive such as ‘Shop for a table’ or an image of a furniture item to shop for. The user is then connected randomly with a human assistant. The assistant can filter the catalog by attributes such as furniture category, price, color, and material, navigate through the filtered results and share their view with the user. As part of the dialog, the user can request to look closer at one of the options, or see other options. In response, the assistant can either zoom into an item, present an alternate view by rotating it, or look at the catalog description to answer further questions. To enable this, the environment is designed to transition between two states: (a) Carousel, which displays three filtered furniture items (top view, Fig. 1); and (b) Focused, which provides a zoomed in view of one item from the carousel view (bottom view, Fig. 1). The conversation continues for 6–12 turns until the user considers that they have reached a successful outcome. How do we define success? Tab. 10 shows example dialogs.
The SIMMC-Fashion (Image) Dataset represents user interactions with an assistant to obtain recommendations for clothing items, e.g., jacket, dress. Conversations are grounded in real-world images that simulate a shopping scene from a user’s point-of-view (POV). At the start of each dialog the user is presented with a randomly selected ‘seed’ image from the catalog to emulate (visually) that they are in the middle of shopping, as well as a sequence of synthetic memories of ‘previously viewed items’. In addition to the user’s context, the assistant has access to a broader catalog that allows for information lookup and item recommendation. We ask the user to browse and explore options by asking the assistant for recommendations based on the shared attributes, preferences, as referred from visual scenes, memories, and assistant-recommend items. The conversation continues for 6–10 turns until the user is assumed to be given a successful recommendation. Please refer to Tab. 11 for example dialogs.
For both datasets, the ground-truth of which items appear in each view is logged and included in the multimodal context. This allows the problem of computer vision to be sidestepped and focus on semantically combining the modalities. The datasets were collected through the SIMMC Platform [Crook et al., 2019], an extension to ParlAI [Miller et al., 2017] for multimodal conversational data collection and system evaluation. Note that even though we focus on English in this work, our data collection framework is language-agnostic and can be easily extended to other languages.
SIMMC-Furniture has dialogs with an average of rounds (or turn pairs) leading to a total of about utterances. Similarly, SIMMC-Fashion consists of dialogs, each around rounds on average, totaling utterances. In addition to these sets, we also collect a smaller, audio-based SIMMC-Furniture dataset ( dialogs) where the dialog exchanges are aural as opposed to written text.
In Fig. 2, we visualize: (a) Distribution of rounds. Dialogs in SIMMC-Furniture range from (shorter ones are omitted from the dataset) to a maximum of rounds, with of the dialogs containing – rounds (Fig. 2(a)). Dialogs in SIMMC-Fashion range from – rounds at an average of rounds per dialog, as shown in Fig. 2(c). We hope that this widespread range will help train models that can handle diverse conversations of varied lengths. (b) Distribution of utterance lengths. For both user and assistant, we tokenize their utterances and plot the distribution in Fig. 2(b). For SIMMC-Furniture, the assistant utterances are slightly longer with higher variance at when compared to those from the user, at . A potential reason is that because the assistant has access to the catalog, it is expected to be more verbose while responding to description related queries (‘User: Tell me more about the brown table’). However, we do not observe a similar trend for SIMMC-Fashion where user and assistant turns average around tokens per utterance (Fig. 2(d)). (c) Catalog coverage. Recall that both SIMMC datasets contain conversations in a shopping scenario grounded in a catalog of furniture and fashion items respectively. SIMMC-Furniture builds on a catalog of items, where each dialog contains around shares of different views between the user and assistant, and each furniture item is shared in roughly dialogs. Similarly, SIMMC-Fashion contains items that appear in dialogs on average, thus providing a rich catalog context to support interesting multimodal dialogs.
SIMMC Dialog Annotations
Building a task-oriented multimodal conversational model introduces many new challenges, as it requires both action and item-level understanding of multimodal interactions. While most of the previous multimodal datasets provide surface-level annotations (e.g., utterance to multimodal action pairs), we believe it is critical to provide the semantic-level fine-grained annotations that ground the visual context, allowing for a more systematic and structural study for visual grounding of conversations. Towards this end, we develop a novel SIMMC ontology that captures the detailed multimodal interactions within dialog flows. Note that these dialog annotations are collected after the dialog data collection stage, with the help of professional linguists. In this section, we describe the proposed SIMMC ontology and the hierarchical labeling language centered around objects (Sec. 4.1 and 4.2), and the multimodal coreference schema that links the annotated language with the co-observed multimodal context (Sec. 4.3).
The SIMMC ontology provides common semantics for both the assistant and user utterances. The ontology is developed in the Resource Development Framework (RDF) and is an expansion of the Basic Formal Ontology [Arp et al., 2015]. It consists of four primary components:
Objects: A hierarchy of objects is defined in the ontology. This hierarchy is a rooted tree, with finer-grained objects at deeper levels. Sub-types are related to super-types via the isA relationship, e.g., sofa isA furniture. Fine-grained objects include user, dress, and sofa.
Activities: A hierarchy of activities are defined as a sub-graph of objects within the ontology. These represent activities the virtual assistant can take like get, refine, and add_to_cart.
Attributes: A given object has a list of attributes which relate that object to other objects, to primitive data types, or to enums. Finer-grained objects inherit the attributes of their parents. There are restrictions on the available types for both the domain and range of attributes. For example, a sofa can be related to a company via the brand attribute. A person can be related to an item of clothing via the attentionOn attribute. The takesArgument attribute relates Activities and the objects they act upon.
Dialog Acts: A hierarchy of dialog acts is also defined as a sub-graph of objects within the ontology. Dialog acts indicate the linguistically motivated purpose of the user or system’s utterance. They define the manner in which the system conveys information to the user and vice versa. Examples of dialog acts include: ask, inform, and prompt. Dialog acts are related to the activities that they act upon via the takesArgument attribute. Tab. 9 lists the activities and dialog acts used in our work.
2 SIMMC Labeling Language
From the SIMMC ontology, we derive a compositional, linearized, and interpretable labeling language for linguistic annotation, allowing for the representation of the natural language utterances as well-formed subgraphs of the ontology [Kollar et al., 2018]. The labeling language consists of intents and slots [Gupta et al., 2006]. Intents are taken to represents instances of the types they are composed of and take one of two forms: 1) dialog_act:activity:object or 2) dialog_act:activity:object.attribute. Only combinations of objects and attributes declared to be valid in the ontology are made available in the labeling language. Within these intents, slots further specify values for attributes of objects, activities, and attribute types. In the basic case, slots take the form of attributes of the intent-level objects and restrict those attributes. More complex cases include slot-in-slot nesting to restrict the type of the embedding slot, object-attribute combinations for type-shifting contexts, i.e., utterances in which an intent-level object is identical to the range of another object’s property, and a system of indexing to restrict objects introduced within the intent. Crucially, the labeling language is speaker agnostic. It makes no distinction in the parses of the user’s utterance versus those of the assistant.
A number of additional conventions are placed on the annotation task to ensure consistency and accuracy, which are detailed in Appendix B. See Tab. 10 and Tab. 11 in Appendix G for annotated dialog examples that show our SIMMC ontology in action for both our datasets.
3 SIMMC Coreference Annotations
Note that the proposed labeling language allows for the annotation of object types in a dialog, which may in turn refer to specific canonical listings from the underlying multimodal contexts. For example, given an annotated utterance “[da:request:get:chair Show me the back of it]”, the annotated object ‘chair’ (it) would refer to a specific catalog item, represented as a item id within the image metadata. To allow for structural grounding between the verbal and visual modalities in a shared catalog, we further annotate the mapping of object type mentions in the annotated utterance to the corresponding item id in the image metadata. The final SIMMC annotations thus capture the semantic relations of objects in multimodal contexts with their corresponding dialog annotations (activities, attributes and dialog acts), as outlined in the proposed SIMMC ontology (Sec. 4.1). We provide the detailed analysis of the datasets and the annotations in Appendix A.
SIMMC Tasks & Metrics
We define several offline evaluation tasks within the SIMMC framework to train conversational models on these new datasets using the fine-grained annotations that are provided. We first provide the general offline evaluation framework for defining SIMMC tasks (Sec. 5.1), and then present three major tasks that we focus on in this paper. These are primarily aimed at replicating human-assistant actions in order to enable rich and interactive shopping scenarios (Sec. 5.2).
Consider a generic SIMMC dialog that is rounds long, where and are the user and assistant utterances, is the domain-specific multimodal context, and is the action (API call) taken by the assistant at round , respectively. Formally, a task is defined as: At each round , given the current user utterance , the dialog history , multimodal context , predict the assistant action along with the free-form, natural language assistant response .
2 SIMMC Tasks
The proposed offline evaluation framework has a three-fold advantage: (a) It accurately represents the scenario encountered by a SIMMC model during deployment. In other words, models trained for the above task can be deployed to interact with humans to provide a situated, interactive, multimodal conversation. (b) Instead of evaluating the performance on the entire dialog, we evaluate models on a per-turn basis with the ground-truth history. This avoids taking the conversation out of the dataset and reduces the dependency on a user simulator, with the caveat of not encouraging the model to be able to learn multiple equally valid routes to satisfy the user’s request. (c) Finally, it facilitates us to define and evaluate several sub-tasks such as action prediction, response generation, and dialog state tracking, within SIMMC, which allows us to bootstrap from prior work on these sub-tasks.
Task 1: Structural API Call Prediction. This task involves predicting the assistant action as an API call along with the necessary arguments, using as inputs. For example, enquiring about an attribute value (e.g., price) for a shared furniture item is realized through a call to the SpecifyInfo API with the price argument. A comprehensive set of APIs for our SIMMC dataset is given in Tab. D. Apart from these APIs, we also include a None API call to catch situations without an underlying API call, e.g., a respond to ‘U: Can I see some tables?’ as ‘A: What color are you looking for?’ does not require any API calls. Action prediction is cast as a round-wise, multiclass classification problem over the set of APIs, measured using accuracy of predicting the action taken by the assistant during data collection. However, we note that there could be several actions that are equally valid in a given context. For instance, in response to ‘U: Show me some black couches.’, one could show black couches ‘A: Here are a few.’ or enquire further about specific preferences ‘A: What price range would you like to look at?’. Since accuracy does not account for the existence of multiple valid actions, we use perplexity (defined as the exponential of the mean log-likelihood) alongside accuracy. To also measure the correctness of the predicted action (API) arguments, we use attribute accuracy compared to the collected datasets.
Task 2: Response Generation. This task measures the relevance of the assistant response in the current turn. We evaluation in two ways, as a: (a) Conditional language modeling problem, where the closeness between the generated and ground-truth response is measured through using BLEU-4 score [Papineni et al., 2002], and, (b) Retrieval problem, where performance of the model to retrieve the ground-truth response from a pool of candidates (randomly chosen unique to each turn) is measured using standard retrieval metrics like recall@k (), mean rank, and mean reciprocal rank.
Task 3: Dialog State Tracking (DST). The dialog annotations collected using the flexible ontology enable us to study dialog state tracking (DST) in SIMMC, aside from providing additional supervision to train goal-driven agents. As mentioned in Sec. 4, the user and assistant utterances are accompanied with a hierarchy of dialog act labels and text spans for the corresponding slots or attributes, if any. The goal of DST is to systematically track the dialog acts and the associated slot pairs across multiple turns. We use the intent and slot prediction metrics (F1), following prior work in DST [Henderson et al., 2014].
Modeling for SIMMC Tasks
We now propose several models building on top of prior work and train them on the tasks formulated in Sec. 5 to benchmark the SIMMC dataset. We define two classes of models for the SIMMC tasks: (1) Assistant models, which aims at mimicking the assistant actions and responses (Task 1 & 2), and (2) User belief tracking model (Task 3) that output semantic parses of user utterances, agnostic of future assistant actions. Our principal Assistant model architecture is illustrated in Fig. 3, which is composed of four main components: Utterance and History Encoder, MultiModal Fusion, Action Predictor, and Response Generator. On the other hand, our user belief model builds upon the state-of-the-art DST models and extend them to accommodate for multimodal input. Inspired by [Hosseini-Asl et al., 2020], we adapt one such model and finetune a pretrained GPT-2 language model [Radford et al., 2019] to both action prediction and belief tracking.
Multimodal Fusion.
where Attention operator for a query over the key (of size ) and value is defined as
Action Predictor.
Response Generator.
Dialog State Tracking (DST).
In contrast to the assistant model that mimics the assistant actions and responses (Task 1 & 2), the user belief model aims to output semantic parses of user-side dialog (), strictly given the multimodal contexts available to the user, and agnostic of future assistant actions. Thus, . We utilize the state-of-the-art DST models from the recent literature: TRADE [Wu et al., 2019], which implements a pointer network that generates text spans for each slot, and an approach similar to SimpleTOD [Hosseini-Asl et al., 2020], which fine-tunes the pre-trained GPT-2 language model to output the user belief state labels [Radford et al., 2019].
In addition, we extend the SimpleTOD model to allow for multimodal input (SimpleTOD+MM). Specifically, we cast the belief tracking problem as a causal language modeling task, where belief labels and multimodal contexts are represented as additional tokens. A single training sequence can then be represented as the concatenation of input and target output , where both and are cast as string tokens of key value pairs. The language model is then fine-tuned to learn the joint probability with for all tokens in a sequence. At inference time, we provide the user input context as a seed for the language model, and parse the generated output to obtain the structural representation of user belief states.
We further extend the SimpleTOD+MM model to Tasks (1) and (2) by adding actions and assistant responses to the concatenation of input and target output, i.e., . At test time, we provide the input context plus oracle belief state and parse the generated response to extract the action and system utterance. We refer to this model as STOD++.
Experiments & Results
Our models are learned on randomly sampled train (), model hyperparameters chosen via early stopping using performance on dev (), and evaluation numbers reported on the unseen testdev (). In addition to the models described in Sec. 6, we consider two simple baselines that use TF-IDF features for utterance and history encoders for action prediction, and LSTM-based language model (LSTM) trained solely on assistant responses, and compare against them.
Dataset-specific Model Details.
We provide details around modeling multimodal context and encoding action (API call) output for each of the SIMMC datasets below.
Supervision.
We learn SIMMC models end-to-end by jointly minimizing the sum of the action prediction and the response generation losses, i.e., . To extract supervision for API call prediction (along with attributes), we utilize both the assistant (Wizard) interface activity during data collection (Sec. 3) and the fine-grained NLU annotations. Our implementation details are in Appendix E.
Results.
Tab. 4 summarizes the performance of SIMMC Assistant models on structural API prediction and response generation.
The key observations are: (a) All SIMMC neural models (HAE, HRE, MN, T-HAE) outperform the baselines (TF-IDF and LSTM) across all metrics for both the datasets. (b) HRE consistently achieves the highest API prediction accuracy for SIMMC-Furniture (, jointly with HAE) and SIMMC-Fashion (, jointly with HAE and MN). STOD++ achieves accuracy on attributes, an overwhelming point improvement over T-HAE for SIMMC-Furniture, benefiting from having access to the oracle belief state where the user requested attributes are formally represented. (c) For response generation, HRE has superior BLEU score for SIMMC-Furniture and HRE for SIMMC-Fashion. Surprisingly, T-HAE has the least BLEU scores amongst SIMMC models perhaps due to resorting to safe, frequent
responses. (d) The confusion matrix for HRE on SIMMC-Furniture (Appendix F) reveals a high confusion between SearchFurniture and None.
This is intuitive as searching for an item or further obtaining user preferences to narrow the search are equally valid actions for their context. Note that the proposed assistant models do not leverage the rich, fine-grained annotations of the SIMMC datasets (understandably so) as they are adaptations of existing state-of-the-art models.
Tab. 5 presents the performance of the state-of-the-art DST models on the SIMMC datasets. It can be seen that the pretrained GPT-2 based SimpleTOD models outperform the TRADE baseline. Note that the original TRADE implementation does not include the dialog act prediction, hence it is not reported here as well. When the multimodal contexts are added as input (SimpleTOD+MM), the performance improves upon the text-only SOTA model (SimpleTOD) on both datasets, especially in the slot prediction metrics. This demonstrates the efficacy of grounding the multimodal contexts for DST, by better resolving the multimodal coreferences. In general, the performance on the SIMMC-Fashion dataset is typically better than on the SIMMC-Furniture dataset. This could be due to the nature of the dialogs in the SIMMC-Fashion dataset, which involves more natural and casual utterances, as evident in the lower annotator agreement as well (Appendix C).
Conclusion
In this work, we presented Situated Interactive Multi-Modal Conversations (SIMMC), an important new direction towards building next generation virtual assistants with evolving multimodal inputs. In particular, we collected two new datasets using the SIMMC platform, and provided the contextual NLU and coreference annotations on these datasets, creating a new SIMMC task for the community to study. We established several strong baselines for some of the tasks enabled by the datasets, showcasing various uses of the datasets in real-world applications. The fine-grained annotations we collected open the door for studying several different tasks in addition to the ones highlighted in this work, which we leave as future work for the community to tackle.
Acknowledgements
We thank Pararth Shah, Oksana Buniak, Semir Shafi, Ümit Atlamaz, Jefferson Barlew, Becka Silvert, Kent Jiang, Himanshu Awasthi, and Nicholas Flores for their invaluable technical contributions to the data collection platforms, annotation schema development, annotation process, tooling and coordination. We also extend many thanks to all the annotators who meticulously labelled these datasets.
References
Appendix A Dataset & Annotation Analysis
Using the unified ontology framework described in Sec. 4.1, we annotate both the user and assistant utterances of the SIMMC datasets. There are effectively dialog acts which are respectively combined with activities for (SIMMC-Furniture) and activities for (SIMMC-Fashion); the latter by design excludes count and rotate. A detailed list with examples is in Appendix, Tab. 9. Not all combinations of dialog acts and activities are observed in our dataset, i.e., about for SIMMC-Furniture and for SIMMC-Fashion respectively. For instance, a request:disprefer utterance is an invalid combination. The key takeaways from Fig. 4 are: (a) inform is the most dominant dialog act ( in SIMMC-Fashion and in SIMMC-Furniture). This is intuitive as conversations in shopping domain require the user to inform the assistant of their preferences, while the assistant informs the user about the item attributes and availability. (b) Interestingly, get is the dominant activity across most dialog acts, where the assistant either gets new items or additional information about existing items that the user is perusing. (c) The relatively low occurrence of the confirm dialog act perhaps arises from the effectiveness of the human assistant agent. This is desirable to avoid learning assistant models that excessively repeat user requests, e.g., repeatedly seek explicit confirm, as this leads to lower user satisfaction. Note that this analysis of the dialog act and activity distribution is per sentence, with an utterance occasionally containing multiple sentences (see Fig. 1 for an example).
User Satisfaction Metrics.
Since SIMMC datasets aim at goal-oriented dialog, we also collect turn-level and dialog-level user satisfaction scores in the range of 1-5 as part of the data collection. The dialog-level user satisfaction scores for the SIMMC-Furniture dataset average at , showing a heavy concentration around 5. Since the dialogs are collected between humans interacting with each other, we hypothesize that the the assistant (wizard) is able to efficiently respond to user requests, leading to high satisfaction scores. Similar trends were observed across different metrics for both datasets. Therefore, we drop further analysis on this front due to the absence of a clear signal in these collected metrics.
Appendix B Details for SIMMC Labeling Language
A number of additional conventions are placed on the annotation task to ensure consistency and accuracy.
Type ambiguity. When an object appears in an utterance, the most fine-grained type is annotated. For example, in the utterance “Show me some dresses”, the token ‘dresses’ needs to be annotated as dress, as opposed to a coarser-grained type clothing. When more than one fine-grained type is possible, the annotator utilizes a parent-level coarse-grained type instead. Thus the assigned type is the finest-grained type that still captures the ambiguity.
Attribute ambiguity. Attributes are annotated when they are unambiguous. When there is uncertainty in the attribute that should be selected for the representation, the annotator falls back to a more generic attribute. For example when asking about an furniture item the user may specify a particular dimension, e.g. I want a couch that is 2 feet wide. In this case, 2 feet can be annotated with the specific attribute width. However, if the dimension is not specified, the more general attribute dimensions would be used.
Attribute inverses. When an attribute can be annotated in two different directions, a canonical attribute is defined in the ontology and used for all annotations. For example, attentionOn and inAttentionOf are inverses. The user object is connected to furniture or clothing objects via attentionOn if the user is looking at an instance of these objects. Inverserly, those same objects are connected to the user object via the inAttentionOf attribute. The former is designated as the canonical attribute in this case, and used for labeling purposes.
Smart prefixes. Attribute slots are prefixed by A and O respectively to indicate whether they serve to restrict the intent-level Activity or Object. This is primarily for human-annotator convenience. For example the attribute amount is an attribute of the Activity GET. The attribute color is an attribute of clothing. Annotating an assistant response like ’I found five green dress’ yields spanning five with a.amount and green with o.color.
Attribute variables. The attribute .info is employed when the speaker’s intent targets more than one attribute simultaneously. The specific attributes being targeted are then identified with the INFO smart prefix. For example, ’What is the color and brand of this skirt?’ is annotated with the intent da:ask:get:skirt.info and the tokens color and brand are labeled as info.color and info.brand respectively.
Appendix C Details for NLU/NLG/Coref Data Annotation
Data were annotated in two stages: (1) NLU/NLG followed by (2) image-based coreference annotations.
During the NLU/NLG stage, annotators were provided full dialog context for a single dialog and asked to annotate both the user and assistant’s utterances. Image context was not available, and annotators were instructed to use dialog context only up to the target utterance. Essentially all (98.4%) annotations were single-annotated by annotators who passed an evaluation test; while 1.6% were double-annotated. In cases of disagreement between two annotators, a third annotator either selected one of the proposed annotations, or overrode both with a new one.
In order to estimate and improve quality, we double-annotated an additional 12.6% of the data after the fact and applied two measures of inter-annotator agreement: exact matches between semantic parses, and a modified F1 score. 50% of fashion and 60% of furniture annotations were exact matches. Given sample sizes, this corresponds respectively to 95% confidence intervals of 49-50% and 58-63%. In contrast to this binary exact measure, the F1 measures can assume values between 0 and 1. Using this measure, furniture annotations were 73.8% similar while fashion annotations were 72.5% similar .
During image-based coreference, annotators were provided all dialog and image context up until the turn in question. Review of the process suggested it was easy enough for the high-skilled pool annotators to perform without quality checks. By their own account, annotators self-reported 98% confidence in their decision to link an object to an intent; and 98% confidence in their specific choice of object given a link was required.
Tab. 6 presents the Object Classes that were made available to the annotators for annotation as well as the attributes of these Classes. Attributes are listed alphabetically and type information is provided. Note that for readability attributes derived via inheritance from supertype to subtype are not repeated. Classes that were exposed to annotators but had no attributes are not presented here. Attribute ambiguity is indicated by indenting.
Tab. 7 presents the Activity Classes that were made available to the annotators for annotation as well as the attributes of these Classes (see Tab. 9 for examples and definitions). Type information and attributes are provided. Note that for readability attributes derived via inheritance from supertype to subtype are not repeated. All Activities had the attributes amount an INTEGER, endTime and startTime (DATE_TIMEs). Only Activities with additional attributes are listed below.
Appendix D API Call List
Tab. D shows the list of all APIs supported in our SIMMC datasets.
See Tab. 10 and Tab. 11 in Appendix G for annotated dialog examples that show our SIMMC ontology in action for both our datasets.
To help the readers track changes to this document, a brief changelog describing the revisions is provided below:
v1: COLING 2020 anonymity period version. Results on random SIMMC data splits.
v2: COLING 2020 camera-ready version. Results on standard SIMMC data splits.