Reinforcement Learning for Relation Classification from Noisy Data

Jun Feng, Minlie Huang, Li Zhao, Yang Yang, Xiaoyan Zhu

Introduction

Relation classification, aiming to categorize semantic relations between two entities given a plain text, is an important problem in natural language processing, particularly for knowledge graph completion and question answering. Most existing works for relation classification adopt supervised learning approaches, either based on traditional handcrafted features (?; ?) or based on the features automatically generated by deep neural networks (?; ?), but all require high-quality annotated data.

In order to obtain large-scale training data, distant supervision (?) was proposed by assuming that if two entities have a relation in a given knowledge base, all sentences that contain the two entities will mention that relation. Although distant supervision is effective to label data automatically, it suffers from the noisy labeling problem. Taking the triple (Barack_Obama, BornIn, United_States) as an example, the noisy sentence “Barack Obamba is the 44th president of the United State” will be regarded as a positive instance by distant supervision and a BornIn relation is assigned to this sentence, although the sentence does not describe the relation BornIn at all.

To address the issue of noisy labeling, previous studies adopt multi-instance learning to consider the noises of instances (?; ?; ?; ?; ?; ?). In these studies, the training and test process is proceeded at the bag level, where a bag contains noisy sentences mentioning the same entity pair but possibly not describing the same relation. As a result, previous studies suffer from two limitations: 1) Unable to handle the sentence-level prediction; 2) Sensitive to the bags with all noisy sentences which do not describe a relation at all.

To better explain the first limitation, we show an example in Figure 1. Bag-level prediction can find the two relations “EmployedBy” and “BornIn” between the entity pair “Barack_Obama” and “United_States”. However, sentence-level prediction is able to further map each relation to the corresponding sentences. As for the second limitation, for each bag, previous bag-level methods retain at least one sentence, even if all the sentences in a given bag are noisy (not describing the relation). Such bags, produced by distant supervision, are quite common. For instance, our investigation on a widely used datasethttp://iesl.cs.umass.edu/riedel/ecml/ shows that 53%53\% out of 100 sample bags have no sentences that describe the relation. Such noisy bags will definitely decrease the performance of relation classification.

In this paper, to handle the above two limitations, we propose a novel relation classification model consisting of two modules: instance selector and relation classifier. By having an explicit instance selectorInstance is referred to a sentence in this paper., we are able to first select high-quality sentences from a sentence bag, and then predict a relation at the sentence level by the relation classifier. To handle the second limitation, our instance selector will filter the entire bag if all sentences are labeled incorrectly. The major challenge here is how to train the two modules jointly, particularly when the instance selector has no explicit knowledge about which sentences are labeled incorrectly.

We address this challenge by casting the instance selection task as a reinforcement learning problem (?). Intuitively, although we do not have an explicit supervision for the instance selector, we can measure the utility of the selected sentences as a whole. Thus, the instance selection process has the following two properties: first, trial-and-error-search, meaning that the instance selector attempts to choose some sentences and obtain feedback (or reward) on the quality of the selected sentences from the relation classifier; second, the feedback from the relation classifier can be obtained only when we finish the instance selection process, which is typically delayed. These two properties naturally inspire us to utilize reinforcement learning techniques.

We propose a new model for relation classification, which consists of an instance selector and a relation classifier. This formalization enables our model to extract relations at the sentence level on the cleansed data.

We formulate instance selection as a reinforcement learning problem, which enables the model to perform instance selection without explicit sentence-level annotations but just with a weak supervision signal from the relation classifier.

Related Work

Relation classification is a common task in natural language processing. Many approaches have been developed, particularly with supervised methods (?; ?; ?). However, such supervised methods heavily rely on high-quality labeled data.

Recently, neural models have been widely applied to relation classification (?; ?; ?; ?) including convolutional neural networks, recursive neural network (?; ?), and long short-term memory network (?; ?; ?). In (?), two levels of attention is proposed in order to better discern patterns in heterogeneous contexts for relation classification.

In general, a large amount of labeled data are required to train neural models, which is quite expensive. To address this issue, distant supervision was proposed (?) by assuming that all sentences that mention two entities of a fact triple describe the relation in the triple. In spite of the success of distance supervision, such methods suffer from the noisy labeling issue. To alleviate this issue, many studies formulated relation classification as a multi-instance learning problem (?; ?; ?; ?). In (?; ?; ?), a sentence-level attention mechanism over multiple instances was proposed and incorrect sentences can be down-weighted. However, such multi-instance learning models all predict relations at the bag level but not at the sentence level, and they can not deal with the bags in which all sentences are not describing a relation at all. There are other approaches to reduce the noise of distant supervision using active learning (?) and negative patterns (?).

Previous methods are all at the bag level but not at the sentence level and as such, they cannot find the exact mapping between a relation and a sentence. Furthermore, these methods are unable to handle the bags in which all the sentences are not describing the relation. To address these issues, we propose a new framework which first selects correct sentences in the framework of reinforcement learning (?; ?) and then predicts relations from each sentence in the cleansed data.

Methodology

We propose a new relation classification framework, which is able to select correct sentences from noisy data for better relation classification. The proposed framework can predict relations at the sentence level from the cleansed data, rather than at the bag level. Sentence-level prediction is more friendly to the tasks that need to comprehend sentences such as question answering and semantic parsing.

Our framework consists of two key modules: the instance selector which selects correct sentences from noisy data, and the relation classifier which predicts relation and updates its parameters with cleaned data. The two modules interacts with each other during the training process.

Formally, we decompose the task of relation classification into two sub-problems in this paper: instance selection and relation classification.

We formulate the instance selection problem as follows: given a set of <<sentence, relation label>> pairs as X={(x1,r1),(x2,r2),…,(xn,rn)}X=\{(x_{1},r_{1}),(x_{2},r_{2}),\dots,(x_{n},r_{n})\}, where xix_{i} is a sentence associated with two entities (hi,ti)(h_{i},t_{i}) and rir_{i} is a noisy relation label produced by distant supervision. The goal is to determine which sentence truly describes the relation and should be selected as a training instance.

The relation classification problem is formulated as follows: given a sentence xix_{i} and the mentioned entity pair (hi,ti)(h_{i},t_{i}), the goal is to predict the semantic relation rir_{i} in xix_{i}. Essentially, the model estimates the probability: pΦ(ri∣xi,hi,ti)p_{\Phi}(r_{i}|x_{i},h_{i},t_{i}).

Overview

The proposed model is based on a reinforcement learning framework and consists of two components: the instance selector and the relation classifier. In the instance selector, each sentence xix_{i} has a corresponding action aia_{i} to indicate whether or not xix_{i} will be selected as a training instance for relation classification. The state sis_{i} is represented by the current sentence xix_{i}, the already chosen sentences among {x1,⋯ ,xi−1}\{x_{1},\cdots,x_{i-1}\}, and the entity pair hih_{i} and tit_{i} in sentence xix_{i}. The instance selector samples an action given the current state according to a stochastic policy. For the relation classifier, it adopts a convolutional architecture to automatically determine the semantic relation for an entity pair in a given sentence. The instance selector distills the training data to the relation classifier to train the convolutional neural network. Meanwhile, the relation classifier gives feedback to the instance selector to refine its policy function. Figure 2 gives an illustration of how the proposed framework works.

With the help of the instance selector, our method directly filters out noisy sentences. Unlike reducing the weights of noisy sentences (?) or retaining one sentence in a bag (?), our method is better at dealing with noisy data. The relation classifier is trained and tested at the sentence level on the cleansed data, whereas previous models treat the sentence bag as a whole and predict relation at the bag level.

Instance Selector

We cast instance selection as a reinforcement learning problem. The instance selector is the agent, who interacts with the environment that consists of data and the relation classifier. The agent follows a policy to decide which action (choosing the current sentence or not) at each state (consisting of the current sentence, the chosen sentence set, and the entity pair), and then receive a reward from the relation classifier at the terminal state when all the selections are made.

As aforementioned, we can obtain a delayed reward from the relation classifier only when the selection on all the training instances are finished. Thus, we can only update the policy function once for each scan of the entire training data, which is obviously inefficient. To obtain more feedbacks and to make the training process more efficiently, we split the training sentence instances X={x1,…,xn}X=\{x_{1},\dots,x_{n}\} into NN bags B={B1,B2,…,BN}\mathbf{B}=\{B^{1},B^{2},\dots,B^{N}\} and compute a reward when we finish data selection in a bag. Each bag corresponds to a distinct entity pair, and each bag BkB^{k} is a sequence of sentences {x1k,x2k,…,x∣Bk∣k}\{x_{1}^{k},x_{2}^{k},\dots,x_{|B^{k}|}^{k}\} with the same relation label rkr^{k}, however, the relation label is noisy. We define the action as selecting a sentence or not according to a policy function. The reward is computed once the selection decisions are completed on one bag. When the training process of the instance selector is completed, we merge all the selected sentences in each bag to obtain a cleansed dataset X^\hat{X}. Then, the cleansed data will be used to train the relation classifier at the sentence level.

State. The state sis_{i} represents the current sentence, the already selected sentences, and the entity pair when making decision on the ii-th sentence of the bag BB. We represent the state as a continuous real-valued vector F(si)\bm{F}(s_{i}), which encodes the following information: 1) The vector representation of the current sentence, which is obtained from the non-linear layer of the CNN for relation classification; 2) The representation of the chosen sentence set, which are the average of the vector representations of all chosen sentences; 3) The vector representations of the two entities in a sentence, obtained from a pre-trained knowledge graph embedding table.

Action. We define an action ai∈{0,1}a_{i}\in\{0,1\} to indicate whether the instance selector will select the ii-th sentence of the bag BB or not. We sample the value of aia_{i} by its policy function πΘ(si,ai)\pi_{\Theta}(s_{i},a_{i}), where Θ\Theta is the parameters to be learned. In this work, we adopt a logistic function as the policy function:

where F(si)\bm{F}(s_{i}) is the state feature vector, and σ(.)\sigma(.) is the sigmoid function with the parameter Θ={W,b}\Theta=\{\bm{W},\bm{b}\}.

Reward. The reward function is an indicator of the utility of the chosen sentences. For certain bag B={x1,…,x∣B∣}B=\{x_{1},\dots,x_{|B|}\}, we sample an action for each sentence, to determine whether the current sentence should be selected or not. We assume that the model has a terminal reward when it finishes all the selection. Therefore we only receive a delayed reward at the terminal state s∣B∣+1s_{|B|+1}. The reward is zero at other states. Therefore, the reward is defined as follows:

where B^\hat{B} is the set of selected sentences, which is a subset of BB, and rr is the relation label of bag BB. As shown in Figure 2, p(r∣xj)p(r|x_{j}) is calculated by the relation classifier which is given by a CNN model. For the special case B^=∅\hat{B}=\emptyset, we set the reward as the average likelihood of all sentences in the training data, which enables our instance selector to exclude noisy bag effectively.

Note that the relation classifier is at the sentence-level since it computes p(r∣x)p(r|x) for each sentence. The reward is computed on a new bag of sentences selected by the instance selector. Essentially, the above reward evaluates the overall utility of all the actions made by the policy. It supervises the instance selector to maximize the average likelihood of the chosen instances, which makes the objective function of the instance selector consistent with the relation classifier.

In the selection process, not only the final action contributes to this reward, but also all the previous actions do. Therefore, this reward is delayed, and can be handled very well by reinforcement learning techniques (?).

Optimization. For a bag BB, we aim to maximize the expected total reward. More formally, our objective function is defined as

where ai∼πΘ(si,ai)a_{i}\sim\pi_{\Theta}(s_{i},a_{i}), si+1∼P(si+1∣si,ai)s_{i+1}\sim P(s_{i+1}|s_{i},a_{i}). The transition functions P(si+1∣si,ai)P(s_{i+1}|s_{i},a_{i}) are equal to 1, since the state si+1s_{i+1} is fully determined by the state sis_{i} and aia_{i}. VΘV_{\Theta} is the value function, and VΘ(s1∣B)V_{\Theta}(s_{1}|B) represents the expected future total reward that we can obtain by starting at certain state s1s_{1} following policy πΘ(si,ai)\pi_{\Theta}(s_{i},a_{i}).

According to the policy gradient theorem (?) and the REINFORCE algorithm (?), we compute the gradient in the following way. For each bag BB, we sample an action for each state sequentially according to the current policy. We then get a sampled trajectory {s1,a1,s2,a2,...,s∣B∣,a∣B∣,s∣B∣+1}\{s_{1},a_{1},s_{2},a_{2},...,s_{|B|},a_{|B|},s_{|B|+1}\} and a corresponding terminal reward r(s∣B∣+1∣B)r(s_{|B|+1}|B). Since we only have a non-zero terminal reward, the value function is the same for all states from s1s_{1} to s∣B∣s_{|B|}, namely vi=V(si∣B)=r(s∣B∣+1∣B)v_{i}=V(s_{i}|B)=r(s_{|B|+1}|B), for i=1,2,...,∣B∣i=1,2,...,|B|. We update the current policy using the following gradient:

Relation classifier

In the relation classifier, we adopt a CNN architecture to predict relations. The CNN network has an input layer, a convolution layer, a max pooling layer and a non-linear layer from which the representation is used for relation classification.

CNN. In order to obtain high-level and abstractive representation of the raw input of a sentence, we apply a CNN structure for relation classification. This can be briefly described as below:

Then, the probability for relation prediction p(r∣x;Φ)p(r|x;\mathbf{\Phi}) is given as follows:

The key difference between our relation classifier and other studies lies in that our classifier performs relation classification at the sentence level. The input to the relation classifier in other studies is a bag of sentences. Instead, the input to ours is just one sentence, since we already filter out noisy sentences with the instance selector.

Loss function. Given the selected training set {X^}\{\hat{X}\} provided by the instance selector, we define the objective function of the relation classifier using cross-entropy as follows:

Model Training

As the instance selector and the relation classifier are correlated mutually, we train them jointly.The complete joint training process is described in Algorithm 1. To optimize the policy network in the instance selector, we use a Monto-Carlo based policy gradient method (?), which favors actions with high sampled reward. To optimize the CNN component, we use a gradient descent method to minimize the objective function (i.e., Eq. 7). We pre-train the model before the joint training process starts. We first pre-train the CNN in the relation classifier, and then pre-train the policy function by computing the reward with the pre-trained CNN, while the parameters of the CNN model are frozen. At last, we jointly train the instance selector and the relation classifier. We found such a pre-training strategy is quite crucial for our method, which is also widely recommended by many other reinforcement learning studies(?).

Algorithm 2 presents the details of the joint training process. The relation classifier provides a mechanism of computing the rewards of the selected sentences to refine the instance selector. The instance selector chooses high-quality data by excluding wrongly labeled sentences to better train the relation classifier. In order to have a stable update, we take advantage of a target policy network and a target CNN with parameter sets Θ′\Theta^{\prime} and Φ′\Phi^{\prime} respectively, similar to (?). The parameters in the target networks are updated much more slowly than the original ones. We update Θ′\Theta^{\prime} and Φ′\Phi^{\prime} by linear interpolation: Θ′←(1−τ)Θ′+τΘ\Theta^{\prime}\leftarrow(1-\tau)\Theta^{\prime}+\tau\Theta and Φ′←(1−τ)Φ′+τΦ\Phi^{\prime}\leftarrow(1-\tau)\Phi^{\prime}+\tau\Phi, where τ≪1\tau\ll 1 is a hyper-parameter.

Experiment

Dataset. To evaluate our model, we adopted a widely used datasethttp://iesl.cs.umass.edu/riedel/ecml/ generated by the sentences in NYTNew York Times, a widely used text corpus. and developed by (?). There are 522,611 sentences, 281,270 entity pairs, and 18,252 relational facts in the training data; and 172,448 sentences, 96,678 entity pairs and 1,950 relational facts in the test data. Among the data, there are 39,528 unique entities and 53 unique relations from Freebase including a special relation NA that signifies no relation between two entities in a sentence.

Word and entity embedding. We adopted word2vec to train the word embeddings on the NYT corpus. For entity embedding, we implemented the TransE model (?) and trained it on a set of Freebase fact triples whose entities have been mentioned in the training and test data.

Model pre-training. As described in Algorithm 2, we pre-trained the relation classifier and instance selector before the joint training process. As the reward is calculated based on the CNN model in the relation classifier, we first pre-trained the CNN model on the entire training data. Then, we fixed the parameters of the CNN model and pre-trained the policy function in the instance selector where the reward is obtained from the fixed CNN model.

Parameter setting. Similar to previous studies, we tuned our model using three-fold cross validation. For the parameters of the instance selector, we set the dimension of entity embedding as 5050, the learning rate as 0.020.02/0.010.01 at the pre-training stage and joint training stage respectively. The delay coefficient τ\tau is 0.0010.001.

For the parameters of the relation classifier, the word embedding dimension dw=50d^{w}=50 and the position embedding dimension dp=5d^{p}=5. The window size of the convolution layer ll is 33. The learning rate of the instance selector is α=0.02\alpha=0.02 both at the pre-training and joint training stage. The batch size is fixed to 160160. The training episode number L=25L=25. We employed a dropout strategy with a probability of 0.50.5 during the training of the CNN component.

Sentence-Level Relation Classification

As discussed previously, the key difference between our method and other models lies in that our method can perform sentence-level relation classification. We conducted manual evaluation on relation classification in this section.

Evaluation settings. We predicted a relation label for each sentence, instead of for each bag. For example, the task in Figure 1 needs to map the first sentence to relation “BornIn” and the second sentence to “EmployedBy”.

Since the data obtained from distant supervision are noisy, we randomly chose 300 sentences and manually labeled the relation type for each sentence to evaluate the classification performance. We adopted accuracy and macro-averaged F1F_{1} as the evaluation metric.

Baselines. We adopted three state-of-the-art baselines:

CNN (?) is a sentence-level classification model. It does not consider the noisy labeling problem.

CNN+Max (?) is a bag-level classification model. It assumes that there is one sentence describing the relation in a bag. It chooses the most correct sentence in each bag.

CNN+ATT (?) is also a bag-level model, similar to CNN+Max. It adopts a sentence-level attention over the sentences in a bag and thus can down weight noisy sentences in a bag.

CNN is a sentence-level model that is trained directly on noisy data. For bag-level models (CNN+Max and CNN+ATT), the training process is the same as the referenced papers. During test, each sentence is treated as a bag and a relation is predicted for each bag. In this scenario, the bag-level relation prediction is exactly the same as the sentence-level prediction. All the baselines were implemented with the source codes released by (?).

Results. Results in Table 1 reveal the following observations.

CNN+RL obtains superior performance than CNN, indicating that filtering noisy data by instance selection benefits the task.

CNN+RL outperforms CNN+Max and CNN+ATT remarkably. It shows the effectiveness of instance selection with reinforcement learning.

The sentence-level models (CNN and CNN+RL) perform much better than the bag-level models (CNN+Max and CNN+ATT), indicating that bag-level models do not perform well for sentence-level prediction.

Instance Selection

We then evaluated the effectiveness of our instance selector from several aspects. First, we evaluated whether the selected data by our instance selector are better for relation classification. Second, we justified the accuracy of selection decision in the selector by manually checking the decisions on sentences. Third, we compared the proposed RL selection strategy in our selector with greedy selection. Last, we assessed whether the selector has the ability of filtering those bags that contain all noisy sentences.

Relation classification on selected data. To measure the quality of the selected data by our instance selector, we performed relation classification experiments on the selected data. We first used our instance selector to select the high-quality sentences from the original data. Then, we trained two state-of-the-art models, CNN and CNN+ATT with two settings. One setting is to train them on the original data, named as CNN(Original) and CNN+ATT(Original). The other setting is to train them on the selected data, which are named as CNN(Selected) and CNN+ATT(Seleted). We compared the performance of CNN(Original) (CNN+ATT(Original)) with CNN(Selected) (CNN+ATT(Selected)) on the relation classification task. The results are compared under the held-out evaluation configuration (?) which provides an approximate measure of relation classification without expensive human annotations. The held-out evaluation compares the predicted relational fact from the test data with the facts in Freebase, but it does not consider the mapping between a relational fact and a sentence.

As shown in Figure 4 and Figure 4, the models trained on the selected data achieve much better performance than the counterparts trained on the original dataset. The results also indicate our instance selector has the ability of filtering out noisy sentences and distilling high-quality sentences, resulting better classification performance.

Accuracy of instance selection decision. To assess how accurate the decision is by the instance selector, we manually checked each sentence selected and rejected by the instance selector in a sampled dataset. For each sentence, the instance selector makes a correct decision if the sentence’s label is correct and our instance selector selects it as a training instance, or, if its label is wrong and our instance selector rejects it. Otherwise, we judged that the instance selector makes a wrong decision.

Specifically, we sampled 300300 sentences from the training data. Our instance selector chooses 6464 sentences as the training instances, among which 4545 sentences are correctly selected. The selector also rejects 236236 instances, and 177177 of them are noisy instances (not describing the relation). To summarize, the accuracy of our instance selector is (45+177)/300=74%(45+177)/300=74\%, which demonstrates the effectiveness of our instance selector.

Different instance selection strategies. To show the necessity for adopting the reinforcement learning framework for instance selection, we compared two instance selection strategies. Specifically, we performed relation classification on the selected data respectively with reinforcement learning (RL) selection and with greedy selection. The greedy selection selects the top NN sentences with the largest likelihood which is estimated by a pre-trained CNN.

During the experiments, we kept the relation classifier untouched while replacing the RL selection by greedy selection. The number of selected instances NN is the same as the RL strategy. As shown in Figure 5, the performance of our instance selector is much better than the greedy strategy on the held-out evaluation. The results show that our RL strategy is reasonable and effective.

Noisy bag filtering. As previous methods cannot filter the bags with all noisy sentences, we validated the ability of our model to filter bags with all noisy sentences. We randomly selected 100100 deleted sentence bags and find that 86%86\% of the bags consist of all noisy sentences. This indicates that our instance selector can exclude the noisy sentences effectively.

Case Study

Table 2 shows two bag examples for instance selection. The first bag has three correct sentences. The second bag has two noisy sentences. It is clearly show that our model can do better instance selection than both instance-weighting with CNN+ATT and maximum likelihood selection with CNN+Max. The second example indicates that our model is able to filter bags with all noisy sentences while other methods fail to do so.

Conclusion and Future Work

In this paper, we propose a novel model for sentence-level relation classification from noisy data using a reinforcement learning framework. The model consists of an instance selector and a relation classifier. The instance selector chooses high-quality data for the relation classifier. The relation classifier predicts relation at the sentence level and provides rewards to the selector as a weak signal to supervise the instance selection process. Extensive experiments demonstrate that our model can filter out the noisy sentences and perform sentence-level relation classification better than state-of-the-art baselines from noisy data.

Further, our solution for instance selection can be generalized to other tasks that employ noisy data or distant supervision. For instance, a possible attempt might be to perform sentiment classification on noisy data (?). We leave this as our future work.

Acknowledgement

This work was partly supported by the National Science Foundation of China under grant No.61272227/61332007.

References