CAIL2019-SCM: A Dataset of Similar Case Matching in Legal Domain
Chaojun Xiao, Haoxi Zhong, Zhipeng Guo, Cunchao Tu, Zhiyuan Liu, Maosong Sun, Tianyang Zhang, Xianpei Han, Zhen Hu, Heng Wang, Jianfeng Xu
Introduction
Similar Case Matching (SCM) plays a major role in legal system, especially in common law legal system. The most similar cases in the past determine the judgment results of cases in common law systems. As a result, legal professionals often spend much time finding and judging similar cases to prove fairness in judgment. As automatically finding similar cases can benefit to the legal system, we select SCM as one of the tasks of CAIL2019.
Chinese AI and Law Challenge (CAIL) is a competition of applying artificial intelligence technology to legal tasks. The goal of the competition is to use AI to help the legal system. CAIL was first held in 2018, and the main task of CAIL2018 Xiao et al. (2018); Zhong et al. (2018b) is predicting the judgment results from the fact description. The judgment results include the accusation, applicable articles, and the term of penalty. CAIL2019 contains three different tasks, including Legal Question-Answering, Legal Case Element Prediction, and Similar Case Matching. Furthermore, we will focus on SCM in this paper.
More specifically, CAIL2019-SCM contains 8,964 triplets of legal documents. Every legal documents is collected from China Judgments Onlinehttp://wenshu.court.gov.cn/. In order to ensure the similarity of the cases in one triplet, all selected documents are related to Private Lending. Every document in the triplet contains the fact description. CAIL2019-SCM requires researchers to decide which two cases are more similar in a triplet. By detecting similar cases in triplets, we can apply this algorithm for ranking all documents to find the most similar document in the database. There are 247 teams who have participated CAIL2019-SCM, and the best team has reached a score of , which is about points higher than the baseline. The results show that the existing methods have made great progress on this task, but there is still much room for improvement.
In other words, CAIL2019-SCM can benefit the research of legal case matching. Furthermore, there are several main challenges of CAIL2019-SCM: (1) The difference between documents may be small, and then it is hard to decide which two documents are more similar. Moreover, the similarity is defined by legal workers. We must utilize legal knowledge into this task rather than calculate similarity on the lexical level. (2) The length of the documents is quite long. Most documents contain more than characters, and then it is hard for existing methods to capture document level information.
In the following parts, we will give more details about CAIL2019-SCM, including related works about SCM, the task definition, the construction of the dataset, and several experiments on the dataset.
Related Work
SCM aims to measure the similarity between legal case documents. Essentially, it is an application of semantic text matching, which is central for many tasks in natural language processing, such as question answering, information retrieval, and natural language inference. Take information retrieval as an example, given a query and a database, a semantic matching model is required to judge the semantic similarity between the query and documents in the database. Moreover, the tasks related to semantic matching have attracted the attention of many researchers in recent decades.
Intuitively traditional approaches calculate word-to-word similarity with vector space model, e.g. term frequency-inverse document frequency Wu et al. (2008), bag-of-words Bilotti et al. (2007). However, due to the variety of words in different texts, these approaches achieve limited success in the task.
Recently, with the development of deep learning in natural language processing, researchers attempt to apply neural models to encode text into distributed representation. The Siamese structure Bromley et al. (1994) for metric learning achieve great success and is widely applied Amiri et al. (2016); Liu et al. (2018); Mueller and Thyagarajan (2016); Neculoiu et al. (2016); Wan et al. (2016); He et al. (2015). Besides, there are many researchers put emphasis on integrating syntactic structure into semantic matching Liu et al. (2018); Chen et al. (2017) and multi-level text matching with attention-aware representation Duan et al. (2018); Tan et al. (2018); Yin et al. (2016).
Nevertheless, most previous studies are designed for identifying the relationship between two sentences with limited length.
2 Legal Intelligence
Researchers widely concern tasks for legal intelligence. Applying NLP techniques to solve a legal problem becomes more and more popular in recent years. Previous works Kort (1957); Keown (1980); Lauderdale and Clark (2012) focus on analyzing existing cases with mathematical tools. With the development of deep learning, more researchers pay much efforts on predicting the judgment result of legal cases Luo et al. (2017); Hu et al. (2018); Zhong et al. (2018a); Chalkidis et al. (2019); Jiang et al. (2018); Yang et al. (2019). Furthermore, there are many works on generating court views to interpret charge results Ye et al. (2018), information extraction from legal text Vacek and Schilder (2017); Vacek et al. (2019), legal event detection Yan et al. (2017), identifying applicable law articles Liu et al. (2015); Liu and Hsieh (2006) and legal question answering Kim et al. (2015); Fawei et al. (2018).
Meanwhile, retrieving related legal documents with a query has been studied for decades and is a critical issue in applications of legal intelligence. Raghav et al. (2016) emphasize exploiting paragraph-level and citation information. Kano et al. (2017) and Zhong et al. (2018b) held a legal information extraction and entailment competition to promote progress in legal case retrieval.
Overview of Dataset
We first define the task of CAIL2019-SCM here. The input of CAIL2019-SCM is a triplet , where are fact descriptions of three cases. Here we define a function which is used for measuring the similarity between two cases. Then the task of CAIL2019-SCM is to predict whether or .
2 Dataset Construction and Details
To ensure the quality of the dataset, we have several steps of constructing the dataset. First, we select many documents within the range of Private Lending. However, although all cases are related to Private Lending, they are still various so that many cases are not similar at all. If the cases in the triplets are not similar, it does not make sense to compare their similarities. To produce qualified triplets, we first annotated some crucial elements in Private Lending for each document. The elements include:
The properties of lender and borrower, whether they are a natural person, a legal person, or some other organization.
The type of guarantee, including no guarantee, guarantee, mortgage, pledge, and others.
The usage of the loan, including personal life, family life, enterprise production and operation, crime, and others.
The lending intention, including regular lending, transfer loan, and others.
Conventional interest rate method, including no interest, simple interest, compound interest, unclear agreement, and others.
Interest during the agreed period, including , , , and others.
Borrowing delivery form, including no lending, cash, bank transfer, online electronic remittance, bill, online loan platform, authorization to control a specific fund account, unknown or fuzzy, and others.
Repayment form, including unpaid, partial repayment, cash, bank transfer, online electronic remittance, bill, unknown or fuzzy, and others.
Loan agreement, including loan contract, or borrowing, “WeChat, SMS, phone or other chat records”, receipt, irrigation, repayment commitment, guarantee, unknown or fuzzy and others.
After annotating these elements, we can assume that cases with similar elements are quite similar. So when we construct the triplets, we calculate the tf-idf similarity and elemental similarity between cases and select those similar cases to construct our dataset. We have constructed 8,964 triples in total by these methods, and the statistics can be found from Table 1. Then, legal professionals will annotate every triplet to see whether or . Furthermore, to ensure the quality of annotation, every document and triplet is annotated by at least three legal professionals to reach an agreement.
Experiments
To access the challenge of the similar cases matching task, we evaluate several baselines on our dataset. The experiment results show that even the state-of-the-art systems perform poorly in evaluating the similarity between different cases.
Baselines. All the baseline models are trained on Large Train and tested on Large Valid and Large Test. We adapt the Siamese framework Bromley et al. (1994) to our scenario with different encoder, e.g. CNN Kim (2014), LSTM Hochreiter and Schmidhuber (1997), Bert Devlin et al. (2019), used for encoding the legal documents. We will elaborate on the details of the framework in the following part.
Given the triplet of fact description, (, , ), we first encode them into distributed vectors with the same encoder and then compute the similarity scores between the query case and the candidate cases , with a linear layer. Assume that each document consisting of words, i.e. .
For the learning objective, we apply the binary cross-entropy loss function with ground-truth label :
Model Performance. We use the accuracy metric in our experiments. Table 2 shows the results of baselines and top 3 participant teams on Large Valid and Large Test dataset, from which we get the following conclusion: 1) The participants achieve promising progress compared to baseline models. 2) Both the baselines systems and participant teams perform poorly on the dataset, due to the lack of utilization of prior legal knowledge. It’s still challenging to utilize legal knowledge and simulate legal reasoning for the dataset.
Conclusion
In this paper, we propose a new dataset, CAIL2019-SCM, which focuses on the task of similar case matching in the legal domain. Compared with existing datasets, CAIL2019-SCM can benefit the case matching in the legal domain to help the legal partitioners work better. Experimental results also show that there is still plenty of room for improvement.