Deepr: A Convolutional Net for Medical Records

Phuoc Nguyen, Truyen Tran, Nilmini Wickramasinghe, Svetha Venkatesh

I Introduction

A major theme in modern medicine is prospective healthcare, which refers to the capability to estimate the future medical risks for individuals. These risks can include readmission after discharge, the onset of specific diseases, and worsening from a condition . Such capability would facilitate timely prevention or intervention for maximum health impact, and provide a major step toward personalized medicine. An important data resource in aiding this process are electronic medical records . Electronic medical records (EMRs) contain a wealth of patient information over time. Central to EMR-driven risk prediction is patient representation, also known as feature engineering. Representing an EMR amounts to extracting relevant historical signals to form a feature vector.

However, feature extraction in EMR is challenging . An EMR typically consists of a sequence of time-stamped visit episodes, each of which has a subset of coded diagnoses, a subset of procedures, lab tests and textual narratives. The data is irregular at patient level. EMR is episodic – events are only recorded when patients visit clinics, and the time gap between two visits is largely random. Representing irregular timing poses a major challenge. EMR varies greatly in length – young patients usually have just one visit for an acute condition, but old patients with chronic conditions may have hundreds of visits. At the same time, the data is regular at local episode level. Diseases tend to form clusters (comorbidity) and the disease progression may be dictated by the underlying biological processes . Likewise treatments may follow a certain protocol or best practice guideline , and there are well-defined disease-treatment interactions . These regularities can be thought as clinical motifs. Thus an effective EMR representation should be able to identify regular clinical motifs out of irregular data.

Existing EMR-driven predictive work often relies on high-dimensional sparse feature representation, where features are engineered to capture certain regularities of the data This feature engineering practice is effort intensive and non-adaptive to varying medical records systems. Automated feature representation based on bag-of-words (BoW) is scalable, but it breaks collocation relations between words and ignores the temporal nature of the EMR, thus it fails to properly address the aforementioned challenges.

In this work we present a new prediction framework called Deepr\boldsymbol{\mathtt{Deepr}} that does not require manual feature engineering. The technology is based on deep learning, a new revolutionary approach that aims to build a multilayered neural learning system like a brain . When fed with a large amount of raw data, the system learns to recognize patterns with little help from domain experts. Deep learning now powers speech recognition in Google Voice, self-driving cars at Google and Baidu, question answering system at IBM (Watson), and smart assistants at Facebook. It already has a great impact on hundreds of millions (if not billions) people. But healthcare has largely been ignored. We hypothesize that a key to apply deep learning for healthcare patient representation which requires a proper handling of the irregular nature of episodes mentioned above . Deepr\mathtt{Deepr} fills the gap by offering an end-to-end technology that learns to represent patients from scratch. It reads medical records, learns the local patterns, adapts to irregular timing, and predicts personalized risk.

The architecture of Deepr\mathtt{Deepr} is multilayered and is inspired by recent convolutional neural nets (CNNs) in natural languages . The most crucial operation occurs at the bottom level where Deepr\mathtt{Deepr} transforms an EMR into a “sentence” of multiple phrases separated by special “words” that represent time gap. Each phrase is an visit episode. As with syntactical grammars and collocation patterns in NLP, there might exist “health grammars” and “clinical patterns” in healthcare. Health grammars refer to latent biological and environmental laws that dictate the global evolution of one’s health over time, e.g., probable progression from “diabetes type II” to “renal failure”. To handle irregular timing, time gaps and transfers are treated as special words. With this representation, an EMR is transformed into a sentence of variable length that retains all important events. The other layers of Deepr\mathtt{Deepr} constitute a CNN, which is similar to those in . First, words are embedded into a continuous vector space. Next, words in sentence are passed through a convolution operation which detects local motifs. Local motifs are then pooled to form a global feature vector, which is passed into a classifier, which predicts the future risk. All components are learned at the same time from data: the data signals are passed from the data to the output, and the training signals are propagated back from the labels to the motif detectors. Hence Deepr\mathtt{Deepr} is end-to-end.

We validate Deepr\mathtt{Deepr} on a large database of 300K patients collected from a hospital chain in Australia. We focus on predicting unplanned readmission within 6 months after discharge. Compared to existing bag-of-words representation, Deepr\mathtt{Deepr} demonstrates a superior accuracy as well as the capacity to learn predictive clinical motifs, and to uncover the underlying structure of the space of diseases and interventions.

To summarize, we claim the following contributions:

A novel representation of irregular-time EMR as a sentence with time gaps and transfers as special words.

A novel deep learning architecture called Deepr\mathtt{Deepr} that (i) uncovers the structure of the disease/treatment space, (ii) discovers clinical motifs, (iii) predicts future risk and (iv) explains the prediction by identifying motifs with strong responses in each record. The system is end-to-end, and its inner working can be inspected and visualized, allowing interpretability and transparency.

An evaluation of these claimed capabilities on a large-scale dataset of 300K patients.

II Background

An electronic medical record (EMR) contains information about patient demographics and a sequence of hospital visits for a patient. Admission information may include admission time, discharge time, lab tests, diagnoses, procedures, medications and clinical narratives. Diagnoses, procedures and medications are discrete entities. For example, diagnoses may be represented using ICD-10 coding schemeshttp://apps.who.int/classifications/icd10/browse/2016/en. For example, in ICD-10, E10 refers to Type 1 diabetes mellitus, E11 to Type 2 diabetes mellitus. The procedures are typically coded in CPT (Current Procedural Terminology) or ICHI (International Classification of Health Interventions) schemes http://www.who.int/classifications/ichi/en/. One of the most important secondary uses of EMR is building predictive models .

Most existing prediction methods on EMRs either rely on manual feature engineering or simplistic extraction . They either ignore long-term dependencies or do not adequately capture variable length . Neither are they able to model temporal irregularity . Capturing disease progression has been of great interest , and much effort has been spent on Markov models . As Markov processes are memoryless, Markov models forget severe conditions of the past when it sees an admission due to common cold. This is undesirable. A proper modeling, therefore, must be non-Markovian and able to capture long-term dependencies.

Deep learning

Deep learning is an approach in machine learning, aiming at producing end-to-end systems that learn from raw data and perform desired tasks without manual feature engineering. The current wave of deep learning was initiated by the seminal work of in 2006, but deep learning has been developed for decades . Over the past few years, deep learning has broken records in cognitive domains such as vision, speech and natural languages . Current deep learning is mostly based on multilayered neural networks . All the networks share a common unit – the neuron – which is a simple computational device that applies a nonlinear transform to a linear function of inputs: i.e., f(x)=σ(b+∑iwixi)f(x)=\sigma\left(b+\sum_{i}w_{i}x_{i}\right). Almost all networks thus far are trained using back-propagation , thus enable end-to-end learning.

There are three main deep neural architectures in practice: feedforward, recurrent and convolutional. Feedforward nets (FFN) pass unstructured information from one end to the other, usually from an input to an output, hence they act as a universal function approximator . Recurrent nets (RNN) model dynamics over time (and space) using self-replicated units. They maintain some degree of memory, and thus have potential to capture long-term dependencies. RNNs are powerful computational machines – they can approximate any program . Convolutional nets (CNN) exploit the repeated local motifs across time and space, and thus are translation-invariant – the capacity often seen in human visual cortex . Local motifs are small piece of data, usually of pre-defined sizes, e.g., a batch of pixels, or a n-gram of words. CNN is often equipped with pooling operations to reduce the resolution and enlarge the motifs.

III 𝙳𝚎𝚎𝚙𝚛𝙳𝚎𝚎𝚙𝚛\mathtt{Deepr}: A Deep Net for Medical Records

In this section, we describe our deep neural net named Deepr\mathtt{Deepr} (short for Deep net for medical Record) for representing Electronic Medical Records (EMR) and predicting the future risk.

Deepr\mathtt{Deepr} is a multilayered architecture based on convolutional neural nets (CNNs). The information flow is summarized in Fig. 1. At the bottom level, Deepr\mathtt{Deepr} sequences the EMR into a “sentence”, or equivalently, a sequence of “words”. Each word represents a discrete object or event such as diagnosis, procedure, or any derived object such as time-interval or hospital transfer. The next layer embeds words into an Euclidean space. On top of the embedding layer is a CNN that reads a small chunk of words in a sliding window to identify local motifs. The local motifs are transformed by Rectified Linear Unit (ReLU), which is a nonlinear function. All the transformed motifs are then max-pooled across the sentence to derive an EMR-level feature vector. Finally, a linear classifier is placed at the top layer for prediction. The entire architecture of Deepr\mathtt{Deepr} can be summarized as a function f(r)f(r) for record rr:

The CNN plays a crucial role as it detects clinical motifs that are predictive. Clinical motifs are co-occurrences of diseases (also known as comorbidity), disease progression, patterns of disease/treatment, and patterns of collocating treatments . However, as CNN is supervised it requires labels, which may not always be available (e.g., new patients with short history). A possible enhancement is through pretraining the embedding layer through a powerful tool known as word2vec . As word2vec is unsupervised and relies on local collocation patterns, clinical motifs can be pre-detected, and then further refined through CNN with supervising signals.

III-B Sequencing EMR

This task refers to transforming an EMR into a sentence, which is essentially a sequence of words. We present here how the words are defined and arranged in the sentence.

Recall that an EMR is a sequence of time-stamped visit episodes. Each episode may contain many pieces of information, but for the purpose of this work, we focus mainly on diagnoses and treatments (which involve clinical procedures and medications). For simplicity, we do not assume perfect timing of each piece, and thus an episode is a finite set of discrete words (diagnoses and treatments). The episode is then sequenced into a phrase. The order of the element in the phrase may follow the pre-defined ordering by the EMR system, for example, primary diagnosis is placed first, followed by secondary diagnoses, followed by procedures. In absence of this information, we may randomize the elements.

Within an episode, occasionally, there are one or more transfers between care providers, for example, separate departments from the same hospital, or between hospitals. In these cases, an admission is a phrase, and an episode is a subset of phrases separated by a transfer event. We create a special word TRANSFER for this event. Between two consecutive episodes, there is a time gap, whose duration is generally randomly distributed. We discretize the time gap into five intervals, measured in months: (0-1], (1-3], (3-6], (6-12], 12+. Each interval is assigned a unique identifier, which is treated as a word. For example, 0-1m is a word for the (0-1] interval gap. With these treatments, an EMR is a sentence of phrases separated by words for transfers or time gaps. The phrases are ordered by their natural time-stamps. For robustness, infrequent words are coded as RAREWORD.

The following is an example of a sentence, where diagnoses are in ICD-10 format (a character followed by digits), and procedures are in digits: 1910 Z83 911 1008 D12 K31 1-3m R94 RAREWORD H53 Y83 M62 Y92 E87 T81 RAREWORD RAREWORD 1893 D12 S14 738 1910 1916 Z83 0-1m T91 RAREWORD Y83 Y92 K91 M10 E86 6-12m K31 1008 1910 Z13 Z83.

Here the phrases are: [1910 Z83 911 1008 D12], [R94 RAREWORD H53 Y83 M62 Y92 E87 T81 RAREWORD RAREWORD 1893 D12 S14 738 1910 1916 Z83], [RAREWORD Y83 Y92 K91 M10 E86], and [K31 1008 1910 Z13 Z83]. The time separators are: [1-3m], [0-1m], and [6-12m]. Note that within each phrase, the ordering of words has been randomized.

III-C Convolutional Net

Convolution

On top of the word embedding layers is a convolutional layer. Each convolution operation reads a sliding window of size 2d+12d+1 and produces pp filter responses as follows:

Pooling

Once the local filter responses are computed by the convolutional layer, we need to pool all the responses to derive a global sentence-level vector. We apply here the max-pooling operator:

Classifier

The final layer of Deepr\mathtt{Deepr} is a classifier that takes the pooled information and predicts the outcome: f(r)=classifier(zˉ(r))f(r)=\text{classifier}(\bar{\boldsymbol{z}}(r)) for record rr. The main requirement is that the classifiers must allow gradient to propagate down to lower layers. Examples include a linear classifier (e.g., logistic regression) or a non-linear parametric classifier (e.g., neural network).

III-D Training

As mentioned in Sec. III-A, the embedding matrix can be pretrained using word2vec. Here we do not need labels, and thus we can exploit a large set of unlabeled data.

III-E Model Inspection and Visualization

Deepr\mathtt{Deepr} facilitates intuitive model inspection and visualization for better understanding:

Identifying frequent and strong motifs

Motifs with large responses in sequences are collected. From this collection, we keep frequent motifs representative for each outcome class.

Computing word similarity

Through embedding xw=E(w)\boldsymbol{x}_{w}=E(w), word similarity can be computed easily, e.g., through cosine S(w,v)=xw⊤xv(∥xw∥∥xv∥)−1S(w,v)=\boldsymbol{x}_{w}^{\top}\boldsymbol{x}_{v}\left(\left\|\boldsymbol{x}_{w}\right\|\left\|\boldsymbol{x}_{v}\right\|\right)^{-1}.

Visualization of similar patients

Patient vectors from Eq. (3) can be used to compute patient similarity. This enables retrieving patients who have similar history and similar future risk likelihood. This is unlike existing methods that compute only similar history, which does not necessarily guarantee similar future. Further, the similarity is not heuristic, and it does not require a heuristic combination of multiple data types (such as diseases and interventions). Fig. 2, for example, shows the distribution of positive and negative classes, in which patient vectors are projected onto 2D using t-SNE . Patients who have similar history and future will stay close together.

Visualization in disease/intervention space

Since words are embedded into vectors, visualization in 2D is through dimensionality reduction tools such as PCA or t-SNE .

IV Implementation

In this section, we document implementation details of Deepr\mathtt{Deepr} on a typical EMR system. For ease of exposition, we assume that diseases are coded in ICD-10 format, but other versions are also applicable with minimal changes.

Data was collected from a large private hospital chain in Australia in the period of July 2011 – December 2015. The data is coded according to Australian Coding Standard (ACS). The ACS dictates that diagnosis coding is based on ICD-10-AMhttps://www.accd.net.au/Icd10.aspx, an Australian adaptation to WHO’s ICD-10 system. Likewise, procedure coding follows ACHI (Australian Classification of Health Interventions). The data consists of 590,546 records (300K unique patients), each corresponds to an admission (defined by an admission time and a discharge time).

The data subset for testing Deepr\mathtt{Deepr} was selected as follows. First we identified 4,993 patients who had at least an unplanned readmission within 6 months from a discharge, regardless of the admitting diagnosis. This constituted the risk group. For each risk case, we then randomly picked a control case from the remaining patients. For each risk/control group, we used 830 patients for model tuning, 830 for testing and the rest for training. A discharge (except for the last one in risk group) is randomly selected as prediction point, from which the future risk will be predicted. See also Fig. 1 for a graphical illustration.

IV-B Implementation Details of 𝙳𝚎𝚎𝚙𝚛𝙳𝚎𝚎𝚙𝚛\mathtt{Deepr}

Deepr\mathtt{Deepr} assumes that episodes are well-defined with an admission time and discharge time. However, it is not always the case due to intra-hospital or inter-hospital transfers. Our implementation links two admissions into an episode if they are separated by less than 12 hours, or by 12-24 hours but with documented transfer.

Words

For robustness, only level 3 ICD-10-AM codes are used. For example, F20.0 (paranoid schizophrenia) would be converted into F20 (schizophrenia). Similarly, the procedures are converted into procedure blocks. Rare words are those occurring less than 100 times in the database.

Word order randomization

For motifs detection, randomization is necessary to generate many potential motifs. We also test a special case where words in a phrase are ordered starting with the primary diagnosis followed by other secondary diagnoses, then by procedures in their natural ordering as defined by the EMR system.

Sentence length

For CNN, the sentences are trimmed to keep the last min(100, len(sentence)) words. This is to avoid the effects of some patients who have very long sentences which severely skew the data distribution. In a typical EMR, this is equivalent to accounting for up-to 10 visits per patient, which cover more than 95% of patients.

Hyper-parameter tuning

Deepr\mathtt{Deepr} has a number of hyper-parameters pre-specified by model users: embedding dimension mm, kernel window size 2d+12d+1, motif size, number of motifs nn per size, number of epochs, mini-batch size, and other classifier-specific settings. Some hyper-parameters can be found through grid search, which finds the best configuration with respect to the accuracy on the development set.

IV-C Baselines

We implemented the bag-of-words representation and regularized logistic regression (BoW+LR). LR has a parameter CC that helps control overfitting. We searched for the best parameter CC using the development data. We used the model with the best parameter to predict the unseen test data. We found the best parameter C=0.1C=0.1, which is equivalent to a prior Gaussian of mean and standard deviation of 0.3330.333.

V Results

We predict unplanned readmission within 6 months after a random index discharge. Table I reports the prediction accuracy for all methods, when trained on data with and without coded time-gaps. Time-gaps coding improves the BoW-based prediction, suggesting the importance of proper sequential handling. However, time-gaps do not affect the accuracy of Deepr\mathtt{Deepr}. This might be due to the convolution, rectification and max-pooling operations (see Sec. III-C), which pick the most powerful convoluted signals in the sequence. The use of word2vec to initialize the embedding matrix also has little contribution toward the accuracy. This could be because word2vec looks only for local collocations in both directions (past and future), whereas the prediction in Deepr\mathtt{Deepr} is more global and of longer time horizon only in the future direction. In either cases with and without word2vec, Deepr\mathtt{Deepr} is superior than the baseline BoW+LR.

Fig. (2) shows how Deepr\mathtt{Deepr} groups similar patients and creates a more linear decision boundary while BoW+LR scatters the patient distribution and has a more complicated decision boundary. Recall that Deepr\mathtt{Deepr} creates the feature vectors using element-wise max-pooling over all the motifs responses, as in Eq. (3). This demonstrates that the motifs, not just individual words, are important to computing similarity between patients. This also suggests that given a new patient Deepr\mathtt{Deepr} is better at querying similar patients in the database when future risk is needed.

V-B Disease/Procedure Semantics

Recall that Deepr\mathtt{Deepr} first embeds words into a vector space. This offers a simple but powerful way to uncover and visualize the underlying structure of the word space (see Sec. III-E). Fig. 3 plots the distribution of diseases on 2D. Deepr\mathtt{Deepr} discovers disease clusters which partly correspond to nodes in the ICD-10 hierarchy. Apart from pregnancy, child birth issues and injuries, the conditions are not totally separately suggesting a complex dependencies in the disease space. The main bock of the disease space has conditions related to heart, blood, metabolic system, respiratory system, nervous system and mental health. A more close examination of most similar conditions to a disease is given in Table II. For example, similar to cesarean section delivery of baby are those related to pregnancy complications (disproportion, failed induction of labor, or diabetes) and corresponding delivery procedures (cesarean section, manipulating fetal presentation, forceps).

We note in passing that we also obtained a similar visualization using only word2vec as in , which is known to detect hidden semantic relationships between words. Deepr\mathtt{Deepr} trained on the embedding matrix initialized by word2vec did not significantly change the relative positions of words. This suggests that Deepr\mathtt{Deepr} also captures the semantic relationship between words.

V-C Filter Responses and Motifs

While the semantics in the previous sub-section reveal the global relative relation between diseases and procedures, they do not explain local interactions (e.g., motifs). Here we compute the local filter responses per sentence, and from there, a collection of strong and frequent motifs is derived.

Table III shows some sentences with strong responses for Filter 1 and 4 for both risk and no-risk class. It can be seen that the sub-sequences Z85.1163.1910 and 1066.1067.I21 respond strongly for the positive class and contribute to the classification result. The first sub-sequence is about cancer history (Z85), biopsy procedure (1163) and cerebral anesthesia (1910). The other sub-sequence is about heart attack (I21) and kidney-related procedures (1066 and 1067).

From strong and frequent filter responses in all sentences, we derive the list of motifs. Table IV lists the motifs with largest weights and highest frequency of occurrence for code chapter E, I and O. The first motif of Filter 45 shows the pattern that treatment removing toxic substances from the blood co-occurred with care involving dialysis and readmission within 1 month. The second motif in the same row discovers the pattern that type-I diabetes patients involve in education about information and management of diabetes. The third motif in the same row shows type-II diabetes patients readmit within 1-3 months. Filter 26 demonstrates the co-occurrence of diseases related to diabetes. The three motifs show that type-II diabetes patients can have complications such as heart failure, vitamin D deficiency and kidney failure. Filters 10 and 35 show diseases and treatments related to the circulatory system, whereas pregnancy and birth related motifs are shown in Filters 2 and 33 in the last two rows.

VI Discussion

We have presented Deepr\mathtt{Deepr}, a new deep learning architecture that provides an end-to-end predictive analytics in healthcare services. Deepr\mathtt{Deepr} reads directly from raw medical records and predicts future outcomes. This departs from the traditional machine learning that relies on expensive manual feature extraction. Deepr\mathtt{Deepr} learns to extract meaningful features by itself without expert supervision. This translates to uncovering the predictive local motifs in the space of diseases and interventions. These capacities are not seen in existing methods.

Deepr\mathtt{Deepr} contributes to the growing literature of predictive medicine in multiple ways. First, it is able to uncover the underlying space of diseases and interventions, showing the relationships between them. The largest disease cluster in Fig. 3 suggests that diseases may interact in a complex way, and current representation of disease hierarchies such as those in ICD-10 may not reflect the true nature of medical disorders. Second, Deepr\mathtt{Deepr} detects predictive motifs of comorbidity, care patterns and disease progression. The motifs suggest a new look into the complex interactions between diseases and between the diseases and cares. Third, similar patients can be retrieved not just using past history, but from likelihood of future risks as well. This would, for example, help to quickly identify an effective treatment regime based on similar patients who responded well to the treatment, or to alert the care team of a potential risk based on similar patients who had these before. Finally, Deepr\mathtt{Deepr} predicts the future risk for a patient and explains why (through means of motifs responses), which is the core of modern prospective healthcare.

With these capabilities, Deepr\mathtt{Deepr} can enable targeted monitoring, treatments and care packaging. This is highly important for chronic disease management that requires an on-going care and evaluation. For health services, a high predictive accuracy of risk will lead to better resources prioritizing and allocation. For patients, accurate risk estimation is an important step toward personalized care. Patients and family will be promoted to become more aware of the conditions and risk, leading to proactive health management and help seeking. Deepr\mathtt{Deepr} is generic and it can be implemented on existing EMR systems. This will enable innovative healthcare practices for better efficiency and outcomes to occur. For example, doctors, when seeing a patient, may consult the machine for a second opinion, with a transparent, evidence-based reasoning. Because they do not miss any piece of information in the database, they are less likely to overlook important signals.

Comparison to recent work on medical records

Deep learning in healthcare has recently attracted great interest. The most popular application is medical imaging using CNNs , motivated by the recent successes in cognitive vision . However, there has been limited work on non-cognitive modalities. On time-series data (e.g., ICU measurements), the main difficulty is the handling missing data with recent work of . In , time-series are modeled using autoencoders (an unsupervised feedforward net) to discover meaningful phenotypes. In , recurrent nets are used, and in , a convolutional net is employed. Deepr\mathtt{Deepr} can be applied on these data, following a discretization of continuous signals into discrete words (e.g., through cut-points).

On routine medical records, Deepr\mathtt{Deepr} is the only method that employs convolutional nets but there exist alternative architectures. Feedforward nets have been used . Recurrent neural networks (RNN) on medical records include Doctor AI and DeepCare . Doctor AI is a RNN adapted for medical events, where both next events and time-gaps are predicted. DeepCare is a sophisticated model that represents time-gaps using a parametric model. Similar to our observation, the authors of DeepCare also noticed an interesting analogy between natural languages and EMR, where EMR is similar to a sentence, and diagnoses and interventions play the role of nouns and modifiers. While DeepCare is powerful on long records, it is less effective in short records, e.g., those with only one or two admissions. Deepr\mathtt{Deepr}, on the other hand, does not suffer from this limitation. Stochastic deep neural nets such as deep Boltzmann machines are used in . Deep non-neural nets have also been suggested in . These methods are likely to be expensive to train and produce prediction.

Embedding of medical concepts has been proposed in contemporary work . In , medical concepts are embedded using word2vec , ignoring time gaps. The Med2Vec in extends word2vec to embed visits. Both word2vec and Med2Vec model local collocations, but do not explicitly model motifs (with precise relative positions). In , a global model known as eNRBM embeds patients into vectors via regularized nonnegative restricted Boltzmann machines . Local motifs are not modeled and and variable record length and time gaps are not properly handled. Discovering local motifs by means of convolutions has been suggested in through matrix factorization. However, the work does not do prediction.

Limitations and future work

There are rooms for future work. First, long-term dependencies are simply captured through a max-pooling operation. This is rather simplistic due to a complex dynamic between care processes and disease processes . A better model should pool information that is time-sensitive (e.g., recent events are more important to distant ones). At present, Deepr\mathtt{Deepr} works exclusively on recorded events such as diagnoses and interventions. Integration with clinical narrative would be highly useful because rich information is buried in unstructured text. This can be done in the same framework of Deepr\mathtt{Deepr} because of the sequential nature of text. Our evaluation has been limited to a common risk known as unplanned readmission. However, Deepr\mathtt{Deepr} is not limited to any specific type of future risk. It can be well applied to predicting the onset or progression of a disease.

References