Gesture Recognition in Robotic Surgery: a Review

Beatrice van Amsterdam, Matthew J. Clarkson, Danail Stoyanov

I Introduction

Robot-assisted minimally invasive surgery (RMIS) is now well established in clinical practice as a means of enhancing surgical instrumentation and ergonomics with respect to conventional laparoscopic approaches . Surgical specialisations like gynecology and urology have driven the adoption of RMIS, but diffusion of the robotic approach is increasing as hospitals become equipped with surgical robotic systems, surgeons trained to use them and both the capital and running costs reduce . In addition to better surgical instrumentation, surgical robots can potentially be platforms for enabling data driven solutions for better OR capabilities .

Besides digital video, surgical robots can capture quantitative instrument motion trajectories during surgical interventions enabling analysis of surgical activity that is not possible with traditional instrumentation. These new data streams have paved the way for data-driven computational models in computer assisted interventions (CAI) and surgical data science (SDS) to facilitate better procedural understanding and intra-operative support, including the possibility of linking back to the robot control loop to achieve surgical automation. The resultant data has also allowed research into automated surgical competence and skill assessment and adaptive surgical training , aiming at quantifying surgical process without expert monitoring.

Understanding surgical process has often been approached by decomposition into pre-defined procedural segments at different granularity levels. Starting from the finest level, surgical processes have been hierarchically decomposed into dexemes, surgemes (generally called gestures, Fig. 1), activities, phases, procedures and states . Unfortunately, the lack of a universal taxonomy and of clear segmentation criteria has led researches to propose different decomposition strategies (e.g. semantic vs event-based labelling) and alternative motion dictionaries , resulting in conceptual inconsistencies and a lack of standardised terminology.

Segmentation at different levels highlights distinct aspects of the surgical procedure, leading to a range of applications. For example, higher-level situation-awareness (state, phase or activity recognition) is well suited for workflow optimization, scheduling and resource management . Recognition of fine-grained motion (dexemes, surgemes) is instead better suited for low-level analysis, assessment and reproduction of dexterous motion. While dexemes describe short motion segments devoid of medical sense (e.g. turning left) , surgemes represent surgical gestures made with a specific purpose (e.g. grabbing the needle) and are the focus of this review. Kinematic trajectory segmentation into gestures has been shown to have potential to describe technical skill development in a more quantitative manner than relying on global performance measures as time to completion or total path length . Gesture recognition also finds application in surgical automation because short motion segments are less complex and therefore easier to learn and generalize than long surgical tasks, and can be regarded as modular blocks of motion to compose and reuse .

Fine-grained analysis of surgical demonstrations is however non-trivial and presents a number of challenges. When generated manually, gesture annotations are costly, time consuming and subject to personal interpretation of multiple participants. Due to subjectivity and smooth transitions between gestures, boundaries between consecutive segments are often not clearly defined and it is difficult, even for a single annotator, to generate consistent segmentations as data grow over time. For this reason significant research effort is aimed at automating the annotation process utilizing both video and instrument kinematics, either individually or in combination . Automation remains very challenging due to the complexity of surgical tasks and the presence of different variability factors, including surgeon’s skill level, operative style, type of procedure and patient-specific anatomy, but could help to efficiently generate the large amount of training data required for robust CAI implementations. Beyond off-line generation of training datasets for surgical automation and training, automatic recognition can be exploited for online applications where context-awareness is essential, such as systems for surgical process monitoring, error detection and intra-operative guidance and assistance of different kinds (e.g. robotic, visual, haptic) . This will have a potentially high impact in robotic surgery, where demands for increased safety and effectiveness are particularly strong .

In this paper, we review existing methods for automatic recognition of fine-grained gestures in RMIS, illustrating how technical challenges have been addressed, what problems still need to be investigated and what could be the future research directions in the field. Comprehensive reviews on higher-level surgical workflow recognition have recently been reported and are outside the scope of this study .

II Methodology and structure

We used Google Scholar, Scopus, PubMed, Web of Science and IEEEXplore for searching the literature titles, abstracts and keywords with the query: “(robotic OR robot-assisted OR JIGSAWS) AND (surgery OR surgical) AND (gesture OR fine-grained OR surgeme OR action OR trajectory) AND (segmentation OR recognition OR parsing) ”, where JIGSAWS represents the first open source dataset for surgical gesture recognition (see Section VII-A). As Google Scholar only allows full-text searches, the 300 most relevant publications were extracted from the 47.100 retrieved items. Integrating the 5 databases and discarding all duplicates, non-English texts and full conference proceedings, our query delivered a total of 499 matches in October 2020.

We adopt the term “recognition” to refer to the joint problem of “segmentation” and “classification” of surgical gestures, i.e. the simultaneous segmentation of surgical demonstration into distinct motion units and the classification of these units into meaningful action categories. While computer-vision methods for classifying videos with known action boundaries have been successful, temporal segmentation of multiple action instances is still a hard-won problem. Classification of trimmed videos also finds limited applicability in surgical scenarios due to gesture annotation costs, thus research has been fairly limited (only 8 studies were retrieved with our query) and is not presented in this review. We also excluded articles focused on coarser (“higher-level”) granularity and studies that use non-novel recognition methods for skill assessment or surgical automation. After a selection process as illustrated in Fig. 2, we identified a total of 52 publications to be reviewed.

We use the field decomposition shown in Fig. 3 to organise our analysis. Supervised approaches are classified into two groups representing major frameworks for time series analysis and data modelling: probabilistic graphical models (Section III) and deep learning (Section IV). We then describe efforts to mitigate the need for manual annotations (Sections V and VI). A subdivision of the methods by data modality is instead highlighted in Tables II, III, IV and V of Section VII.

III Supervised learning - graphical models

Inspired by time series analysis for speech recognition, structured temporal models such as probabilistic graphical models of temporal processes have been extensively used for gesture recognition in robotic surgery . Similarly to the natural language, where syntactic rules govern the generation of words and sentences, the surgical language can be decomposed into surgical motion units at different granularity levels, and probabilistic grammars can be identified to describe the structure of higher-level surgical tasks (Fig. 4).

There is often significant flexibility in choosing the structure and parameterization of a graphical model. Given a sequence of observations and corresponding output labels, generative models such as Hidden Markov Models (HMM) and Linear Dynamical Systems (LDS) explicitly attempt to capture the joint probability distribution over inputs and output. Discriminative models like Conditional Random Fields (CRF), on the other hand, attempt to model the conditional distribution of the output labels given the input data. The advantage with respect to generative models is that dependencies in the input features don’t need to be modelled, so that conditional models can have simpler structure .

HMMs were first used to classify motion segments obtained by thresholding the kinematic signals, which had limited recognition rate but could be significantly improved with appropriate data transformation including normalization to zero-mean and unit-variance, and Linear Discriminant Analysis (LDA) . Normalization compensates for different units of measure in the kinematic features, while LDA reduces the data dimensionality and model complexity, which is important for robust parameter estimation when only little data is available. LDA’s ability to enhance the separation between classes is another desirable property for fine-grained gesture recognition, where different classes exhibit similar motion and visual appearance.

Highly discriminative feature extraction is key to the recognition problem, so that even simple frame-wise classifiers, such as the Naïve Bayes classifier, can provide high recognition accuracy . Following , LDA was successfully applied to the high-dimensional kinematic signals recorded from the da Vinci surgical robot in combination with a continuous HMM for gesture recognition, composed of basic HMMs representing different surgemes and trained individually . Recognition rate of such models further improves promoting discrimination between sub-gestures (dexeme) rather than entire gestures (surgeme), as this helps to capture the evolution and internal variability of each segment . Automatic derivation of the optimal HMM topology (i.e. the optimal number of hidden states) can also help to discover skill-related and context-specific sub-gestures. While yielding accurate recognition on new trials of seen users, data-derived HMMs may also introduce user-specific sub-gestures and therefore overfit on observed subjects .

Alternative data representation and reduction techniques investigated for recognition include Factor-Analyzed Hidden Markov Models (FA-HMM), which yield better recognition compared to standard HMMs with LDA . Sparse Hidden Markov Models (SHMM) were also applied to the recognition of surgical gestures . Available implementations of SHMM represent each surgeme with a different hidden state, whose observations are modelled as sparse linear combinations of atomic motions. The advantage with respect to conventional discrete or Gaussian observation models relies in the sparsity property, which guarantees more robust learning with high-dimensional signals such as the robot kinematics. While each surgical gesture can be associated to a different dictionary that is learnt independently , more compact solutions propose unique motion dictionaries shared among all gestures .

Finally, interaction forces measured between the robotic tool tip and the environment have been employed as complementary data modality to the robot kinematics . Force information allows one to assign physical interpretation to HMM state observations.

III-B Linear Dynamical Systems

A fundamental drawback of the HMM is its reliance on discrete switching states to model temporal dependencies in the input signals, thus failing to capture more complex correlations in the continuous kinematic and video data. Switched Linear Dynamical Systems (S-LDS) were then proposed to model temporal dependencies between adjacent observations . S-LDSs approximate non-linear systems, such as demonstrations of complex surgical tasks, using a suitable number of linear systems. The data distribution within each gesture is thus modelled through the linear evolution of a continuous hidden state, while the transition between different regimes relies on a discrete switching hidden state. Thanks to improved dynamics modelling, recognition with S-LDS outperformed FA-HMM significantly .

S-LDS models can be further extended to capture the dynamics of past observations samples, thus becoming Switched Vector Auto-Regressive models (S-VAR). By imposing specific structure on the model parameters, with certain parameters depending on the discrete latent state and others which are time-invariant, S-VAR models can learn both gesture-dependent and global dynamics in the kinematic signals, resulting in improved recognition accuracy .

III-C Conditional Random Fields

CRFs were first used to recognize surgical gestures in combination with video data. Inspired by related work on video-based surgeme classification , bag of spatio-temporal visual features extracted around space-time interest points or dense trajectories were combined with the kinematic signals in input to a Markov/semi-Markov recognition model (MsM-CRF), which is able to capture frame-level and segment-level dependencies in the input signals .

Long-range temporal dependencies were also modelled with the Skip-Chain conditional random field (SC-CRF) , incorporating a new skip-length data potential into the CRF energy function to capture sample dependencies at d frames of distance, where d is a skip-length hyper-parameter. In addition, a set of semantic features were extracted from the laparoscopic videos based on the distance of each surgical tool to the closest object in the surgical environment. The authors argued that in constrained environments such as surgical training stations, where the set of objects is known, semantic features encoding object-based information could help to discriminate surgical gestures better than abstract visual features as in .

As for HMMs, efforts were spent on improving data representation and feature extraction for CRFs. Inspired by , kinematic signals were modelled as sparse linear combinations of atomic motions from an over-complete motion dictionary shared among all gestures . The motion dictionary was learnt jointly with a SC-CRF to capture temporal dependencies and generate more discriminative signal representations, which is valuable to fine-grained analysis. Similarly to , surgical kinematics were also analysed at the sub-gesture level. Temporal convolutional filters were used to model the evolution of “latent action primitives” composing each surgical gesture and capture variability within surgemes . The real kinematic data were then replaced with virtual kinematic signals estimated from the video data through a spatial CNN, addressing the situation where kinematic information is available during training, but only video data are available for testing . Convolutional filters were later employed in a similar but more powerful multi-stage framework based on deep convolutional neural networks for spatio-temporal feature extraction from the raw video data (ST-CNN) . Nevertheless, feature extraction strategies as in and are only able to capture local temporal correlations in the observations. Modelling of long-range dependencies is entirely entrusted to the high-level temporal classifier (SC-CRF and sM-CRF respectively). A summary of the presented graphical models for surgeme recognition is reported in Table I.

IV Supervised learning - deep learning

Recently, Deep Neural Networks (DNN) have emerged as the go-to choice for powerful feature extractors that can be learnt automatically from the data, in contrast with hand-crafted filters based on domain-specific prior knowledge . In DNNs, layers of computational units are stacked to generate features at different semantic levels and with increased spatio-temporal receptive fields. This is advantageous for action recognition because long-range spatial and temporal dependencies encoded in the input data can be captured already in early processing stages, easing the subsequent recognition task.

Temporal Convolutional Network (TCN) is one of the first deep convolutional models proposed for fine-grained surgeme recognition . TCN uses a hierarchy of temporal convolutions and pooling layers to capture long-range temporal dependencies in the input data (Fig. 5a), thus learning action dependencies in surgical demonstrations better than competing methods and producing well structured predictions. Integration of bilinear pooling with learnable weights further boosts recognition performance incorporating second-order statistics adaptable to the data . As opposed to RNNs, TCN predictions are computed simultaneously for each time stamp and training is much faster.

Built on a similar encoder-decoder backbone, more recent solutions rely instead on two-stream processing: one stream to capture multiscale contextual information and action dependencies in the input sequence, the other to extract low-level information for precise action boundary identification . Basic temporal convolutions were replaced with a set of Deformable Temporal Residual Modules (TDRM) to capture the temporal variability of human actions and merge the two streams at multiple processing levels . Alternatively, atrous (aka dilated) temporal convolutions and pyramid pooling were used to explicitly generate multi-scale encoding of the input data, that were merged with the low-level local features just before the decoding phase . Both architectures were able to improve consistently both frame-wise and segmental evaluation scores (see Section VII for a description and explanation of the most common evaluation metrics) upon the competing methods.

While pooling operations help to increase the temporal receptive field of a network, they are also responsible for partial loss of fine-grained information and less precise identification of the gesture boundaries. Stacking multiple layers of dilated convolution with increasing dilation factor was proposed as alternative strategy to model long range temporal dependencies between surgical gestures . Gesture predictions can be further refined in a multi-task, multi-stage framework where each stack of dilated convolutions is applied to the output of the previous stage, and the whole system is trained on the sum of all stage losses, as well as on the auxiliary task of surgical skill score prediction . Alternatively, two convolutional stages can be bridged with a self-attention kernel to encode local and global temporal dependencies in the input signals, returning highly accurate predictions with little training cost .

All aforementioned architectures represent high-level temporal models that can process efficiently only low-dimensional signals such as kinematic data or pre-encoded visual features. The latter are often derived from the original video frames through spatial encoding techniques , albeit unable to capture the dynamics in surgical demonstrations. To address this issue, spatio-temporal feature extraction strategies from short video snippets were proposed using 3D CNNs or Spatial Temporal Graph Convolutional Networks (ST-GCN) . While 3D CNNs operate directly on the video data, achieving superior recognition performance to related spatial and spatio-temporal models, ST-GCN are applied on spatio-temporal skeleton representations of the surgical tools, which are automatically estimated from the video clips. Skeleton-based features are independent of the background and thus robust to scene and context variations, making them suitable for recognition across different surgical tasks or datasets. Clip-based methods like 3D CNNs and ST-CNNs have however limited span of attention, which prevents them from learning long-range temporal dependencies, yielding suboptimal segmental scores. Integration with higher-level convolutional or recurrent models could further boost segmentation performance.

Finally, efforts were spent on improving recognition robustness to inter-surgeon data variability, associated with heterogeneity factors such as surgical style and expertise level, which could hinder generalization performance on unobserved subjects (see Sections VII-C and VIII-B for more details). Robustness to the subject’s identity was obtained via adaptive feature normalization , where the feature statistics are estimated over data sub-sets recorded from the same surgeon aiming to reduce inter-surgeon data variability and ease the recognition process.

IV-B Recurrent Neural Networks

Recurrent Neural Networks (RNN) have been used to capture long-term non-linear dynamics in surgical kinematic data . In-depth analysis and comparison of different architectures (simple Bidirectional RNNs, Bidirectional LSTMs, Bidirectional GRUs and Bidirectional Mixed History RNNs) showed that LSTMs and GRUs, which were conceived to alleviate the vanishing gradient problem and can be trained more robustly, are less sensitive to hyper-parameter choice and achieve the best recognition performance .

Since predictions are computed sequentially in time, RNNs can naturally handle input signals of different duration and can operate in real time. They also maintain memory mechanisms so that predictions depend in principle on all the previous time stamps (Fig. 5b). Their temporal span of attention, however, is in practice much shorter and they fail to capture multi-scale temporal dependencies in the observations . To address this issue, Multi-Scale Recurrent Neural Network (MS-RNN) extracts multi-scale temporal information from the kinematic signals by means of hierarchical convolutions with wavelets, which are then given as additional inputs to a LSTM for gesture recognition. MS-RNN is trained end-to-end to generate compact descriptors and lessen LSTM susceptibility to overfitting.

Alternative approaches have been recently proposed to regularize training of recurrent architectures and avoid overfitting, such as using data augmentation to improve the network robustness to rotation or multi-task learning to jointly recognize surgical gestures and estimate the progress of the surgical task, in order to introduce ordering relationships and temporal context explicitly into the recognition process .

IV-C Hybrid Convolutional-Recurrent Networks

Temporal Convolutional and Recurrent Network (TricorNet) represents the first hybrid network combining convolutional and recurrent paradigms to recognise surgical gestures. TricorNet features a temporal convolutional encoder similar to to learn local motion evolution, followed by a hierarchy of bidirectional LSTMs to capture long-range temporal dependencies and decode the prediction sequence. Since recurrent cells are in principle able to carry long-term memory of the past observations, the authors argued that hybrid solutions could better model high-level action dependencies than purely convolutional systems with fixed-size local receptive field.

Rather than combining different models sequentially, Fusion-KV performes multi-modal learning with multiple recognition models (CNN-TCN, TCN, LSTM, Random Forest, Support Vector Machine) operating in parallel on three different data sources: kinematics, video and system events collected from the robotic platform (e.g. camera follow, instrument follow, surgeon head in/out of the console, etc). Predictions from all streams are fused through a weighted voting scheme that takes into account recognition performance of each individual stream on the training data. The strengths of different methods and data are thus combined to deliver improved recognition performance. Fusion-KV was later incorporated in a mult-task sequence-to-sequence framework for surgical gesture and instrument trajectory prediction, where gesture labels and instrument positions are estimated over multiple steps in the future given a past observation window . Compact features are first extracted from multiple data sources (full video frames, region-of-interest around surgical tools and kinematic data) and concatenated into a single multi-modal feature vector. Multi-modal features for an observation window in the past and corresponding gesture labels estimated with Fusion-KV are then given as input to an attention-based sequence-to-sequence decoder for gesture prediction, showing 10-step prediction rates comparable to state-of-the-art recognition (0-step prediction) models.

IV-D Deep Reinforcement Learning

Surgeme recognition can alternatively be modelled as a sequential decision-making process that can be learnt with Reinforcement Learning (RL), as recently proposed by . In their work, an intelligent agent observes each surgical sequence from the beginning. At each time stamp, it selects an appropriate step size and gesture label according to a specific policy, and then moves step-size frames ahead, classifying all the frames in-between as the selected gesture (Fig. 5c). The reward function for policy learning is driven by both classification error and chosen step size, introducing a bias towards larger steps to promote prediction smoothness. While the system was able to achieve competitive results with respect to several state-of-the-art techniques, it sometimes delivered uncertain predictions, especially for rare gestures or at the segmentation boundaries. A tree-search component was therefore introduced to refine the low-confidence predictions from the policy network by looking into the future, improving recognition of the action boundaries .

V Unsupervised learning

Because of the lack of labelled examples, unsupervised approaches are mostly focused on the segmentation sub-problem, i.e. the identification of motion unit boundaries within surgical demonstrations.

Transition State Clustering (TSC) , a multilevel generative model based on hierarchical Gaussian Mixture Model (GMM), was applied to surgical kinematic data to identify segmentation points pruning spurious segments generated by inconsistent motion, under the assumption that action order is partially consistent across surgical demonstrations. The kinematic signals were then integrated with visual features extracted with a pre-trained CNN, improving segmentation performance . TCN methodology was later extended using Dense Convolutional Encoder-Decoder Netowork (DCED-Net) for unsupervised feature extraction from surgical videos . Trained on the task of image reconstruction, DCED-Net extracts discriminative visual features using specific convolutional blocks, called Dense Blocks, to reduce information loss during dimensionality reduction. Segmentation results were further improved with a majority vote strategy to eliminate spurious transition points exploiting both kinematic and visual information, and an iterative algorithm to merge similar adjacent segments .

More complex generative models were later proposed to describe the joint probability distribution of observed kinematic data and latent action sequences. Surgical demonstrations were for instance modelled as Markov decision processes learnt through imitation learning with options , where complex policies are broken down into simpler low-level primitives called options. The model parameters were estimated with a policy-gradient algorithm aimed to maximize the likelihood of the available data and simultaneously segment them into different motion primitives. Alternative solutions such PRISM (PRocedure Identification with a Segmental Mixture model) were developed under the assumption that all demonstrations are described by a common latent sequence of actions. While generalizable to non-Markov processes, PRISM performance degrades significantly with increased action ordering variability.

Other clustering techniques group the data samples based on feature similarities and temporal constraints. Spectral clustering was tailored to temporal segmentation through specific regularization terms considering local dependencies in the motion data, where consecutive samples are likely to belong to the same action class . Good results were obtained with bottom-up clustering, where the sequential nature of motion signals can be exploited by merging neighbouring segments according to various merging and stopping criteria . In , surgical demonstrations were first segmented into a set of very fine-grained motion primitives (dexeme) based on the persistence and dissimilarity of consecutive segments between critical points in the kinematic trajectories. A descriptive signature was then created to represent each dexeme and used for classification into surgeme categories. Starting instead from the finest possible segmentation of the kinematic signals , neighboring segments were merged iteratively using multiple compaction scores based on different distance measures (similarity between PCA results, distance between segment centres, dynamic time warping distance) and fuzzy membership scores to model gradual transitions between surgical gestures.

Rather than grouping similar data samples or segments together, opposite approaches search for optimal segmentation criteria. Spatio-temporal and variance properties of the kinematic data can be exploited to identify candidate segmentation points around the peaks of distance and variance profiles computed from the original signals . The method’s disadvantage, however, is in the tedious tuning of threshold values.

Finally, zero-shot learning (ZSL) approaches have also shown potential for surgical gesture recognition . ZSL consists in learning a set of high-level semantic attributes describing instances of specific action classes and exploiting them to recognize instances of new classes, provided with their attribute description (signature) but no labelled examples. modelled surgical gestures with dynamic signatures represented by a set of attributes that change in time according to action-specific rules. Given a surgical demonstration, the presence or absence of those attributes were detected on each video frame and compared to the set of signatures for action matching, achieving better results than fully supervised unstructured baseline models.

VI Semi-supervised learning

While unsupervised approaches mitigate significantly labelling costs and issues, as labels are only necessary for model testing, they still show a significant gap in performance with respect to the supervised methods. A midway solution is represented by semi-supervised learning , where a small pool of annotated demonstrations is used for model training. Such demonstrations can be used to transfer gesture labels to a large set of unlabelled observations previously aligned with DTW , or to initialize clustering of unlabelled data and avoid tedious parameter tuning . These systems are however highly sensitive to large data variability, showing rapid drop in recognition accuracy with increased heterogeneity in terms of gesture ordering or surgical skill level respectively.

Promising results were obtained with RNNs within a two-phase pipeline comprising unsupervised representation learning and semi-supervised recognition . An RNN-based generative model was first trained to learn the full joint distribution over the kinematic signals, an auxiliary task that needs no additional annotations but can extract informative features from the input data. These features were then used by a bidirectional LSTM for gesture recognition, which was only trained on a very small number of labelled demonstrations. Decoupling between representation learning and gesture recognition was later solved via semi-supervised interleaved training . Meaningful visual embeddings were learnt using triplet loss to pull together images in the same action segment and push away samples from different segments, where action pseudo-labels for unlabeled training sequences were iteratively predicted by the current estimate of a recurrent network trained via cross-entropy loss. Thanks to robust representation learning, the gap in performance with respect to equivalent fully supervised systems was limited.

Semi-supervised learning was also explored in combination with transfer learning between surgical tasks . One-dimensional projections of the kinematic signals derived from their Self-Similarity Matrix (SSM), which preserves motion patters in the data, were used to learn segmentation policies of specific surgical tasks (e.g. knot-tying) from annotated demonstrations of other tasks (e.g. suturing). In contrast to , the class of the generated segments was not predicted.

VII Benchmark datasets and evaluation protocols

Early techniques for surgical gesture segmentation and skill assessment were validated on study-specific data using different evaluation metrics, hindering comparative research and progress tracking in the field. To address this issue, the JHU-ISI Gesture and Skill Assessment Working Set (JIGSAWS) was released as first public dataset containing video and kinematic recordings of robotic surgical demonstrations, as well as gesture and skill annotations.

JIGSAWS contains synchronized video and kinematic data captured from the da Vinci Surgical System (dVSS) during the execution of elementary surgical tasks, including 39 suturing, 36 knot-tying and 20 needle-passing demonstrations (Fig. 6). All trials are performed on a training phantom by eight surgeons with different robotic surgical experience: experts reported more than 100 hours of training, intermediate between 10 and 100, and novice less than 10 .

Video data: Stereo videos from the da Vinci endoscopic cameras are recorded at 30Hz and at a resolution of 640 x 480. Demonstrations are all observed from the same point of view, as camera movement and zoom are completely absent. Calibration parameters are not provided.

Kinematic data: Kinematic trajectories of the two master tool manipulators (MTMs) and patient side manipulators (PSMs) are recorded in sync with the video data. The pose of each arm is represented by a local reference frame linked to the end-effector, and their motion is described by 19 kinematic variables that include Cartesian positions, a rotation matrix, linear velocities, angular velocities and a gripper angle. Kinematic features of different manipulators are all referred to a common coordinate system.

Annotations: A common dictionary comprised of 15 action classes is used to describe all three surgical tasks in the dataset. Ground truth annotations for each trial include the sequence of gestures, boundary timestamps and the demonstrator’s expertise level.

VII-B Experimental setup

Two standard cross-validation setups have been used to analyze performance and generalization capability of the available recognition systems:

Leave-one-supertrial-out (LOSO): five validation folds are created taking one different trial from each user. The LOSO scheme can be used to evaluate model generalization to new trials performed by known surgeons.

Leave-one-user-out (LOUO): eight validation folds are created taking all trials performed by the same user. The LOUO scheme can be used to evaluate model generalization to new and unknown surgeons, which is more challenging due to subject-specific style variability.

Notably, deep learning schemes typically limit extensive cross-validation due to high computational requirements and long training time and most recent studies have used only LOUO to discriminate different methods. LOUO is considered as the gold standard because it does not allow approaches to overfit to surgeon-specific features, which is also important for surgical skill assessment where networks should not learn to classify based on the surgeon but on the execution quality.

VII-C Evaluation metrics

Evaluation metrics for action recognition can be broadly divided into three categories highlighting different strengths and weaknesses of the proposed methods:

Frame-wise metrics: standard classification metrics such as average accuracy, precision, recall and F1-score, have been widely used to assess frame-wise recognition performance. These were however conceived for non-sequential data and are not sufficient for an exhaustive analysis of temporal predictions and robust method comparison. Predictions with similar accuracy might in fact show large qualitative differences in terms of action ordering and over-segmentation (i.e. prediction of many insignificant action boundaries).

Segmental metrics: the two most common segmental metrics, which have been specifically designed to evaluate the temporal structure of action predictions, are the edit score and the segmental F1 score with threshold k/100 (F1@k). The edit score evaluates action ordering but not their timing, penalizing out-of-order predictions and over-segmentation. It is computed as the Levenshtein distance between true and predicted label sequences, either normalized at the dataset level (Edit), where the common normalization factor is the maximum number of segments in any ground-truth sequence, or at the instance level (Edit*), where the normalization factor for each sequence is the maximum number of segments in either the ground-truth or the prediction . F1@k also penalizes over-segmentation, but it additionally evaluates the temporal overlap between predictions and ground truth segments with reduced sensitivity to slight temporal shifts, compensating for annotation noise around the segment boundaries. F1@k is obtained by computing the intersection over union (IoU) overlap score between each predicted segment and the corresponding ground truth segment of the same class. That prediction is considered as true positive (TP) if the IoU is above a threshold τ=k/100\tau=k/100, otherwise it is a false positive (FP). TPs and FPs are then used to compute the final F1 score . Normalized mutual information (NMI) has been used less frequently to measure alignment between two action sequences in unsupervised settings , as this score is not affected by permutations of cluster labels.

Unsupervised metrics: instead of comparing predictions to their ground truth, unsupervised metrics such as the silhouette score (SS) have been used to measure the compactness of action clusters generated by unsupervised recognition methods .

Other metrics used in the reviewed articles include Segmentation Accuracy (Seg-Acc) , Temporal Structure Score (TSS) and detection scores based on temporal tolerance windows . Formulas and pseudocode of the most common evaluation metrics are reported in Appendix A.

VII-D Results

Results reported on the suturing demonstrations of JIGSAWS are grouped into supervised graphical models (Table II), supervised deep learning (Table III), semi-supervised (Table V) and unsupervised (Table IV) methods. Despite referring to the same dataset, attention should be paid to the input data when comparing different methods, especially when the performance gap is limited. Improvements in segmental scores, for example, could be partially attributed to the adoption of lower sampling rate rather than to the method itself. Small accuracy gaps could also be linked to specific data pre-processing (filtering, normalization, dimensionality reduction) and kinematic and visual channel selection. Given the small dataset size and similarity between certain results, more rigorous comparative analysis could be supported by statistical testing to highlight significative performance gaps. Evaluation on larger surgical datasets should however be the main goal of future researches.

VIII Discussion and open questions

Automatic recognition of surgical gestures is a complex task and a number of considerations affect the current state of the art in available methods as detailed below.

Our comparative analysis on the JIGSAWS dataset revealed that recent deep learning models are able to capture complex temporal dependencies in surgical motion, leading to state-of-the-art recognition performance. In particular, hierarchical temporal convolutions and recurrent modules can model multi-scale temporal information for robust classification and boundary identification, especially when integrated with temporal attention mechanisms . Only a few studies, however, have considered the continuous nature of surgical motion and explicitly modelled gradual transitions between action segments . Transition handling has been alternatively addressed by under-penalizing classification errors to take into account uncertainties in ground truth annotations . Promising results have also been achieved with reinforcement learning, but related work in this area is still very limited and requires deeper analysis.

As for the lookahead time, the majority of reported methods relies on a variable number of future frames for robust recognition. While real-time performance or ability to predict future samples are not always needed and pursued, further efforts in this direction could broaden the applicability of such algorithms within the surgical theatre, supporting systems for surgical process monitoring and anomaly detection and prevention .

VIII-B Discriminative feature extraction

Robust analysis and understanding of fine-grained surgical gestures is especially complicated by the high variability of data with same action class. Gestures represent generic blocks of motion which can be shared across different tasks, phases and procedures, and significant variation comes from such context including the instrument types and the patient-specific environment. Other variability factors comprise the surgeon’s skill level, experience and individual style, which can significantly alter the duration, kinematics and order of actions in a subject-specific manner . For example, novice surgeons tend to make different mistakes and perform multiple gesture attempts and adjustment motions, while expert demonstrations are generally faster and smoother. Even among experts, differences might be notable in terms of surgical style, which are then reported among novice surgeons mimicking their tutors.

Similarity of data with different action class is also problematic because the amount of available information to discriminate between fine-grained gestures is limited. Considering the suturing task, for example, the same surgical tools and objects (a needle and a thread) are operated within the same environment during the whole task execution, conferring similar visual appearance to all the constituent segments (Fig. 4). Even the kinematic patterns of different gestures are sometimes very similar (e.g. when a surgeon attempts to push a needle through the tissue multiple times, the gestures “positioning needle tip on insertion point” and “pushing needle through the tissue” can be easily confused). Additional data such as tool usage information and video of the whole operating room, which can help to discriminate high-level surgical phases , are instead redundant for fine-grained analysis.

Discriminative feature extraction from surgical video and kinematics has thus been an important research focus area. Most of the reviewed methods, however, analyze all data samples independently or consider local neighborhoods in time. Only few studies integrate discriminative feature extraction and high-level temporal modelling through end-to-end training , thus embedding long-range temporal information in the low-level features. Nevertheless, end-to-end training of large architectures such as CNN-LSTMs, often applied to the analysis of high-level surgical workflow , is problematic on very small datasets like JIGSAWS.

Discriminative information could also derive from optical flow motion features, which have aided surgical phase recognition and action classification , or from semantic analysis of the video data . Semantic visual features could be potentially used to represent information about the surgical environment and manipulated objects that are not explicitly captured by abstract video representations or by the kinematic data, such as identification and localization of anatomical structures, measures of tissue deformation, pose or state of manipulated objects, presence of bleeding, smoke or other surgical tools operated by a surgical assistant. In addition, semantic information could also be gained indirectly through multi-task or transfer learning with appropriate auxiliary tasks. Multi-task learning has been extensively explored at coarse granularity levels, where systems detecting surgical tools or predicting the remaining time of surgery have also shown improved phase recognition capabilities . Gesture recognition networks have only recently been enhanced with parallel branches for progress , skill scores or surgical tool trajectory prediction, and other auxiliary tasks could potentially be conceived to aid recognition at fine granularity (e.g. tasks based on semantic image segmentation, object tracking or anomaly detection).

VIII-C Multi-modal data integration

The integration of multi-modal data could also play an important role in improving the performance of available recognition systems, as different sensor modalities often measure complementary information. Kinematic data, for example, describe poses and velocities of the surgical tools in the three-dimensional space, whereas visual features carry implicit complementary information about the surgical environment and instrument-tissue interactions. The combination of kinematic and visual features has in fact yielded consistent improvement over the two individual modalities . Force and torque signatures have also been used for motion decomposition and skill assessment , even though this information may not be readily available from system manufacturers. This lack of data access has also meant that multi-modal data integration for surgeme recognition has not been investigated thoroughly.

Kinematic and visual information have been combined at the input level , at intermediate level or in the final prediction and pruning stages . Concatenation of multi-modal data at low processing levels might be suboptimal and inefficient, as the two sensing modalities are not directly comparable due to differences in cardinality, semantics and stochasticity (e.g. visual features are typically much larger than kinematic samples and more noisy). A splicing weight between kinematic and visual features has been used to find appropriate importance levels in different surgical tasks , but it entails extensive tuning and can not adapt dynamically to environmental changes. Class-specific weights for multi-modal late fusion are also unable to capture intra-gesture dynamic properties. Improved integration could be achieved with flexible systems that seek timely multi-modal cooperation at multiple processing levels, with adaptive weighting or knowledge transfer from the strongest to the weakest modality.

VIII-D Translational research

Despite recent advances, the field and applicability to real surgical scenarios have been greatly hindered by the lack of large and diverse datasets of annotated demonstrations, which are essential for robust training of modern recognition systems based on machine learning and deep learning. Current methods need testing on varied surgical tasks and procedures, as well as on data from real interventions with complex environments, blood, specularities, camera motions, illumination changes, occlusions and higher variability in motion and gesture ordering . International collaborations for clinical data collection and sharing would not only accelerate the generation of such datasets, but also allow the representation of variations in human physiology linked to ethnicity as well as a wider spectrum of surgical techniques. Development of a common fine-grained surgical language with standardized decomposition criteria for different procedures and target anatomical structures is then needed in order to integrate information from multiple data sources.

It is however not trivial to define a comprehensive gesture dictionary and meaningful segmentation criteria which are consistent in the presence of unconstrained movement and generalizable to a variety of anatomies and surgical styles. The characterization of “surgical gesture” is in fact an open problem itself, in particular for bimanual operations. Datasets like JIGSAWS regard fine-grained gestures as short action segments performed by either of the two arms, while the motion of the other arm is ignored. This strategy allows for faster labelling, but suffers from ambiguities on complex demonstrations where the two arms might cooperate or perform different actions at the same time. Such ambiguity can be resolved with finer-grained analysis and multi-label segmentation with separate classification of right and left arm motions. Multi-label analysis is burdensome but more precise and can additionally account for action combinations .

The optimal decomposition strategy could also depend on application requirements. A robotic collaborative assistant might not need to recognize highly detailed surgical motions and localize precise action boundaries in order to provide the surgeon with timely support. On the other side, robust understanding of task workflow, bimanual cooperation and action transitions are required to achieve surgical automation and successfully execute complex action sequences in real case scenarios. As for surgical skill assessment, the proper decomposition strategy could depend on the annotators’ evaluation technique. Identification of known skill-related gestures and failure modes while ignoring style-related variability might be beneficial.

Even when the action dictionary is well defined, however, manual labelling of surgical gestures is extremely costly, time consuming, prone to errors and inconsistencies. To mitigate data requirements, the recognition problem can be tackled in a semi-supervised or unsupervised manner, where labels are only necessary for model testing. Even if not yet able to outperform supervised methods, these have a lot of potential to capitalise on large amounts of unlabelled data .

When collection of clinical records is problematic, one of the most viable strategies to generate large-scale surgical data is represented by virtual reality simulators , which can nowadays reproduce real surgical environments with impressive realism, as well as provide complementary information (e.g. surgical tool kinematics or semantic segmentation of the 3D scene) to aid recognition. Further efforts could then be directed towards new strategies for knowledge transfer between surgical environments (simulation vs real surgery), as well as between surgical tasks or surgical robots , in order to exploit cross-domain information to improve generalization performance.

Finally, an emerging and closely related research sub-field is represented by the detection and forecast of surgical errors in fine-grained gestures. It has been recently shown that common gesture-specific errors can be identified in real-time if provided with corresponding real-time gesture labels . Public availability of error annotations will hopefully stimulate future research in this relatively unexplored field, which is fundamental for several CAI implementations including workflow monitoring and surgical automation.

IX Conclusion

In this paper, we have presented a comprehensive analysis and critical review on the recognition of surgical gestures from intraoperatively collected information. Automatic recognition of surgical gestures is challenging due to the variability, complexity and fine-grained structure of surgical motion. The digitization of surgical video and increased availability of sensor information in robotic instrumentation has driven the emergence and evolution data-driven recognition models. While increasing the robustness of automatic recognition, deep-learning-based models still face many challenges compounded by the limited availability and small size of surgical training datasets that do no capture the domain variability of clinical practice across diverse and sometimes procedure specific activities. The field would also benefit from standardized definition of the segmentation criteria for different applications and surgical procedures to support comparative analysis and research across multiple datasets. Technically, unsupervised and semi-supervised approaches, which do not require large amounts of ground truth labels, are important directions for future research in order to enable scaling of methods across surgical procedures and allow adaptability to the evolution of surgical technique and instrumentation.

Appendix A Formulas of the most common evaluation metrics

Accuracy (Acc): Given a sequence of length N,

where NcNc is the number of correctly labelled frames.

Edit score (Edit): Given a sequence of labels τ\tau and corresponding ground truth γ\gamma:

where DD is the Levenshtein distance between τ\tau and γ\gamma, NN is either the maximum number of segments in τ\tau or γ\gamma (instance level normalization), or in any ground truth sequence (dataset level normalization).

Segmental F1 score with threshold k/100 (F1@k): Given a sequence of predicted segments PSPS and corresponding ground truth segments TSTS:

TPTP = FPFP = NusedN_{used} = for each predicted segment PSjPS_{j} in PSPS: cc = label of PSjPS_{j} TScTS^{c} = list of TSTS segments of class cc IoUallIoU_{all} = IoU(PSj,TSc)IoU(PS_{j},TS^{c}) IoUiIoU_{i} = max(IoUall)max(IoU_{all}) TSicTS^{c}_{i} = segment in TScTS^{c} corresponding to max(IoUall)max(IoU_{all}) if IoUi>k/100IoU_{i}>k/100 and TSicTS^{c}_{i} not usedused: TPTP = TP+1TP+1 Flag TSicTS^{c}_{i} as usedused NusedN_{used} = NusedN_{used} + 1 else: FPFP = FP+1FP+1 FNFN = length(TS)length(TS) - NusedN_{used} precisionprecision = TP/(TP+FP)TP/(TP+FP) recallrecall = TP/(TP+FN)TP/(TP+FN) F1@kF1@k = 2∗(precision∗recall)/(precision+recall)∗1002*(precision*recall)/(precision+recall)*100

where IoUIoU is the intersection over union overlap score, TPTP are true positives, FPFP false positives and FNFN false negatives. Detailed implementation is publicly available .

Normalized Mutual Information (NMI): Given two sequences of labels τ\tau and γ\gamma:

where II is the mutual information and HH the entropy.

Silhouette Score (SS): The SS for a single data sample ii is defined as:

where a(i)a(i) is the distance (e.g. Euclidean distance) between that sample and the mean of its own cluster, while b(i)b(i) is the distance between that sample and the mean of the nearest cluster it is not a part of. SSSS is the average Silhouette score over all Ns samples:

References