DISK: Learning local features with policy gradient
Michał J. Tyszkiewicz, Pascal Fua, Eduard Trulls
Introduction
Related work
Related Work
The process of extracting local features usually involves three steps: finding a keypoint, estimating its orientation, and computing a description vector. In traditional methods such as SIFT [Lowe04] or SURF [Bay08], this involves many hand-crafted heuristics. The first wave of local features involving deep networks featured descriptors learned from patches extracted on SIFT keypoints [Zagoruyko15, Han15, Simo-Serra15] and some of their successors, such as HardNet [Mishchuk17], SOSNet [sosnet], and LogPolarDesc [Ebel19], are still state-of-the-art. Other learning-based methods focus on keypoints [Verdie15, Savinov17, keynet] or orientations [Yi16a], \mtor merge the two notions entirely [implicitly-matched].
These methods attack a single element of this process. Others have developed end-to-end-trainable pipelines [Yi16b, superpoint, Ono18, Dusmanu19, Revaud19] that can optimize the whole process and, hopefully, improve performance. However, they either use inexact approximations to the true objective [superpoint, Revaud19], break differentiability [Ono18] or make big assumptions, such as extrema in descriptor space making good features [Dusmanu19].
Three recent approaches are attempting to bridge the gap between training and inference in a spirit close to ours. GLAMpoints [Truong19a] seeks to estimate homographies between retinal images and use Reinforcement Learning (RL) methods to find keypoints that are correctly matched by SIFT descriptors. Since matching is deterministic, Q-learning can be used to regress for the expected reward of each keypoint, rather than optimize directly in policy space. Using hand-crafted descriptors and only addressing the detection problem was motivated by domain-specific requirements of strong rotation equivariance, which most learned models lack. While it makes sense in the specific scenario it was developed for, it limits what the method can do. \mtSimilarly, [succinct-interest-points] also uses handcrafted descriptors and learns to predict the probability that each pixel would be successfully matched with those. Their approach therefore inherits many of the limitations of GLAMpoints.
Reinforced Feature Points [reinforced-feature-points] address the more difficult issue of learning with a general non-differentiable objective for the purpose of camera pose estimation, with RANSAC in the loop. Unfortunately, supervising all detection and matching decisions with a single reward means that this approach suffers from weak training signal, an endemic RL problem, and has to rely on pre-trained models from [superpoint] that can only be fine-tuned. Our method can be seen as a relaxation of their approach, where we train for a surrogate objective: finding many correct feature matches. This allows for substantially more robust training from scratch and yields better downstream results.