TRAK: Attributing Model Behavior at Scale

Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, Aleksander Madry

Introduction

Training data is a key driver of model behavior in modern machine learning systems. Indeed, model errors, biases, and capabilities can all stem from the training data [ilyas2019adversarial, gu2017badnets, geirhos2019imagenet]. Furthermore, improving the quality of training data generally improves the performance of the resulting models [huh2016makes, lee2022deduplicating]. The importance of training data to model behavior has motivated extensive work on data attribution, i.e., the task of tracing model predictions back to the training examples that informed these predictions. Recent work demonstrates, in particular, the utility of data attribution methods in applications such as explaining predictions [koh2017understanding, ilyas2022datamodels], debugging model behavior [kong2022resolving, shah2022modeldiff], assigning data valuations [ghorbani2019data, jia2019towards], detecting poisoned or mislabeled data [lin2022measuring, hammoudeh2022identifying], and curating data [khanna2019interpreting, liu2021influence, jia2021scalability].

However, a recurring tradeoff in the space of data attribution methods is that of computational demand versus efficacy. On the one hand, methods such as influence approximation [koh2017understanding, schioppa2022scaling] or gradient agreement scoring [pruthi2020estimating] are computationally attractive but can be unreliable in non-convex settings [basu2021influence, ilyas2022datamodels, akyurek2022towards]. On the other hand, sampling-based methods such as empirical influence functions [feldman2020neural], Shapley value estimators [ghorbani2019data, jia2019towards] or datamodels [ilyas2022datamodels] are more successful at accurately attributing predictions to training data but require training thousands (or tens of thousands) of models to be effective. We thus ask:

Are there data attribution methods that are both scalable and effective in large-scale non-convex settings?