dattri-LLM: A Unified and Efficient Library for Training Data Attribution at LLM Scale
dattri-LLM is a new library that addresses the challenges of training data attribution at large language model scale by improving efficiency and compatibility.
dattri-LLM is a new library that addresses the challenges of training data attribution at large language model scale by improving efficiency and compatibility.
Most scalable TDA methods rely on per-example gradients, whose computation and use at LLM scale pose challenges in efficiency, compatibility, and extensibility.
Abstract: Training data attribution (TDA) estimates the contribution of individual training examples to model outputs. Most scalable TDA methods rely on per-ex…