flâneur — a map of the web's best reading

II: Theoretical foundations

ml-data-tutorial.org · 2,515 words · saved by 1 readers

Recall from the first chapter that our focus is going to be a task called “datamodeling” or “predictive data attribution,” where our goal is to predict how models will behave when trained on different training datasets. It turns out that in classical statistics, this problem has been studied extensively. Predicting such data counterfactuals is of significant interest in, for example, cross-validation, the bootstrap, etc. Why “predictive data attribution”/“datamodeling”? The problem of predicting model changes as a function of training dataset changes is an age-old problem, especially in statistics: names like “influence functions” (Hampel 1974), “von Mises calculus” (von Mises 1947), “regression analysis” (Pregibon 1981), and “infinitesimal jackknife” (Jaeckel 1972), all refer to this problem (or variants thereof, or methods for solving it). We use the names “predictive data attribution” or “datamodeling” as a catchall for all of these terms, emphasizing the connection to data attribut

II: Theoretical foundations July 18, 2024 --> Recall from the first chapter that our focus is going to be a task called “datamodeling” or “predictive data attribution,” where our goal is to predict how models will behave when trained on different training datasets. It turns out that in classical statistics, this problem has been studied extensively . Predicting such data counterfactuals is of significant interest in, for example, cross-validation, the bootstrap, etc. Why “predictive data attribution”/“datamodeling”? The problem of predicting model changes as a function of training dataset chan

Explore this link on the map →

related reading