Model Training Blocking Rules - Splink
Model Training Blocking Rules choose which record pairs from a dataset get considered when training a Splink model. These are used during Expectation Maximisation (EM), where we estimate the m probability (in most cases). The aim of Model Training Blocking Rules is to reduce the number of record pairs considered when training a Splink model in order to reduce the computational resource required. Each Training Blocking Rule define a training "block" of records which have a combination of matches and non-matches that are considered by Splink's Expectation Maximisation algorithm. The Expectation Maximisation algorithm seems to work best when the pairwise record comparisons are a mix of anywhere between around 0.1% and 99.9% true matches. It works less efficiently if there is a huge imbalance between the two (e.g. a billion non matches and only a hundred matches). Note Unlike Prediction Rules, it does not matter if Training Rules excludes some true matches - it just needs to generate examp
Blocking for Model Training ¶ Model Training Blocking Rules choose which record pairs from a dataset get considered when training a Splink model. These are used during Expectation Maximisation (EM), where we estimate the m probability (in most cases). The aim of Model Training Blocking Rules is to reduce the number of record pairs considered when training a Splink model in order to reduce the computational resource required. Each Training Blocking Rule define a training "block" of records which have a combination of matches and non-matches that are considered by Splink's Expectation Maximisati
Explore this link on the map →related reading
- What are Blocking Rules? - Splinkmoj-analytical-services.github.io
- 4. Estimating model parameters - Splinkmoj-analytical-services.github.io
- 3. Blocking - Splinkmoj-analytical-services.github.io
- 3. Blocking - Splinkmoj-analytical-services.github.io
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- GitHub - google-research/tuning_playbook: A playbook for systematically maximizing the performance of deep learning models. · GitHubgithub.com
- Defining and customising comparisons - Splinkmoj-analytical-services.github.io
- Go smol or go home | Harm de Vriesharmdevries.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Competing with sampling — Alignment Research Centeralignment.org
- Fermi estimate of future training runsdanieldewey.net