Choosing comparators and thresholds - Splink
When building a Splink model, one of the most important aspects is defining the Comparisons and Comparison Levels that the model will train on. Each Comparison Level within a Comparison should contain a different amount of evidence that two records are a match, which the model can assign a Match Weight to. When considering different amounts of evidence for the model, it is helpful to explore fuzzy matching as a way of distinguishing strings that are similar, but not the same, as one another. This guide is intended to show how Splink's string comparators perform in different situations in order to help choosing the most appropriate comparator for a given column as well as the most appropriate threshold (or thresholds). For descriptions and examples of each string comparators available in Splink, see the dedicated topic guide. There are three main classes of string comparator that are considered within Splink: where String Similarity Scores are scores between 0 and 1 indicating how simil
Choosing String Comparators ¶ When building a Splink model, one of the most important aspects is defining the Comparisons and Comparison Levels that the model will train on. Each Comparison Level within a Comparison should contain a different amount of evidence that two records are a match, to which the model can assign a match weight. When considering different amounts of evidence for the model, it is helpful to explore fuzzy matching as a way of distinguishing strings that are similar, but not the same, as one another. This guide is intended to show how Splink's string comparators perform in
Explore this link on the map →related reading
- Defining and customising comparisons - Splinkmoj-analytical-services.github.io
- 4. Estimating model parameters - Splinkmoj-analytical-services.github.io
- ML vs. Score matchingarxiv.org
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- score matching (hyvarinen)jmlr.org
- Levenshtein distance - Wikipediaen.wikipedia.org
- 3. Blocking - Splinkmoj-analytical-services.github.io
- 3. Blocking - Splinkmoj-analytical-services.github.io
- Regularized by Score Matching (LeCun)papers.nips.cc
- GitHub - open-compass/opencompass: OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets. · GitHubgithub.com
- GitHub - ansonyuu/matchmaking: Embedding space of names clustered based on their interests using the sentence-transformers all-MiniLM-L6-v2 model · GitHubgithub.com
- What are Blocking Rules? - Splinkmoj-analytical-services.github.io