How to Scale K-Means Clustering with just ClickHouse SQL
Recently, when helping a user who wanted to compute centroids from vectors held in ClickHouse, we realized that the same solution could be used to implement K-Means clustering. They wanted to solve this at scale across potentially billions of data points while ensuring memory could be tightly managed. In this post, we give implementing K-means clustering using just SQL a try and show that it scales to billions of rows. In the writing of this blog, we became aware of the work performed by Boris Tyshkevich. While we use a different approach in this blog, we would like to recognize Boris for his work and for having this idea well before we did! As part of implementing K-Means with ClickHouse SQL, we cluster 170M NYC taxi rides in under 3 minutes. The equivalent scikit-learn operation with the same resources takes over 100 minutes and requires 90GB of RAM. With no memory limitations and ClickHouse automatically distributing the computation, we show that ClickHouse can accelerate machine le
Introduction # Recently, when helping a user who wanted to compute centroids from vectors held in ClickHouse, we realized that the same solution could be used to implement K-Means clustering. They wanted to solve this at scale across potentially billions of data points while ensuring memory could be tightly managed. In this post, we give implementing K-means clustering using just SQL a try and show that it scales to billions of rows. In the writing of this blog, we became aware of the work performed by Boris Tyshkevich. While we use a different approach in this blog, we would like to recognize
Explore this link on the map →related reading
- k-means clustering - Wikipediaen.wikipedia.org
- CS 1110 Fall 2024cs.cornell.edu
- Clustering Algorithms: K-Means, EMC and Affinity Propagation | Toptal®toptal.com
- ClickHouse vs. Elasticsearch: The Mechanics of Count Aggregations | ClickHouseclickhouse.com
- kMeans: Initialization Strategies- kmeans++, Forgy, Random Partition | Analytics Vidhyamedium.com
- How to Determine the Optimal K for K-Means? | by Khyati Mahendru | Analytics Vidhya | Mediummedium.com
- k8s-1m Overviewbchess.github.io
- In-depth: ClickHouse vs Snowflake - PostHogposthog.com
- How to cluster images based on visual similarity | Towards Data Sciencetowardsdatascience.com
- How to Rack 30 Petabytes of Storage | blogsi.inc
- ClickHouse - Wikipediaen.wikipedia.org
- Hierarchical clustering - Wikipediaen.wikipedia.org