k8s-1m Overview
Several years ago at OpenAI I helped author Scaling Kubernetes to 7500 Nodes which remains one of the CNCF’s most popular blog posts. Alibaba made a post about running Kubernetes clusters with 10K nodes. Google made a post about 15K nodes with Bayer Crop Science. Fast forward to today, GKE supports running some clusters up to 65K nodes, and AWS recently announced support for clusters up to 100K nodes. In online forums and in my own conversations with peers, I’ve encountered a lot of debate about how big a Kubernetes cluster can get. What tends to be lacking from these discussions is hard data and evidence-backed justifications. I’ve worked with engineers reluctant to push things beyond what they’ve seen before because they’re fearful or uncertain of what may go wrong. Or when something does go wrong, the response is to scale down the cluster rather than understand and address the bottleneck. The spirit of the k8s-1m project is to identify the hard blockers to scalability. What are the
k8s-1m Overview k8s-1m Overview This is an effort to create a fully functional Kubernetes cluster with 1 million active nodes. Table of Contents Why? Components Networking Pod IPs IPv4-only external service dependencies Network Policies Network flow needs (# of TCP connections) Managing state kube-apiservers vs etcd Meeting the QPS needs for a 1M node cluster etcd is too slow Reduce durability and eliminate replicas Reduce the interface Txn-Put mem_etcd: custom in-memory etcd Watch() Watches per node Update() Caching and locking Garbage collection Scheduler Basic design: shard on nodes The pai
Explore this link on the map →saved by
related reading
- Cluster Architecture | Kuberneteskubernetes.io
- Deploying R with kubernetes – Notes from a data witchblog.djnavarro.net
- Kubernetes production readiness checklistlearnk8s.io
- The feedback loops behind Kubernetes — PlanetScaleplanetscale.com
- Best practices for right-sizing your Apache Kafka clusters to optimize performance and cost | AWS Big Data Blogaws.amazon.com
- How to Rack 30 Petabytes of Storage | blogsi.inc
- Paul Butler – The hater’s guide to Kubernetespaulbutler.org
- Kubernetes CPU Throttling | Temporaltemporal.io
- Carving The Scheduler Out Of Our Orchestrator · The Fly Blogfly.io
- What we learned after running Airflow on Kubernetes for 2 years | by Alexandre Magno Lima Martins | Apache Airflow | Mediummedium.com
- We rebuilt the Kubernetes control loop on Postgres | Atlasflowatlasflow.com
- Production Twitter on One Machine? 100Gbps NICs and NVMe are fast - Tristan Humethume.ca