Why data scientists shouldn’t need to know Kubernetes
Recently, there have been many heated discussions on what the job of a data scientist should entail [1, 2, 3]. Many companies expect data scientists to be fu...
[ Hacker News discussion , Twitter thread ] Recently, there have been many heated discussions on what the job of a data scientist should entail [ 1 , 2 , 3 ]. Many companies expect data scientists to be full-stack, which includes knowing lower-level infrastructure tools such as Kubernetes (K8s) and resource management. This post is to argue that while it’s good for data scientists to own the entire stack, they can do so without having to know K8s if they leverage a good infrastructure abstraction tool that allows them to focus on actual data science instead of getting YAML files to work . The
saved by
related reading
- Streamlit • A faster way to build and share data appsstreamlit.io
- What I Learned From The Modern Data Stack Conference 2021 - James Lejameskle.com
- Scaling App Infrastructure with Kubernetes & Microservicesrtinsights.com
- Sweatshop data is overmechanize.work
- From Data Engineer to YAML Engineerjuhache.substack.com
- What we learned after running Airflow on Kubernetes for 2 years | by Alexandre Magno Lima Martins | Apache Airflow | Mediummedium.com
- Airflow Architecture: Key Components & Best Practiceshevodata.com
- Prefect - Workflow Orchestration for Data, ML, and Agentsprefect.io
- The Revenge of the Data Scientist – Hamel’s Blog - Hamel Husainhamel.dev
- Designing Data Science Tools at Spotify: Part 2medium.com
- Kubernetestheswissbay.ch
- Reflections On Data Science In Big Tech | Varun's blogvarunprajan.github.io