How to process extremely large (> 100TBs) data sets without burning millions
I first worked with >100 TB pipelines when I joined the Core Growth team at Facebook in 2016 back when it was the best place to work. The first three months were filled with ice cream, bike rides, and lots of fun! Then suddenly my dream paradise turned intense when my boss, Jitender told me, I was to own the push, email, and SMS notification data for all of Facebook. I was excited about this challenge when I learned: EcZachly Data Engineering Newsletter is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Facebook sends 50 BILLION NOTIFICATIONS every single day Facebook has five channels they send notifications through Jewel, Push, Email, SMS, and Logged Out Push We need low latency for optimal machine learning The notification filtering machine learning degraded significantly if it was delayed > 1 day. It’s a complex space that balances between spammy and engaging! If you prefer this in video form, here’s a viral vi
Explore this link on the map →