In this talk, we will present an in-house system we built for serving large datasets generated through Hadoop/Hive jobs. The system has been in production for several months and is serving Greater than 50T of data across numerous use cases with over 1M cumulative qps. A large chunk of these use cases serve our online user traffic at low latencies. We will present the system architecture in detail and discuss why we chose to not use off the shelf solutions.
Presenter:
Varun Sharma